How we talk: Generating Singaporean conversations from scratch

10 minute read

Published:

Reposted 25 August 2026. The original blog can be found here

Large language models (LLMs) are incredibly powerful — but they’re only as useful as the data they’re trained on. That’s why we need relevant, high-quality, robust data that we can fine-tune LLMs with, so that these LLMs can understand how we ask questions in our daily lives and respond appropriately.

In our case, we want to train models that could engage meaningfully in conversations with Singaporeans from different walks of life, grounded in the information and culture of our Home Team. But where would the data come from? We have to create our own.

Here, we share how we built our own conversational dataset in the Home Team context—starting with web scraping, then layering on realistic problem statements and personas to simulate Q&A interactions between Singaporeans.

What’s the context?

We begin by scoping the data available to us — a pool of information that we can tap on to provide context for the interactions — which would have to be factual, objective, and credible domain-specific information about the Ministry of Home Affairs (MHA), and the different Home Team departments and agencies.

We obtained this by scraping the open-sourced websites of MHA, Singapore Police Force (SPF), Singapore Civil Defence Force (SCDF), and Home Team Science and Technology Agency (HTX). This allowed us to gather text data from different types of material — newsroom articles, interviews, advisories, e-services, annual reports, etc. — that covers a wide range of information that a member of public might find useful in their understanding of what we do as One Home Team.

Source websites scraped for context

We processed and cleaned the raw scraped data, which we then divided into chunks of ~2000 tokens and uploaded to a vector database.

What would they want to know?

Now that we have the context, we need to figure out what the public might be curious about, regarding what the Home Team does.

We first sampled a subset of the scraped data to ensure that the various sections of the departments’ websites are well-represented. This subset was then used to generate a series of questions that a member of public might ask with regard to the given context. Here’s a sneak peek of the prompt given:

…Based on the retrieved content provided below, identify a key challenge, issue, or consequence of an action related to the topic that a member of public might face. Craft a concise and well-structured problem statement that highlights the core issue in a clear and actionable manner…

After generating the questions, we also had to validate their quality. Given the large number of questions, we employed an initial prompt-based checker with GPT-4o before moving on to manual checking.

Question generation pipeline

This process allowed us to generate 2373 questions based on various contexts across the four departments and agencies. Here are some examples of the questions:

  • How can I explore volunteer opportunities with the Home Team to contribute to Singapore’s safety and security?
  • What is the Rescue 995 microsite, and what type of content does it provide?

Who would they be?

Now we have the context and the questions. But as we all know, our colloquial variety of English — also known as Singlish — is an important part of the Singaporean diaspora, and we have to account for this variation when preparing a dataset that can help the model generalise well to our unique speech patterns.

Persona generation pipeline

We turned to Persona Hub to account for this uniqueness (as well as to scale the dataset). The idea is that by attaching a ‘persona’ to the prompt when we generate the conversations (think: roleplay), we are able to obtain variations in language use and different perspectives, even when the conversation is about the same topic.

In order to obtain Singapore-specific personas, we filtered through the billion personas on Persona Hub on a set of criteria:

  • Is the persona relevant to Singapore? (e.g., Singaporean, persona living in Singapore, etc.)
  • Is the persona non-human? → if not, modify to human persona
  • Is the persona region specific? (e.g. man living in England…) → modify to be Singapore specific

We also explored generating additional personas using the demographic metadata of the speakers from the National Speech Corpus, such as age or occupation. The data was also helpful in augmenting some of the generic personas found in Persona Hub. Diversity of the set was also enhanced by comparing persona embeddings and selecting semantically distant personas.

Persona distribution analysis

By generating and manually evaluating sample conversations with different personas of various demographics, we identified the education and occupation of a given persona to be the main contributor in driving the overall tone of the conversation. With this information, we also re-sampled under-represented demographic groups, such as those without tertiary education. Some of the personas generated are:

  • A Singaporean tourism strategist with a bachelor’s degree in tourism management, leveraging augmented reality to enhance visitor experiences in Singapore’s attractions.
  • A young aspiring Singaporean musician with a diploma in music, inspired by the history and influence of protest songs.
  • A data scientist with a bachelor’s degree in computer science, working on predictive analytics for Singapore’s traffic congestion issues.

We finalised our list with a persona pool of 300.

How would they talk?

Now we have the problem statements and the personas, and we are left with the task of actually generating the conversation between the persona and the assistant, based on the topic given in the problem statement. Cross-multiplying these two lists helped us to scale our dataset — we are looking at a maximum of 300 personas × 2373 problem statements = 711,900 unique conversations to be generated.

However, scaling while maintaining data quality is not as easy as it sounds. We still faced challenges in making sure that each generated conversation is as realistic and diverse as possible.

We started off with providing a problem statement and the relevant text chunk, as well as a persona, and generating a three-turn conversation — one turn consisting of one question from the persona and one answer from the assistant — per API call with GPT-4o. This vanilla approach comes with a number of issues that we have addressed, as detailed in the section below.

Issue #1: Repeated conversation flow

Conversations with the same problem statement tend to follow the same flow and coverage of topics, despite the text chunk having different angles or topics to expand on.

We addressed this by also providing the history of five previously generated conversations as references during generation. We also introduced random shuffling of persona set and a running average of embedding similarity, for which a threshold can be specified to stop generation once exceeded.

Issue #2: Hallucinations

There can be minor hallucinations, with facts or claims that were not directly supported by the text chunk.

In addition to prompting for strict adherence to the content of the text chunk, we also employed a first-cut hallucination checker with GPT-4o where we prompted the model to check the assistant’s answers against the respective text chunks, followed by human evaluation for hallucinated content.

Issue #3: Prompt adherence

Lastly, as with most generative models, there were instances where the output does not adhere strictly to the instructions given in the prompt:

  • Concluding statements: Since the prompt specifies a three-turn interaction, GPT-4o often treats the final turn as a concluding turn, where the persona or the assistant would wrap the conversation up with a comment or summary, which was unnatural.
  • Speech style: Assistant mimicked the persona’s tone and speaking style, which sometimes led to unprofessional responses. This is especially noticeable when the assistant aligns too closely with personas that have distinct or informal language patterns.
  • Referencing ‘text chunk’: Assistant would sometimes make reference to the presence of a ‘text chunk’ ad verbatim, revealing its awareness of the setup and breaking the illusion of a natural conversation.

The above issues have been minimised via iterations of prompt editing and human evaluation. We also added greater variability by uniformly sampling the number of turns per conversation from a range of 2 to 5. After dropping or regenerating conversations that did not meet our criteria, we finalised our dataset with 529,924 conversations in total. Here’s one of the conversations:

User: Eh, u know if ICA got use any new machins or someting to make checkpoint faster? Now always q so long, sian leh.

Assistant: Yes, ICA has implemented the New Clearance Concept which incorporates AI-driven systems to verify the identities of travelers. This initiative allows travelers to pass through immigration checkpoints more efficiently, thereby reducing queue times and enhancing overall clearance speed.

User: Wah like dat can reali save time lor. Bt AI how it check identity ah? Safe or not?

Assistant: AI systems used in the ICA’s New Clearance Concept are designed to verify identities through advanced algorithms, ensuring accuracy and security. These systems enhance the reliability of identity checks without compromising safety, as they operate with stringent protocols to protect personal information and prevent misuse.

User: Oh ok. Den HTX got use similar ting for other areas ah? Or only ICA side?

Assistant: HTX uses AI across various Home Team projects beyond ICA. For example, they leverage AI for resource optimization, reducing response times, and building tools like Natural Language Processing systems and automated speech recognition designed to accommodate Singapore’s multilingual colloquialisms. These initiatives collectively enhance efficiency and capabilities across the Home Team.

What’s next?

There are many considerations to ensure the quality of a synthetic dataset, especially when we want to generate them at scale. We are met with many learning points throughout this journey, and are again reminded of the challenges of creating novel, quality data in this increasingly data-scarce landscape.

This is by no means the ‘perfect dataset’ — the context we have is up-to-date only to the point of data extraction, and we would have to revise them from time to time with new initiatives and technologies from the Home Team.

As we move forward with using this dataset for fine-tuning and evaluating conversational LLMs for the Home Team context, we hope to leave you with food for thought on how best to synthesise your own datasets as well.


Many thanks to our interns — Clarabelle, Rachel, and Hazel for data collection; Brian, Caleb, Neleh for doing the heavy lifting with the personas and problem statements; Amber and Yuan Qi for cleaning up the dataset. Kudos to Shisheng, Jason, and Calvin for insightful discussions in developing this dataset too!