StarDrinks Test Set

An English and Korean test set for evaluating LLMs and speech assistants in a drink ordering scenario

About

StarDrinks is a test set in English and Korean for evaluating LLMs and speech assistants in a realistic drink ordering scenario. Task-oriented systems are often assessed under controlled conditions that fail to capture the variability of real user requests: diverse named entities, drink types, sizes, customizations, and brand-specific terminology, as well as spontaneous speech phenomena such as hesitations and self-corrections.

The dataset contains speech utterance features, transcriptions, and annotated slots, and supports three evaluation settings: speech-to-slots (SLU), transcription-to-slots (NLU), and speech-to-transcription (ASR). Collected from authentic drink orders, it provides a linguistically rich benchmark for measuring model robustness and generalization to previously unseen named entities.

Downloading the data

Access requires agreeing to a custom license: please sign the license agreement available on the dataset page above before downloading.

Citing us

When using our dataset, please cite the following paper:

@inproceedings{zanon-boito-etal-2026-stardrinks,
    title = "{S}tar{D}rinks: An {E}nglish and {K}orean Test Set for {SLU} Evaluation in a Drink Ordering Scenario",
    author = "Zanon Boito, Marcely  and
      Brun, Caroline  and
      Kim, Inyoung  and
      Proux, Denys  and
      Ait-Mokhtar, Salah  and
      Lagos, Nikolaos  and
      Meunier, Jean-Luc  and
      Calapodescu, Ioan",
    booktitle = "Proceedings of the 2026 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC)",
    month = may,
    year = "2026",
    address = "Palma de Mallorca, Spain",
    publisher = "European Language Resources Association",
    url = "https://arxiv.org/abs/2604.26500",
}