StarDrinks Test Set
An English and Korean test set for evaluating LLMs and speech assistants in a drink ordering scenario
About
StarDrinks is a test set in English and Korean for evaluating LLMs and speech assistants in a realistic drink ordering scenario. Task-oriented systems are often assessed under controlled conditions that fail to capture the variability of real user requests: diverse named entities, drink types, sizes, customizations, and brand-specific terminology, as well as spontaneous speech phenomena such as hesitations and self-corrections.
The dataset contains speech utterance features, transcriptions, and annotated slots, and supports three evaluation settings: speech-to-slots (SLU), transcription-to-slots (NLU), and speech-to-transcription (ASR). Collected from authentic drink orders, it provides a linguistically rich benchmark for measuring model robustness and generalization to previously unseen named entities.
Downloading the data
- Dataset: NAVER LABS Europe
Access requires agreeing to a custom license: please sign the license agreement available on the dataset page above before downloading.
Citing us
When using our dataset, please cite the following paper:
@inproceedings{zanon-boito-etal-2026-stardrinks,
title = "{S}tar{D}rinks: An {E}nglish and {K}orean Test Set for {SLU} Evaluation in a Drink Ordering Scenario",
author = "Zanon Boito, Marcely and
Brun, Caroline and
Kim, Inyoung and
Proux, Denys and
Ait-Mokhtar, Salah and
Lagos, Nikolaos and
Meunier, Jean-Luc and
Calapodescu, Ioan",
booktitle = "Proceedings of the 2026 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC)",
month = may,
year = "2026",
address = "Palma de Mallorca, Spain",
publisher = "European Language Resources Association",
url = "https://arxiv.org/abs/2604.26500",
}