最近,像豆包、千问 Omni等新一代语音对话模型的惊艳表现引发了广泛关注,它们展示了极其流畅的 语音交互能力。另一方面,具备“听、看、说”的智能眼镜也犹如雨后春笋般的越来越普及。然而,这些模型大多仍依赖于较为理想的使用环境。戴上 AI 智能眼镜,“行走的大模型”的“自然交互”还能否经受住真实世界的考验?
目前的语音处理系统在面对智能眼镜带来的独特挑战时,依然面临瓶颈:智能眼镜随佩戴者穿梭于办公室和嘈杂街头等高度动态的声学环境中;在真实的社交场景下,系统不仅要应对第一视角下的复杂非稳态噪声,更要处理频繁的抢话、重叠以及长篇幅的语用逻辑。这正是目前穿戴式语音系统从“演示”走向“商用”必须跨越的难题。
为了打破这一瓶颈,推动技术迈向真实的“类人”交互水平,由西北工业大学 ASLP 实验室联合上海交通大学、南京大学、中国科学技术大学、南洋理工大学、华为、希尔贝壳、Rokid等多家单位,发起 SmartGlasses (Egocentric Speech Interaction on AI Glasses) 挑战赛。首届挑战赛将在语音旗舰会议 IEEE SLT 2026 上举办。
竞赛试题
本次挑战赛聚焦两大核心社交交互场景,并针对每个场景从声学识别与语义理解两个维度进行深度评测。所有评测数据均基于商用 AI 眼镜在真实物理环境中采集。
赛道 一:双人对话理解
本赛道聚焦佩戴者与另一人的日常面对面交流,场景涵盖从安静室内到嘈杂户外等多种物理环境,重点考察系统在高度动态环境下的综合语音处理能力。
ASR 维度:要求模型具备强大的鲁棒性,能够从佩戴者的第一视角音频中,精准分离并转写出双方的对话内容。系统需要有效克服由佩戴者移动和开放环境带来的噪声干扰,并实现精准的说话人日志。
SLU 维度:侧重于评估模型在自然口语表达下的语义解析能力。要求系统能够在真实交谈中准确捕捉事实细节,并在存在噪声干扰的情况下追踪对话的逻辑流向。
赛道 二:多人会议理解
本赛道聚焦佩戴者参与的多人讨论或会议场景,数据采集自真实的办公与社交环境,包含了高频的语音重叠、快速的话题切换以及复杂的多人互动。
ASR 维度:要求系统在真实的多人语音重叠、抢话和环境干扰下,依然能保持高准确率的语音识别,并准确判定每个发言者的身份与时间边界。
SLU 维度:系统不仅要做到“听清”,还要能够跨越不同的说话人提取关键信息,梳理长篇幅、多轮次的对话脉络,最终获得客观、准确的理解。
SmartGlasses挑战赛 Website
参赛报名
大赛面向国内高等院校及科研院所在读学生 (含全日制本科生、硕士研究生、博士研究生)
*每支参赛团队人数为1-5人,只限1位老师指导完成。
*本届大赛各获奖队伍需要在第21届全国人机语音通讯学术会议(NCMMSC2026)的赛事特殊议题上以论文形式提交参赛技术方案说明,入选的论文会邀请在大会赛事特殊议题报告上作分享展示。具体会议时间与投稿要求另行通知。
赛程安排
The tentative timeline for running the challenge is as follows:
- 2026-04-15: Registration Opens
- 2026-05-07: Release of Training Set, Validation Set
- 2026-06-01: Registration Closes
- 2026-06-12[Updated!]: Release of Test Set
- 2026-06-26[Updated!]: Results Submission Deadline
- 2026-07-03: System Description Submission Deadline
- 2026-07-08: SLT Official Paper Submission Deadline
- 2026-09-01: Paper Notification
终稿提交规范
本次赛事由一支兼具专业性与行业影响力的团队全程组织护航,确保赛事公平、高效开展。核心组织者包括:
- 谢磊 西北工业大学
- 肖龙帅 华为技术有限公司
- 陈谐 上海交通大学
- 杜俊 中国科学技术大学
- 王帅 南京大学
- 薛浏蒙 南京大学
- Eng Siong Chng 南洋理工大学
- 付中华 西北工业大学
- 周军 Rokid
- 卜辉 希尔贝壳
- 徐昕 希尔贝壳
- 高德辉 西北工业大学
- 廖育杰 西北工业大学
- 赵致闲 西北工业大学
- 朱毅可 西北工业大学
联系邮箱
gdh@mail.nwpu.edu.cn
liaoyujie@mail.nwpu.edu.cn
zxzhao@mail.nwpu.edu.cn
注册指南
Step 1: Registration
Please complete your registration by filling out a registration form.
- Preferred: Google Registration Form
- Alternative (Mainland China): If you cannot access Google Forms, or if repeated submissions are unsuccessful, please use the Tencent Registration Form.
After you submit the registration form, the organizing committee will send a confirmation email within 1 business day. Please check your inbox in time; if you do not receive it, please check your spam folder first, or contact us via the email addresses below.
Step 2: Dataset Access
After successful registration and agreeing to the challenge rules, participating teams will be granted access to the SmartGlasses dataset, and the download link will be sent via email.
Contact Information
If you have any questions, please contact the organizers:
数据集
The organizing committee has provided independent download links for each track. After downloading and extracting the files, taking Track 1 as an example, the file structure of the dataset's root directory (SmartGlasses-Track1) is as follows (the structure for Track 2 is identical):
SmartGlasses-Track1/
├── Train/ # Training set
│ └── Part1/
│ ├── audio/ # Audio files (.wav)
│ └── textgrid/ # Text and timestamp annotation files (.TextGrid)
├── Dev/ # Development (Validation) set
│ └── Part1/
│ ├── audio/
│ ├── textgrid/
│ └── QA/ # QA annotation files (.json) - Provided only in the Dev set
└── data.jsonl # Global data index and metadata file
Detailed folder descriptions
·audio/: Contains the four-channel dialogue audio files (.wav format) for the respective track.
·textgrid/: Contains the .TextGrid annotation files corresponding to the audio. These files include speaker-level timestamp boundaries and their corresponding text transcriptions.
·QA/: (Provided only in the Dev set.) This folder contains the .json files used for the objective multiple-choice evaluation. Each JSON file corresponds to an audio segment and includes the question, options, and ground-truth answer required for the evaluation.
Special note on QA data: The multiple-choice QA data in the Dev set is automatically generated by large language models (LLMs) and has not undergone strict human verification; therefore, it may contain minor flaws or noise. We are fully open-sourcing this data and its answers to serve as reference examples, helping participating teams build their pipelines and debug their models. For the final hidden test set, we will use a similar approach to construct complex speech understanding and reasoning multiple-choice questions. However, all test questions and ground-truth answers will undergo strict human review and refinement to ensure absolute fairness and scientific rigor in the final evaluation.
Global index file: data.jsonl
A data.jsonl file is provided in the root directory of each track. This file serves as the global index for the entire dataset, where each line represents the metadata for a single data sample.
Note: The additionally collected data will be released later as Part 2. Part 2 is merely a chronological update in the release schedule; its data format and folder structure will be exactly identical to the current Part 1.
Dataset statistics
Below is the statistical overview of all released data (Part 1, Part 2, and Test):
Track 1: Dyadic Dialogue Understanding
| Split | Sessions | Total Duration (hrs) | Avg. Duration (sec) |
|---|---|---|---|
| Train (Part 1) | 332 | 29.12 | 315.71 |
| Train (Part 2) | 55 | 4.81 | 314.53 |
| Dev (Part 1) | 66 | 5.59 | 304.82 |
| Dev (Part 2) | 15 | 1.27 | 305.01 |
| Test | 50 | 4.16 | 299.80 |
| Total | 518 | 44.95 | 312.39 |
Track 2: Multi-party Meeting Understanding
| Split | Sessions | Total Duration (hrs) | Avg. Duration (sec) |
|---|---|---|---|
| Train (Part 1) | 105 | 34.13 | 1170.09 |
| Train (Part 2) | 25 | 7.70 | 1108.58 |
| Dev (Part 1) | 21 | 7.06 | 1210.47 |
| Dev (Part 2) | 15 | 4.45 | 1068.29 |
| Test | 30 | 8.69 | 1042.21 |
| Total | 196 | 62.03 | 1139.33 |
Smart Glasses Microphone Array Layout

Channel-to-Microphone Mapping
The 4 channels of the audio files are mapped to the physical micro-electro-mechanical systems (MEMS) microphone array integrated onto the smart glasses frames as follows:
·Channel 1 (mic1): Right temple, rear position
·Channel 2 (mic2): Right temple, front position
·Channel 3 (mic3): Left temple, front position
·Channel 4 (mic4): Left temple, rear position
Physical Array Geometry
The spatial coordinates and geometric constraints of the acoustic centers of the four microphones are specified below:
Horizontal Projection Displacements:
·Intra-temple separation (Right): The axial distance between mic1 and mic2 is 47 mm.
·Intra-temple separation (Left): The axial distance between mic3 and mic4 is 50 mm.
·Inter-temple span (Front): The cross-lateral distance between mic2 and mic3 is 145 mm.
·Inter-temple span (Rear): The cross-lateral distance between mic1 and mic4 is 146 mm.
Vertical & Lateral Offsets:
·Compared to the right-rear microphone (mic1), the right-front microphone (mic2) features a positive vertical elevation of 10 mm and an outward lateral offset of 1 mm along the frame thickness direction.
·The acoustic centers of mic1, mic3, and mic4 reside on the same horizontal reference plane (zero vertical offset).
·The left and right temples are orthogonal to the plane of the lenses.
·The baseline connecting mic1 and mic4 is strictly parallel to the plane of the lenses.
排行榜
The official leaderboards are now open for a total of 8 days. For this challenge, each task (TSA-ASR and SLU) under Track 1 and Track 2 is completely independent. The final rankings for each task will be calculated independently, and challenge prizes will be distributed separately. Please double-check and submit your results via the dedicated portals for the corresponding tasks. For submission procedures, packaging rules, and format examples, please refer to the detailed guidelines on each respective Leaderboard page.
📍 Submission Portals
Track 1: Dyadic Dialogue Understanding
Track 2: Multi-Party Meeting Understanding
⚠️ Important Evaluation Instructions for Task 2 (SLU)
To ensure fairness and scientific rigor, Task 2 adopts a two-phase evaluation mechanism and requires the latest data format:
- Data Update Requirement: You must use the latest test set data, which now includes the complete Question IDs, for inference and result formatting. [Download Link]
- Phase 1 (First 6 Days): Evaluation is based on the current question bank containing complete Question IDs. Leaderboard scores during this phase serve only as a periodic reference and will not be counted toward the final ranking.
- Phase 2 (Last 2 Days): The organizing committee will release the final batch of brand-new QA questions, which will be merged with the Phase 1 data to form the complete question bank. At that time, the supplemental QA data and the specific submission links for Phase 2 will be available directly on the Phase 1 Leaderboard pages linked above. This website will also be updated accordingly. The test scores from this phase (i.e., on the complete question bank) will serve as the sole criterion for the final ranking of the SLU task.
微信公众号
联系我们
商务合作:bd@aishelldata.com
技术服务:tech@aishelldata.com
联系电话:+86-010-80225006
公司地址:
北京市海淀区海淀大悦信息科技园D5-A501
开源数据
