歧义数据集

所属章节:LLM的数据集;原始行号:300;资料类型:数据集或语料资源

事实边界:本文在 2026-08-04T10:33:06.244Z 对 assignment 中的 2 个原始链接逐一核验。可访问页面支持的事实会明确写出;历史目录描述、无法访问页面及页面未披露的信息不自动视为当前事实,统一标记“待核实”。

项目简介

“歧义数据集”来自 docs/a.md 第 300 行,归入“LLM的数据集”。原目录留下的文字是:能否正确的消除歧义是衡量大语言模型的一个重要指标。不过一直没有一个标准化的衡量方法,这篇论文提出了一个包含1,645个具有不同种类歧义的数据集及对应的评估方法。

本轮核验后,首要可确认的正式名称或页面标题为“GitHub - alisawuffles/ambient: Code and data associated with the AmbiEnt dataset in "We're Afraid Language Models Aren't Modeling Ambiguity" (Liu et al., 2023)”,页面类型为“GitHub 代码或资料仓库;arXiv 论文页面”。目录名称用于保持索引兼容,不代表外部项目当前仍沿用完全相同的命名。

实时核验摘要

共检查 2 个原始 URL,其中 2 个取得可确认身份与内容的证据,0 个因失效、限制、内容不足或错误而跳过。当前可确认的核心信息是:Code and data associated with the AmbiEnt dataset in "We're Afraid Language Models Aren't Modeling Ambiguity" (Liu et al., 2023) - alisawuffles/ambient

技术或资料范围:论文方法、实验设置与数据范围需以正文为准。版本/更新时间:版本或更新时间待核实。。这些信息均对应本次访问时页面状态,不承诺未来不变。

项目特色

Code and data associated with the AmbiEnt dataset in "We're Afraid Language Models Aren't Modeling Ambiguity" (Liu et al., 2023) - alisawuffles/ambient

上述特色只取自可访问页面的项目说明、数据卡、论文摘要或仓库元数据;若证据仅来自标题或简介,本文不会进一步外推性能、质量、适用规模或商业可用性。

核心功能

Code and data associated with the AmbiEnt dataset in "We're Afraid Language Models Aren't Modeling Ambiguity" (Liu et al., 2023) - alisawuffles/ambient

原始目录的历史描述可帮助理解收录原因,但功能是否仍存在、默认是否启用、是否依赖外部服务以及输出质量如何,必须以当前文档和实际测试为准。未在页面中明确出现的能力均待核实。

技术栈与资料范围

当前可确认:论文方法、实验设置与数据范围需以正文为准。

数据集采用前应检查样本结构、划分方式、生成或采集来源、许可、敏感信息与去重污染问题。 对模型、数据或课程中出现的参数规模、性能数字和数据量,应回到具体版本与测试条件核对;本文不把目录里的历史数字重新包装为当前结论。

适用场景

作为“LLM的数据集”下的数据集或语料资源,它适合用于建立候选资料清单、复核原始来源、开展小规模验证或比较同类方案。实际采用前,应先明确目标是阅读、复现、下载数据、运行模型还是接入产品,再判断当前入口是否满足要求。

生产系统、高风险行业、公开服务或涉及个人数据的任务不能只依据目录条目决策,还需完成安全、隐私、许可证、成本、性能和运维评估。

安装或使用入口

以原始页面当前提供的使用、注册、下载或文档入口为准;未确认到统一安装命令。

只有在官方页面明确给出时才应采用具体命令、依赖版本、模型权重或 API 参数。若入口是论文、视频、课程或数据页面,它本身不等同于可安装的软件包。

实践建议

  1. 先确认最终 URL、页面标题和发布主体与项目相符,再下载代码、模型或数据。
  2. 把原始目录描述视为历史线索,逐项对照当前 README、数据卡、论文版本或产品说明。
  3. 数据集采用前应检查样本结构、划分方式、生成或采集来源、许可、敏感信息与去重污染问题。
  4. 用隔离环境做最小化验证,记录依赖版本、输入样例、输出结果和失败条件。
  5. 在分发、商用或公开部署前单独核验主项目及所有上游模型、数据和依赖的许可。

优势与限制

优势:该条目保留了项目在原目录中的定位与全部来源入口;本轮又补充了逐站状态、最终 URL、页面身份、更新线索与可确认特色,便于后续复核。

限制:网页可访问不等于项目可复现,仓库未归档不等于活跃维护,页面更新时间也不等于正式版本发布日期。未实际运行代码、下载完整数据或观看完整视频的部分仍需验证;不可访问来源不作推测。

维护状态

维护状态待核实。

维护判断仅依据页面是否归档、可见更新时间或发布记录;issue 响应速度、兼容性、路线图和长期支持承诺均待核实。

许可证与资料可信度

当前可确认许可证:待核实。若这里显示“待核实”,表示本轮可访问证据没有给出明确、可归属到当前项目的许可证,而不是默认允许自由使用。

资料可信度按“官方仓库/论文/官方产品页优先”的原则评估。短链、转载视频、聚合页或个人博客可作为线索,但不能替代主项目的许可证、版本和使用条款。

原始链接与逐站访问结果

原始项目名:歧义数据集;slug:project-300原始章节:LLM的数据集;原始行号:300。

  1. 原始 URL:https://github.com/alisawuffles/ambient

    状态:verified;HTTP:200;最终 URL:https://github.com/alisawuffles/ambient

    页面标题:GitHub - alisawuffles/ambient: Code and data associated with the AmbiEnt dataset in "We're Afraid Language Models Aren't Modeling Ambiguity" (Liu et al., 2023);核验时间:2026-08-04T10:32:48.527Z

    证据摘要:正式名称/标题:GitHub - alisawuffles/ambient: Code and data associated with the AmbiEnt dataset in "We're Afraid Language Models Aren't Modeling Ambiguity" (Liu et al., 2023);页面类型:GitHub 代码或资料仓库;核心信息:Code and data associated with the AmbiEnt dataset in "We're Afraid Language Models Aren't Modeling Ambiguity" (Liu et al., 2023) - alisawuffles/ambient;技术或资料范围:待核实;维护/更新:维护状态待核实。 版本或更新时间待核实。;许可证:待核实;特色证据:Code and data associated with the AmbiEnt dataset in "We're Afraid Language Models Aren't Modeling Ambiguity" (Liu et al., 2023) - alisawuffles/ambient

    跳过原因:

  2. 原始 URL:arxiv.org/abs/2304.14399

    状态:verified;HTTP:200;最终 URL:https://arxiv.org/abs/2304.14399

    页面标题:We're Afraid Language Models Aren't Modeling Ambiguity;核验时间:2026-08-04T10:33:06.243Z

    证据摘要:正式名称/标题:We're Afraid Language Models Aren't Modeling Ambiguity;页面类型:arXiv 论文页面;核心信息:Ambiguity is an intrinsic feature of natural language. Managing ambiguity is a key part of human language understanding, allowing us to anticipate misunderstanding as communicators and revise our interpretations as listeners. As language models (LMs) are increasingly employed as dialogue interfaces and writing aids, handling ambiguous language is critical to their success. We characterize ambiguity in a sentence by its effect on entailment relations with another sentence, and collect AmbiEnt, a linguist-annotated benchmark of 1,645 examples with diverse kinds of ambiguity. We design a suite of tests based on AmbiEnt, presenting the first evaluation of pretrained LMs to recognize ambiguity and disentangle possible meanings. We find that the task remains extremely challenging, including for GPT-4, whose generated disambiguations are considered correct only 32% of the time in human evaluation, compared to 90% for disambiguations in our dataset. Finally, to illustrate the value of ambiguity-sensitive tools, we show that a multilabel NLI model can flag political claims in the wild that are misleading due to ambiguity. We encourage the field to rediscover the importance of ambiguity for NLP.;技术或资料范围:论文方法、实验设置与数据范围需以正文为准;维护/更新:论文资料通常按版本修订,不能据此推断配套代码仍在维护。 Submission history From: Alisa Liu [ view email ] [v1] Thu, 27 Apr 2023 17:57:58 UTC (7,649 KB) [v2] Fri, 20 Oct 2023 05:46:14 UTC (8,199 KB) Full-text links: Access Paper: View a PDF of the pa;许可证:待核实;特色证据:Ambiguity is an intrinsic feature of natural language. Managing ambiguity is a key part of human language understanding, allowing us to anticipate misunderstanding as communicators and revise our interpretations as listeners. As language mod

    跳过原因:

核验清单

  1. slug、原始项目名、章节、行号与 assignment 保持一致。
  2. 全部 2 个原始链接均已保留,并各自对应一条 sourceChecks 记录。
  3. 正式名称、核心功能、技术范围、入口、维护、版本和许可证只写入可确认事实。
  4. 至少一项项目特色有页面证据;无证据时明确写出“暂未从可访问资料确认特色”。
  5. 失效、登录限制、拦截、超时、无关或身份不明的页面已记录原因,没有据此猜测。

发布路径