中文多模态数据集「悟空」
正式名称(本次可确认):Noah-Wukong Dataset;页面类型:数据集官网/产品页。
所属章节:多模态;manifest 索引:400;原始行号:614;实时核验时间:2026-08-04T10:25:12.668Z。
项目简介
About The Noah-Wukong dataset is a large-scale multi-modality Chinese dataset. The dataset contains 100 Million <image, text> pairs Images in the datasets are filtered according to the size ( > 200px for both dimensions ) and aspect ratio ( 1/3 ~ 3 ) Text in the datasets are filtered according to its language, length and frequency. Privacy and sensitive words are also taken into consideration. Examples Announcement 2022/01/30 Dataset is now available for download 2022/02/14 Paper for Noah-Wukong Dataset is released
原始目录对本项目的说明为:华为诺亚方舟实验室开源大型,包含1亿图文对 本文保留该历史描述,但只把当前可访问的官方仓库、文档、论文或产品页中能直接确认的内容写成事实;没有证据的性能、规模、作者归属和商业可用性不作推断。
实时核验摘要
本次共核验 1 个原始链接,其中 1 个可确认或经重定向后可确认,0 个跳过/错误。事实陈述优先来自首个可确认的官方入口。
- 可确认正式名称:Noah-Wukong Dataset;
- 可确认页面类型:数据集官网/产品页;
- 技术栈或主要实现:PyTorch;
- 许可证:待核实;
- 版本:待核实。
项目特色
About The Noah-Wukong dataset is a large-scale multi-modality Chinese dataset. The dataset contains 100 Million <image, text> pairs Images in the datasets are filtered according to the size ( > 200px for both dimensions ) and aspect ratio ( 1/3 ~ 3 ) Text in the datasets are filtered according to its language, length and frequency. Privacy and sensitive words are also taken into consideration. Examples Announcement 2 这一点来自本次可访问页面的仓库描述、README 或页面摘要;其具体效果和适用范围仍需通过样例或数据验证。
特色判断只用于说明它与同章节其他资源的差异,不等同于“更先进”或“更适合生产”。若页面只给出项目自述,仍需用自己的数据、环境和评价标准复核。
核心功能
About The Noah-Wukong dataset is a large-scale multi-modality Chinese dataset. The dataset contains 100 Million <image, text> pairs Images in the datasets are filtered according to the size ( > 200px for both dimensions ) and aspect ratio ( 1/3 ~ 3 ) Text in the datasets are filtered according to its language, length and frequency. Privacy and sensitive words are also taken into consideration. Examples Announcement 2022/01/30 Dataset is now available for download 2022/02/14 Paper for Noah-Wukong Dataset is released
从资料边界看,本项目应先作为“多模态”候选资源评估。具体输入、输出、接口、训练流程、推理方式和异常处理只有在页面证据明确时才能确认;README 未说明的能力均为待核实。
技术栈与资料范围
可确认技术:PyTorch。页面主题标签:未从页面标签确认。
页面资料摘要显示:About The Noah-Wukong dataset is a large-scale multi-modality Chinese dataset. The dataset contains 100 Million <image, text> pairs Images in the datasets are filtered according to the size ( > 200px for both dimensions ) and aspect ratio ( 1/3 ~ 3 ) Text in the datasets are filtered according to its language, length and frequency. Privacy and sensitive words are also taken into consideration. Examples Announcement 2022/01/30 Dataset is now available for download 2022/02/14 Paper for Noah-Wukong Dataset is released at Arxiv 2022/03/28 Benchmark models is now released 2022/03/28 Wukong-Test is available for downlo
适用场景
评估时需同时核对图像与文本的来源、配对质量、许可范围、模型输入规范和跨模态指标,尤其不能把历史榜单或样例效果当作当前生产能力。
更稳妥的使用方式是先把项目限定在其页面明确支持的任务和数据域内,用少量可人工审查的样本验证,再决定是否扩展。若涉及医疗、司法、客服或公开数据发布,还应单独完成隐私、偏差、授权和人工兜底审查。
安装或使用入口
未从可访问页面确认稳定的安装命令;应从原始页面的 README、下载区或使用说明开始,不根据项目名称补写命令。
首要入口仍是下方完整保留的原始链接。若仓库已归档、命令引用旧依赖、模型或数据另行托管,建议在隔离环境中固定分支或提交号,并记录实际下载来源,避免把今天可见的说明误当作长期稳定接口。
实践建议
- 先确认任务定义和输入输出,与自己的数据格式做最小闭环测试;
- 记录访问日期、分支、提交或数据版本,并保存依赖清单;
- 对样例之外的长文本、空输入、噪声数据、编码和领域外输入做失败测试;
- 在再分发、训练或商用前核对代码许可证、数据条款及第三方素材授权;
- 将页面宣传、历史榜单和目录旧描述视为线索,不替代当前复现实验。
优势与限制
本条目的优势在于提供了可追溯的项目入口,并能从实时页面确认名称、类型及部分技术或数据线索。未取得可比较的仓库活跃度快照;本文不以页面热度推断项目质量。
限制是:网页状态会变化,仓库更新时间不等于维护承诺,README 自述不等于独立评测;无法访问的链接、没有明确版本的命令、未展示的许可证和未复现的指标都必须继续标记为待核实。
维护状态、版本与许可证
维护状态:官方仓库当前未标记为归档;最近推送时间为 待核实。这只能说明仓库状态,不能单凭一次更新时间断言仍在积极维护。
版本/更新时间:正式版本号与最近更新时间待核实。
许可证:待核实。即使仓库元数据显示 SPDX 标识,也应在实际使用前阅读仓库内 LICENSE、数据说明和第三方依赖条款;没有明确许可时,不应默认允许复制、再分发或商用。
原始链接与逐站访问结果
- 原始链接 1
https://wukong-dataset.github.io/wukong-dataset/
-
站点 1:已核验
https://wukong-dataset.github.io/wukong-dataset/- HTTP 状态:未取得
- 最终地址:
https://wukong-dataset.github.io/wukong-dataset/ - 页面标题:Noah-Wukong Dataset
- 访问时间:2026-08-04T10:25:12.668Z
- 证据摘要:About The Noah-Wukong dataset is a large-scale multi-modality Chinese dataset. The dataset contains 100 Million <image, text> pairs Images in the datasets are filtered according to the size ( > 200px for both dimensions ) and aspect ratio ( 1/3 ~ 3 ) Text in the datasets are filtered according to its language, length and frequency. Privacy and sensitive words are also taken into consideration. Examples Announcement 2022/01/30 Dataset is now available for download 2022/02/14 Paper for Noah-Wukong Dataset is released at Arxiv 2022/03/28 Benchmark models is now released 2022/03/28 Wukong-Test is available for download 2022/05/11 Chinese lables for dataset are now available at download 2022/07/05 Implementation code is available on Pytorch version. Citation @misc{gu2022wukong, title={Wukong: 100 Million Large-scale Chinese Cross-modal Pre-training Dataset and A Foundation Framework}, author=
- 跳过原因:无
核验清单
- 原项目名、slug、章节、原始行号和全部原始链接已保留;
- sourceChecks 数量 1,与 sourceLinks 数量 1 一致;
- 正式名称、页面类型、核心功能、技术范围、入口、维护、更新时间和许可证均按证据边界表述;
- 无法确认的信息已标记“待核实”,失败站点已记录原因;
- 规范文章链接:/articles/wukong-dataset-polished-guide。



