LLMScore

所属章节:多模态LLM;原始行号:294;资料类型:数据集或语料资源

事实边界:本文在 2026-08-04T10:32:41.739Z 对 assignment 中的 2 个原始链接逐一核验。可访问页面支持的事实会明确写出;历史目录描述、无法访问页面及页面未披露的信息不自动视为当前事实,统一标记“待核实”。

项目简介

“LLMScore”来自 docs/a.md 第 294 行,归入“多模态LLM”。原目录留下的文字是:LLMScore是一种全新的框架,能够提供具有多粒度组合性的评估分数。它使用大语言模型(LLM)来评估文本到图像生成模型。首先,将图像转化为图像级别和对象级别的视觉描述,然后将评估指令输入到LLM中,以衡量合成图像与文本的对齐程度,并最终生成一个评分和解释。我们的大量分析显示,LLMScore在众多数据集上与人类判断的相关性最高,明显优于常用的文本-图像匹配度量指标CLIP和BLIP。

本轮核验后,首要可确认的正式名称或页面标题为“LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis Evaluation”,页面类型为“arXiv 论文页面;GitHub 代码或资料仓库”。目录名称用于保持索引兼容,不代表外部项目当前仍沿用完全相同的命名。

实时核验摘要

共检查 2 个原始 URL,其中 2 个取得可确认身份与内容的证据,0 个因失效、限制、内容不足或错误而跳过。当前可确认的核心信息是:Existing automatic evaluation on text-to-image synthesis can only provide an image-text matching score, without considering the object-level compositionality, which results in poor correlation with human judgments. In this work, we propose LLMScore, a new framework that offers evaluation scores with multi-granularity compositionality. LLMScore leverages the large language models (LLMs) to evaluate text-to-image models. Initially, it transforms the image into image-level and object-level visual descriptions. Then an evaluation instruction is fed into the LLMs to measure the alignment between the synthesized image and the text, ultimately generating a score accompanied by a rationale. Our substantial analysis reveals the highest correlation of LLMScore with human judgments on a wide range of datasets (Attribute Binding Contrast, Concept Conjunction, MSCOCO, DrawBench, PaintSkills). Notably, our LLMScore achieves Kendall's tau correlation with human evaluations that is 58.8% and 31.2% higher than the commonly-used text-image matching metrics CLIP and BLIP, respectively.

技术或资料范围:论文方法、实验设置与数据范围需以正文为准。版本/更新时间:Submission history From: Yujie Lu [ view email ] [v1] Thu, 18 May 2023 16:57:57 UTC (7,498 KB) Full-text links: Access Paper: View a PDF of the paper titled LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis Evaluation, by Yujie Lu and 4 other authors。这些信息均对应本次访问时页面状态,不承诺未来不变。

项目特色

Existing automatic evaluation on text-to-image synthesis can only provide an image-text matching score, without considering the object-level compositionality, which results in poor correlation with human judgments. In this work, we propose LLMScore, a new framework that offers evaluation scores with multi-granularity compositionality. LLMScore leverages the large language models (LLMs) to evaluate text-to-image models. Initially, it transforms the image into image-level and object-level visual descriptions. Then an evaluation instruction is fed into the LLMs to measure the alignment between the synthesized image and the text, ultimately generating a score accompanied by a rationale. Our substantial analysis reveals the highest correlation of LLMScore with human judgments on a wide range of datasets (Attribute Binding Contrast, Concept Conjunction, MSCOCO, DrawBench, PaintSkills). Notably, our LLMScore achieves Kendall's tau correlation with human evaluations that is 58.8% and 31.2% higher than the commonly-used text-image matching metrics CLIP and BLIP, respectively.

上述特色只取自可访问页面的项目说明、数据卡、论文摘要或仓库元数据;若证据仅来自标题或简介,本文不会进一步外推性能、质量、适用规模或商业可用性。

核心功能

Existing automatic evaluation on text-to-image synthesis can only provide an image-text matching score, without considering the object-level compositionality, which results in poor correlation with human judgments. In this work, we propose LLMScore, a new framework that offers evaluation scores with multi-granularity compositionality. LLMScore leverages the large language models (LLMs) to evaluate text-to-image models. Initially, it transforms the image into image-level and object-level visual descriptions. Then an evaluation instruction is fed into the LLMs to measure the alignment between the synthesized image and the text, ultimately generating a score accompanied by a rationale. Our substantial analysis reveals the highest correlation of LLMScore with human judgments on a wide range of datasets (Attribute Binding Contrast, Concept Conjunction, MSCOCO, DrawBench, PaintSkills). Notably, our LLMScore achieves Kendall's tau correlation with human evaluations that is 58.8% and 31.2% higher than the commonly-used text-image matching metrics CLIP and BLIP, respectively.

原始目录的历史描述可帮助理解收录原因,但功能是否仍存在、默认是否启用、是否依赖外部服务以及输出质量如何,必须以当前文档和实际测试为准。未在页面中明确出现的能力均待核实。

技术栈与资料范围

当前可确认:论文方法、实验设置与数据范围需以正文为准。

多模态项目需分别核对模型权重、输入输出模态、显存与依赖、训练数据授权及生成内容风险。 对模型、数据或课程中出现的参数规模、性能数字和数据量,应回到具体版本与测试条件核对;本文不把目录里的历史数字重新包装为当前结论。

适用场景

作为“多模态LLM”下的数据集或语料资源,它适合用于建立候选资料清单、复核原始来源、开展小规模验证或比较同类方案。实际采用前,应先明确目标是阅读、复现、下载数据、运行模型还是接入产品,再判断当前入口是否满足要求。

生产系统、高风险行业、公开服务或涉及个人数据的任务不能只依据目录条目决策,还需完成安全、隐私、许可证、成本、性能和运维评估。

安装或使用入口

论文页面本身不是安装入口;若作者提供代码,应以论文中的官方代码链接为准。

只有在官方页面明确给出时才应采用具体命令、依赖版本、模型权重或 API 参数。若入口是论文、视频、课程或数据页面,它本身不等同于可安装的软件包。

实践建议

  1. 先确认最终 URL、页面标题和发布主体与项目相符,再下载代码、模型或数据。
  2. 把原始目录描述视为历史线索,逐项对照当前 README、数据卡、论文版本或产品说明。
  3. 多模态项目需分别核对模型权重、输入输出模态、显存与依赖、训练数据授权及生成内容风险。
  4. 用隔离环境做最小化验证,记录依赖版本、输入样例、输出结果和失败条件。
  5. 在分发、商用或公开部署前单独核验主项目及所有上游模型、数据和依赖的许可。

优势与限制

优势:该条目保留了项目在原目录中的定位与全部来源入口;本轮又补充了逐站状态、最终 URL、页面身份、更新线索与可确认特色,便于后续复核。

限制:网页可访问不等于项目可复现,仓库未归档不等于活跃维护,页面更新时间也不等于正式版本发布日期。未实际运行代码、下载完整数据或观看完整视频的部分仍需验证;不可访问来源不作推测。

维护状态

论文资料通常按版本修订,不能据此推断配套代码仍在维护。

维护判断仅依据页面是否归档、可见更新时间或发布记录;issue 响应速度、兼容性、路线图和长期支持承诺均待核实。

许可证与资料可信度

当前可确认许可证:待核实。若这里显示“待核实”,表示本轮可访问证据没有给出明确、可归属到当前项目的许可证,而不是默认允许自由使用。

资料可信度按“官方仓库/论文/官方产品页优先”的原则评估。短链、转载视频、聚合页或个人博客可作为线索,但不能替代主项目的许可证、版本和使用条款。

原始链接与逐站访问结果

原始项目名:LLMScore;slug:llmscore原始章节:多模态LLM;原始行号:294。

  1. 原始 URL:https://arxiv.org/abs/2305.11116

    状态:verified;HTTP:200;最终 URL:https://arxiv.org/abs/2305.11116

    页面标题:LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis Evaluation;核验时间:2026-08-04T10:32:39.361Z

    证据摘要:正式名称/标题:LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis Evaluation;页面类型:arXiv 论文页面;核心信息:Existing automatic evaluation on text-to-image synthesis can only provide an image-text matching score, without considering the object-level compositionality, which results in poor correlation with human judgments. In this work, we propose LLMScore, a new framework that offers evaluation scores with multi-granularity compositionality. LLMScore leverages the large language models (LLMs) to evaluate text-to-image models. Initially, it transforms the image into image-level and object-level visual descriptions. Then an evaluation instruction is fed into the LLMs to measure the alignment between the synthesized image and the text, ultimately generating a score accompanied by a rationale. Our substantial analysis reveals the highest correlation of LLMScore with human judgments on a wide range of datasets (Attribute Binding Contrast, Concept Conjunction, MSCOCO, DrawBench, PaintSkills). Notably, our LLMScore achieves Kendall's tau correlation with human evaluations that is 58.8% and 31.2% higher than the commonly-used text-image matching metrics CLIP and BLIP, respectively.;技术或资料范围:论文方法、实验设置与数据范围需以正文为准;维护/更新:论文资料通常按版本修订,不能据此推断配套代码仍在维护。 Submission history From: Yujie Lu [ view email ] [v1] Thu, 18 May 2023 16:57:57 UTC (7,498 KB) Full-text links: Access Paper: View a PDF of the paper titled LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis Evaluation, by Yujie Lu and 4 other authors;许可证:待核实;特色证据:Existing automatic evaluation on text-to-image synthesis can only provide an image-text matching score, without considering the object-level compositionality, which results in poor correlation with human judgments. In this work, we pro

    跳过原因:

  2. 原始 URL:https://github.com/YujieLu10/LLMScore

    状态:verified;HTTP:200;最终 URL:https://github.com/YujieLu10/LLMScore

    页面标题:GitHub - YujieLu10/LLMScore: LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis Evaluation;核验时间:2026-08-04T10:32:41.191Z

    证据摘要:正式名称/标题:GitHub - YujieLu10/LLMScore: LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis Evaluation;页面类型:GitHub 代码或资料仓库;核心信息:LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis Evaluation - YujieLu10/LLMScore;技术或资料范围:待核实;维护/更新:维护状态待核实。 版本或更新时间待核实。;许可证:待核实;特色证据:LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis Evaluation - YujieLu10/LLMScore

    跳过原因:

核验清单

  1. slug、原始项目名、章节、行号与 assignment 保持一致。
  2. 全部 2 个原始链接均已保留,并各自对应一条 sourceChecks 记录。
  3. 正式名称、核心功能、技术范围、入口、维护、版本和许可证只写入可确认事实。
  4. 至少一项项目特色有页面证据;无证据时明确写出“暂未从可访问资料确认特色”。
  5. 失效、登录限制、拦截、超时、无关或身份不明的页面已记录原因,没有据此猜测。

发布路径