基于视觉注意力掩码的CT报告直接偏好优化方法及装置

By using visual attention masks and reference-guided generation strategies, combined with a composite loss function to optimize the CT report generation model, the problem of factual illusion caused by reliance on prior linguistic knowledge in CT report generation is solved, thereby improving the clinical accuracy and hallucination suppression efficiency of the model.

CN122091065BActive Publication Date: 2026-07-17BEIHANG UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIHANG UNIV
Filing Date
2026-04-23
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing CT report generation methods rely excessively on prior linguistic knowledge, leading to factual illusions, and lack a mechanism to construct unbiased samples by combining visual evidence, thus failing to improve the clinical accuracy of the model without requiring a large amount of additional labeled data.

Method used

By blocking access to visual information through a visual attention masking mechanism, the model relies on prior linguistic knowledge to generate non-preferred samples. A reference-guided generation strategy ensures that the preferred and non-preferred reports are consistent in syntactic structure. Combined with a composite loss function to optimize the model, high-quality CT reports are generated.

Benefits of technology

It effectively suppresses false positive and false negative hallucinations, improves the clinical accuracy of the model without requiring a large amount of labeled data, avoids shortcut learning, and improves hallucination suppression efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122091065B_ABST
    Figure CN122091065B_ABST
Patent Text Reader

Abstract

本申请公开了一种基于视觉注意力掩码的CT报告直接偏好优化方法及装置,涉及医学影像处理技术领域,其中,方法包括:通过视觉注意力掩码机制主动阻断视觉信息访问,使得模型在阻断视觉的情况下,依赖自身语言先验知识产生幻觉输出,从而构成非优选报告,有效抑制假阳性和假阴性幻觉,并通过参考引导生成策略,确保优选与非优选报告在句法结构上保持高度一致,使DPO训练聚焦于区分事实性内容与幻觉性描述,避免捷径学习,显著提升幻觉抑制效率。由此,解决了现有技术过度依赖语言先验知识而导致事实幻觉,且缺乏结合视觉证据构建非偏好样本的机制,无法在无需大量额外标注数据的情况下实现模型临床准确性的提升等问题。
Need to check novelty before this filing date? Find Prior Art