基于视觉注意力掩码的CT报告直接偏好优化方法及装置
By using visual attention masks and reference-guided generation strategies, combined with a composite loss function to optimize the CT report generation model, the problem of factual illusion caused by reliance on prior linguistic knowledge in CT report generation is solved, thereby improving the clinical accuracy and hallucination suppression efficiency of the model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIHANG UNIV
- Filing Date
- 2026-04-23
- Publication Date
- 2026-07-17
AI Technical Summary
Existing CT report generation methods rely excessively on prior linguistic knowledge, leading to factual illusions, and lack a mechanism to construct unbiased samples by combining visual evidence, thus failing to improve the clinical accuracy of the model without requiring a large amount of additional labeled data.
By blocking access to visual information through a visual attention masking mechanism, the model relies on prior linguistic knowledge to generate non-preferred samples. A reference-guided generation strategy ensures that the preferred and non-preferred reports are consistent in syntactic structure. Combined with a composite loss function to optimize the model, high-quality CT reports are generated.
It effectively suppresses false positive and false negative hallucinations, improves the clinical accuracy of the model without requiring a large amount of labeled data, avoids shortcut learning, and improves hallucination suppression efficiency.
Smart Images

Figure CN122091065B_ABST