一种基于自我改进的大型视觉语言模型人类偏好对齐方法
By using a self-supervised framework for self-improving visual language models, and leveraging visual enhancement to generate diverse candidate answers and iteratively optimize them, the problem of large visual language models outputting results that do not conform to human preferences in medical diagnosis is solved, achieving efficient and low-cost preference alignment and model performance improvement.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG UNIV
- Filing Date
- 2025-10-10
- Publication Date
- 2026-07-17
AI Technical Summary
Existing large-scale visual language models face challenges in aligning generated content with human preferences, especially in medical diagnosis where outputs are biased, misleading, and do not conform to medical standards. Existing methods are costly and fail to fully utilize visual-text interaction.
We propose a self-improving human preference alignment method (SHAPE) for large visual language models. Through a self-supervised framework, we utilize visual enhancement to generate diverse candidate answers and iteratively optimize the model, constructing a high-quality preference dataset to achieve self-improvement of the model.
It achieves efficient preference alignment without manual annotation, generates more comprehensive text that conforms to human preferences, improves the accuracy and efficiency of the model in medical diagnosis, and significantly reduces costs.
Smart Images

Figure CN121525832B_ABST