一种基于自我改进的大型视觉语言模型人类偏好对齐方法

By using a self-supervised framework for self-improving visual language models, and leveraging visual enhancement to generate diverse candidate answers and iteratively optimize them, the problem of large visual language models outputting results that do not conform to human preferences in medical diagnosis is solved, achieving efficient and low-cost preference alignment and model performance improvement.

CN121525832BActive Publication Date: 2026-07-17ZHEJIANG UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG UNIV
Filing Date
2025-10-10
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing large-scale visual language models face challenges in aligning generated content with human preferences, especially in medical diagnosis where outputs are biased, misleading, and do not conform to medical standards. Existing methods are costly and fail to fully utilize visual-text interaction.

Method used

We propose a self-improving human preference alignment method (SHAPE) for large visual language models. Through a self-supervised framework, we utilize visual enhancement to generate diverse candidate answers and iteratively optimize the model, constructing a high-quality preference dataset to achieve self-improvement of the model.

Benefits of technology

It achieves efficient preference alignment without manual annotation, generates more comprehensive text that conforms to human preferences, improves the accuracy and efficiency of the model in medical diagnosis, and significantly reduces costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121525832B_ABST
    Figure CN121525832B_ABST
Patent Text Reader

Abstract

基于自我改进的大型视觉语言模型人类偏好对齐方法,包括:自动构建偏好数据集;从视觉问答数据集中获取“图像‑问题”对;对图像施加多种视觉增强,并结合问题,驱动参考模型生成一组候选答案;将这组候选答案汇总提炼成一个更全面的“优胜”文本,同时将参考模型对于原始图片和问题的文本答案作为“落败”文本,由此构成“图像‑询问‑‘优胜’文本‑‘落败’文本”格式的偏好数据;基于步骤1构建的偏好数据集,应用直接偏好优化算法对目标视觉语言模型进行对齐微调,在微调过程中保持视觉编码器参数冻结,仅在扩模态对齐模块和语言解码器中引入低秩适配层(LoRA)进行训练;模型迭代式自我改进;微调完成后,将优化过的模型作为新的参考模型,并重复步骤1的数据构建流程,以生成质量更高的偏好数据用于下一轮优化;如此循环往复,实现模型对齐能力的持续自我提升。本发明实现了无需人工标注的自监督偏好对齐,生成了更全面且高质量的“优胜”文本。
Need to check novelty before this filing date? Find Prior Art