视觉语言模型的两阶段微调及解耦推理方法和装置
By employing decoupled learning and decoupled inference strategies for panoramic and subject views, the bias problem in context processing of visual language models is resolved, achieving a balance between improved base class recognition performance and new class generalization ability, thus enhancing the robustness of the model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NAT UNIV OF DEFENSE TECH
- Filing Date
- 2026-05-09
- Publication Date
- 2026-07-17
AI Technical Summary
Existing visual language models struggle to distinguish between semantic cues and interference biases when processing visual context, resulting in an inability to simultaneously address both context-dependent and context-independent samples, and a conflict between base class learning and the ability to generalize to new classes.
By employing view-specific cue learning, base class optimal decision synthesis, and decoupled reasoning strategies, global context and local subject details are captured through decoupled learning of panoramic and subject views, and weighted fusion is performed on the base class. Evidence theory fusion is used on the new class to achieve differentiated reasoning.
It effectively addresses the double-edged sword effect of context, improves base class recognition performance while maintaining the generalization ability of new classes, and enhances the robustness of the model in open-world scenarios.
Smart Images

Figure CN122154841B_ABST