A method for constructing a visual language action model based on an elimination Adaboost
By constructing a lightweight weak classifier pool based on the elimination system Adaboost to process multimodal inputs in parallel, the problems of insufficient generalization ability and real-time performance of VLA models in rare scenarios are solved, and efficient adaptation and real-time response to embodied tasks are achieved.
Patent Information
- Application Number
- CN202610130374.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-30
- Publication Date
- 2026-06-05
- Estimated Expiration
- 2046-01-30
AI Technical Summary
Existing vision-language-action (VLA) models lack generalization ability in rare scenarios and have insufficient real-time ensemble learning capabilities, making it difficult to meet the application requirements of embodied tasks.
We employ an Adaboost-based elimination approach, using multimodal datasets to generate a lightweight pool of weak classifiers. Combined with multi-stream parallel inference technology, we optimize action sequence output to improve the model's generalization ability and real-time performance.
It significantly improves the model's adaptability to complex scenarios, reduces inference latency and GPU memory requirements, and meets the real-time requirements of embodied tasks.
Smart Images

Figure CN121598175B_ABST
Abstract
Citation Information
Patent Citations
Data set classification method and system based on unbalanced incremental learning
CN119669862A
Road safety early warning method and system based on mixed precision quantification visual large model
CN120808095A