A method for constructing a visual language action model based on an elimination Adaboost

By constructing a lightweight weak classifier pool based on the elimination system Adaboost to process multimodal inputs in parallel, the problems of insufficient generalization ability and real-time performance of VLA models in rare scenarios are solved, and efficient adaptation and real-time response to embodied tasks are achieved.

CN121598175BActive Publication Date: 2026-06-05SHENZHEN BEAUTIFUL RUBIKS CUBE ROBOT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610130374.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-30
Publication Date
2026-06-05
Estimated Expiration
2046-01-30

AI Technical Summary

Technical Problem

Existing vision-language-action (VLA) models lack generalization ability in rare scenarios and have insufficient real-time ensemble learning capabilities, making it difficult to meet the application requirements of embodied tasks.

Method used

We employ an Adaboost-based elimination approach, using multimodal datasets to generate a lightweight pool of weak classifiers. Combined with multi-stream parallel inference technology, we optimize action sequence output to improve the model's generalization ability and real-time performance.

Benefits of technology

It significantly improves the model's adaptability to complex scenarios, reduces inference latency and GPU memory requirements, and meets the real-time requirements of embodied tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121598175B_ABST
    Figure CN121598175B_ABST
Patent Text Reader

Abstract

The application discloses a method for constructing a visual language action model based on an elimination Adaboost, and relates to the technical field of embodied intelligence and multi-modal machine learning. The method for constructing the visual language action model comprises the following steps: training a multi-modal shared encoding module constructed based on a VLM component, fixing parameters of the VLM component after the training is completed, initializing weights of training samples, extracting feature vectors of the training samples based on the multi-modal shared encoding module and inputting the feature vectors into a weak classifier for training, iteratively training the weak classifier based on an elimination Adaboost, and forming a dynamic lightweight weak classifier pool; parallel processing input content of the multi-modal shared encoding module based on the weak classifier pool, fusing action sequence output based on Adaboost weights of the weak classifier, and performing smoothing processing on the action sequence based on EMA, so that the construction of the visual language action model is completed. The scheme disclosed by the application effectively improves the generalization ability of a VLA model and enhances the real-time performance of integrated reasoning of the VLA model.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Data set classification method and system based on unbalanced incremental learning

    CN119669862A

  • Road safety early warning method and system based on mixed precision quantification visual large model

    CN120808095A