Robot VLA model optimization method and system based on video representation alignment
Patent Information
- Application Number
- CN202611358422.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-09-03
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]本发明的目的在于,提供一种基于视频表征对齐的机器人VLA模型优化方法及系统,以解决现有机器人VLA模型不涉及时序动态先验、无法利用上下文的问题
本发明提供的基于视频表征对齐的机器人VLA模型优化方法,首先利用视频基础模型对包含当前观测帧在内的帧视频片段进行特征提取,得到
帧时序动态特征作为对齐信号,为后续表征对齐提供包含物体运动规律和物理交互先验的监督目标。其次,获取VLA模型多层注意力主干第
层输出的视觉令牌特征,并利用动态解码器将该视觉令牌特征转换为与所述
帧时序动态特征逐帧对应的
帧视频预测特征,实现从VLA模型特征空间到视频基础模型特征空间的映射。在此基础上,确定每帧视频预测特征与对应帧时序动态特征之间的帧级特征对齐损失和时序对比对齐损失,其中帧级特征对齐损失用于约束每一帧预测特征与对应帧真实特征在语义空间中的相似性,时序对比对齐损失用于强制不同帧之间保持时序区分性,两者联合有效解决了单帧VLA模型与多帧视频表征不对称导致的对齐后表征易坍缩为时序平均化静态特征的问题。最后,利用上述两种损失对VLA模型和动态解码器进行联合训练,使视频基础模型所蕴含的时序动态先验通过表征对齐隐式迁移至VLA模型内部,训练完成后动态解码器和视频基础模型被丢弃,推理阶段仅使用VLA模型原始推理流程,不引入任何额外参数和计算开销。在保持VLA模型高效推理优势的同时赋予其时序动态理解能力,且不修改基础VLA模型架构,具有良好的插件式兼容性。该优化方法能够解决现有VLA模型不涉及时序动态先验、无法利用上下文的问题。
Smart Images

Figure CN122840146A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method and system for optimizing robot VLA models based on video representation alignment, belonging to the interdisciplinary field of artificial intelligence and robotics. Background Technology
[0002] Vision-Language-Action (VLA) is the core solution for general robot control. It achieves end-to-end control by mapping visual observations and language commands into robot-executable actions.
[0003] Existing VLA models typically include a visual encoder, a text segmenter, a multi-layer attention backbone, and an action head. Their Vision-Language Model (VLM) backbone is primarily pre-trained on large-scale static image-text data, thus possessing rich semantic understanding and scene reasoning capabilities. Typical VLA models include the open-source Vision-Language-Action Model (OpenVLA), the Robotic Transformer 2 (RT-2) model, and the aVision-Language-Action Flow Model for General Robot Control. A Vision-Language-Action Model with Open-World Generalization In this type of model, during inference, the visual encoder encodes the current visual observation to generate a visual token, the text segmenter segments the language instruction into language tokens, the visual tokens and language tokens interact across modalities through a multi-layer attention backbone to refine the representation, and finally the action head predicts multi-step action blocks based on the refined representation, and the robot executes the action blocks to complete the control task.
[0004] However, the VLM backbone of existing VLA models takes single-frame images as input, failing to capture object motion patterns, physical interaction causality, and temporal dynamic priors. This limits their performance in long-term manipulation, dynamic scenes involving moving targets, and tasks requiring prediction of object trajectories. Furthermore, the inference process of existing VLA models relies solely on the current single-frame visual observation, unable to utilize contextual information from past moments to aid current decision-making, further restricting their performance in complex dynamic environments. Summary of the Invention
[0005] The purpose of this invention is to provide a method and system for optimizing robot VLA models based on video representation alignment, so as to solve the problems that existing robot VLA models do not involve temporal dynamic priors and cannot utilize context.
[0006] A first aspect of the present invention provides a method for optimizing a robot VLA model based on video representation alignment, comprising: S1. Using video basic models to... Feature extraction is performed on the video segments to obtain... Frame temporal dynamic characteristics; The video frame contains the currently observed frame.
[0007] S2, Obtain the VLA model's multi-layer attention backbone. The visual token features output by the layer are transformed using a dynamic decoder. Frame-based video prediction features; Frame video prediction features and The frame-by-frame dynamic features correspond to the time sequence.
[0008] S3. Determine the frame-level feature alignment loss and temporal contrast alignment loss between the predicted features of each video frame and the temporal dynamic features of the corresponding frame.
[0009] S4. Jointly train the VLA model and the dynamic decoder using frame-level feature alignment loss and temporal contrast alignment loss to obtain the optimized VLA model.
[0010] Specifically, the frame-level feature alignment loss and temporal contrast alignment loss between the predicted features of each video frame and the temporal dynamic features of the corresponding frame are determined, including: Determine the cosine similarity between the predicted features of each video frame and the temporal dynamic features of the corresponding frame.
[0011] The frame-level feature alignment loss and temporal comparison alignment loss are determined based on cosine similarity.
[0012] Specifically, the VLA model and dynamic decoder are jointly trained using frame-level feature alignment loss and temporal contrast alignment loss, including: Determine the weighted sum of frame-level feature alignment loss and temporal contrast alignment loss.
[0013] The total loss is determined based on the weighted sum and the loss due to actions.
[0014] The parameters of the VLA model and the dynamic decoder are jointly trained using the total loss.
[0015] Specifically, the VLA model includes a visual encoder, a text segmenter, and a multi-layer attention backbone. Therefore, obtaining the first... The visual token features output by the layer include: The current observation frame is encoded using a visual encoder to generate a visual token.
[0016] Use a text segmenter to segment language instructions into language tokens.
[0017] This study utilizes a multi-layered attention backbone to enable cross-modal interaction between visual and linguistic tokens, and obtains the first-order attention value of the multi-layered attention backbone. Visual token features output by the layer.
[0018] Specifically, a dynamic decoder is used to convert visual token features into... Frame video prediction features include: Amplifying the hidden dimensions of visual token features using a dynamic decoder times, .
[0019] Magnify the hidden dimension Projecting the visual token features of multiples onto the video base model In the multiple dimensions, the token features of the video model are obtained.
[0020] Reshape the video model token features into the video base model. 1 eigenvector, to obtain Frame video prediction features.
[0021] Specifically, S1 includes: S11, to Noise perturbation is applied to frame video segments to obtain noise latent features.
[0022] S12. Input the language instructions and noise latent features into the diffusion Transformer of the video base model to extract the first... Intermediate features of the layer.
[0023] S13, regarding the first Spatial average pooling is performed on the intermediate features of the layer to obtain Frame-time dynamic characteristics.
[0024] Specifically, for Noise perturbation is applied to frame video segments to obtain noise latent features, including: S111, will Frame video segments are compressed into clean latent variables using a video variational autoencoder based on the video fundamental model.
[0025] S112. Apply noise perturbation to the clean latent variables to obtain the noise latent features.
[0026] A second aspect of the present invention provides an optimization system based on the above-described robot VLA model optimization method, comprising: The temporal feature determination module is used to utilize the video base model to determine... Feature extraction is performed on the video segments to obtain... Frame temporal dynamic characteristics; The video frame contains the currently observed frame.
[0027] The video feature determination module is used to obtain the first multi-layer attention backbone of the VLA model. The visual token features output by the layer are transformed using a dynamic decoder. Frame-based video prediction features; Frame video prediction features and The frame-by-frame dynamic features correspond to the time sequence.
[0028] The loss determination module is used to determine the frame-level feature alignment loss and temporal contrast alignment loss between the predicted features of each video frame and the temporal dynamic features of the corresponding frame.
[0029] The joint training module is used to jointly train the VLA model and the dynamic decoder using frame-level feature alignment loss and temporal contrast alignment loss to obtain an optimized VLA model.
[0030] The robot VLA model optimization method and system based on video representation alignment of the present invention has the following advantages compared with the prior art: The robot VLA model optimization method based on video representation alignment provided by this invention first utilizes the video base model to optimize the VLA model including the current observation frame. Feature extraction is performed on the video segments to obtain... Frame-time dynamic features serve as alignment signals, providing a supervisory target that includes prior knowledge of object motion patterns and physical interactions for subsequent alignment representations. Secondly, the VLA model's multi-layer attention backbone is obtained... The visual token features output by the layer are then converted using a dynamic decoder into features similar to those described above. Frame-by-frame corresponding temporal dynamic features The system predicts video features from the VLA model feature space to the video base model feature space. Based on this, it determines frame-level feature alignment loss and temporal contrast alignment loss between the predicted features of each frame and the corresponding frame's temporal dynamic features. The frame-level feature alignment loss constrains the similarity between each frame's predicted features and the corresponding frame's ground truth features in the semantic space, while the temporal contrast alignment loss forces temporal discriminability between different frames. Together, they effectively address the problem of alignment collapse into temporally averaged static features caused by the asymmetry between single-frame VLA model and multi-frame video representations. Finally, the VLA model and dynamic decoder are jointly trained using these two losses, implicitly transferring the temporal dynamic priors inherent in the video base model to the VLA model through representation alignment. After training, the dynamic decoder and video base model are discarded; the inference phase uses only the original VLA model inference process without introducing any additional parameters or computational overhead. This approach maintains the efficient inference advantage of the VLA model while endowing it with temporal dynamic understanding capabilities, without modifying the basic VLA model architecture, and exhibits good plug-in compatibility. This optimization method can solve the problems of existing VLA models not involving temporal dynamic priors and being unable to utilize context.
[0031] The optimization system provided by this invention, based on the aforementioned robot VLA model optimization method, utilizes a coordinated approach involving a temporal feature determination module, a video feature determination module, a loss determination module, and a joint training module. It leverages the temporal dynamic priors inherent in the video base model to optimize the representation alignment of the VLA model, enabling the VLA model to possess temporal dynamic understanding capabilities after training. This system does not modify the architecture or inference process of the base VLA model and can be directly integrated into various types of VLA models, exhibiting excellent plug-and-play characteristics. This system effectively addresses the technical problem of existing VLA models, which rely solely on single-frame input, resulting in performance limitations in long-term manipulation and dynamic scenarios. Attached Figure Description
[0032] Figure 1 This is a flowchart illustrating the robot VLA model optimization method based on video representation alignment provided in an embodiment of the present invention. Detailed Implementation
[0033] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0034] The first aspect of this invention provides a method for optimizing a robot VLA model based on video representation alignment, such as... Figure 1 As shown, it includes: S1. Using video basic models to... Feature extraction is performed on the video segments to obtain... Frame temporal dynamic characteristics; The video frame contains the currently observed frame.
[0035] Specifically, S1 includes: S11, to Noise perturbation is applied to frame video segments to obtain noise latent features.
[0036] Specifically, for Noise perturbation is applied to frame video segments to obtain noise latent features, including: S111, will Frame video segments are compressed into clean latent variables using a video variational autoencoder based on the video fundamental model.
[0037] In an embodiment of the present invention, A video clip is a frame containing the currently observed frame. and its time neighborhood A video clip of a frame.
[0038] First, a frozen video base model is selected, which includes a Video Variational Autoencoder (VAE) and a Diffusion Transformer (DiT). In practical applications, the video base model can be Wan, Cosmos, etc. Then, the current observed frame will be used... Centered Frame video clips are compressed into clean latent variables using a frozen video VAE encoder. .
[0039] S112. Apply noise perturbation to the clean latent variables to obtain the noise latent features.
[0040] In this embodiment of the invention, the clean latent variables are matched along the flow matching path used during pre-training. Apply controlled noise perturbation to construct noise latent variables : (1) in, Let be a noise vector randomly drawn from a standard normal distribution, and ,in It is the identity matrix; For the preset disturbance level, and In practical applications, The empirically optimal value is 0.3.
[0041] S12. Input the language instructions and noise latent features into the diffusion Transformer of the video base model to extract the first... Intermediate features of the layer.
[0042] In this embodiment of the invention, the noise latent variable Language instructions and preset disturbance level Input to the frozen DiT, extract the first... Intermediate features of the layer : (2) This intermediate feature Compared to the first and final layers of DiT, it retains richer spatiotemporal dynamic information, making it suitable for subsequent frame-level feature alignment and temporal comparison alignment.
[0043] S13, regarding the first Spatial average pooling is performed on the intermediate features of the layer to obtain Frame-time dynamic characteristics.
[0044] In this embodiment of the invention, according to time index intermediate features Split into Frame features Spatial average pooling is performed separately to obtain Frame temporal dynamic features This serves as a signal for subsequent alignment. Frame temporal dynamic features for: (3) in, For time indexing, ; Indicates DiT's first The output of the first layer Frame features; Indicates the first Frame number Video features of a spatial location.
[0045] The temporal dynamic features of the video base model introduced in the above manner, through representation alignment and transfer to the VLA model in subsequent training, can effectively reduce the dependence of policy learning on action-annotated data. Experimental results show that, in a low-resource scenario using only 5% of the demo data, the technical solution of this invention improves the test success rate of the LIBERO-Long task from 17.2% to 50.2%, an improvement of 33 percentage points. These results demonstrate that the temporal dynamic prior provided by the video base model significantly improves the sample efficiency of the VLA model, enabling it to learn effective manipulation strategies even with limited action-annotated data.
[0046] S2, Obtain the VLA model's multi-layer attention backbone. The visual token features output by the layer are transformed using a dynamic decoder. Frame-based video prediction features; Frame video prediction features and The frame-by-frame dynamic features correspond to the time sequence.
[0047] Specifically, the VLA model includes a visual encoder, a text segmenter, and a multi-layer attention backbone. Therefore, obtaining the first... The visual token features output by the layer include: The current observation frame is encoded using a visual encoder to generate a visual token.
[0048] In this embodiment of the invention, a visual encoder is used to measure the current time. Observation frames Encode and generate Visual token ,in For index variables, ; This serves as a visual identifier to distinguish it from language tokens. In practical applications, the visual encoder can be a SigLIP pre-trained language-image encoder or a DINO encoder based on object detection.
[0049] In an embodiment of the present invention, It is an RGB image. ,in For the real number field, Image height, The width is 3, and the number of RGB color channels is 3.
[0050] Use a text segmenter to segment language instructions into language tokens.
[0051] In this embodiment of the invention, a text segmenter is used to segment language instructions. Word segmentation is Language tokens ,in For index variables, ; This serves as a language identifier, distinguishing it from visual tokens. In practical applications, the specific implementation of the text segmenter depends on the underlying VLA model used; different VLA models can be configured with different types of segmenters.
[0052] This study utilizes a multi-layered attention backbone to enable cross-modal interaction between visual and linguistic tokens, and obtains the first-order attention value of the multi-layered attention backbone. Visual token features output by the layer.
[0053] In this embodiment of the invention, the above-mentioned A visual token and Multiple language tokens are input into a multi-layer attention backbone. Cross-modal interaction is achieved through the self-attention mechanism of each attention layer, enabling visual tokens to integrate semantic information from language instructions and progressively refine the representation. The first language token is then obtained from the multi-layer attention backbone. The visual token features output by the layer are used as input to the dynamic decoder, and are converted through dynamic prediction. Frame video prediction feature sequence In practical applications, the specific architecture type of the multi-layer attention backbone depends on the underlying VLA model used; for example, a GPT-style architecture containing only a decoder can be adopted.
[0054] In this embodiment of the invention, the VLA model also includes an action head. The length of the refined representation prediction based on the output of the last layer of the multi-layer attention backbone is [length to be filled in]. Action blocks: (4) Action head The output action blocks are used to control the robot to perform corresponding manipulation actions. Among them, Indicates from time The Beginning of the Future A sequence of actions at each time step; The mapping function for the action head; The preset action block length, The value can be adjusted according to the time granularity of the actual control task. In practical applications, .
[0055] Specifically, a dynamic decoder is used to convert visual token features into... Frame video prediction features include: Amplifying the hidden dimensions of visual token features using a dynamic decoder times, .
[0056] In this embodiment of the invention, a lightweight dynamic decoder is employed. It consists of 2-4 layers of a multilayer perceptron (MLP) and uses activation functions and linear projection layers. Preferably, the MLP in this embodiment of the invention has 2 layers; the activation function is a Gaussian Error Linear Unit (GELU). In practical applications, the activation function can also be other nonlinear activation functions such as a Rectified Linear Unit (ReLU).
[0057] In this embodiment of the invention, the multi-layer attention backbone is... Each visual token feature output by the layer The dimension is MLP will Magnified 4 times, visual token features Mapped to The hidden dimension is used to expand the feature representation space to improve mapping accuracy. In practical applications, the magnification factor... As long as the dynamic decoder has sufficient capacity for feature mapping and maintains its lightweight characteristics, it is sufficient.
[0058] Magnify the hidden dimension Visual token features are projected onto the video base model. In the multiple dimensions, the token features of the video model are obtained.
[0059] In this embodiment of the invention, the hidden dimension of the video base model is . The hidden dimension is Visual token features Project to Dimension.
[0060] Reshape the video model token features into the video base model. 1 eigenvector, to obtain Frame video prediction features.
[0061] In this embodiment of the invention, the above-mentioned Remodeling Each visual token has a unique feature vector. corresponding Frame video prediction features: (5) in, For dynamic decoders The mapping function; The number of visual tokens; Indicates the first The visual token corresponds to the first The video prediction features of the frame have the following dimensions: , Remodeled The feature vectors are respectively compared with those extracted from the video base model. Frame temporal dynamic features Frame-by-frame correspondence is used for subsequent calculation of frame-level feature alignment loss and temporal comparison alignment loss.
[0062] S3. Determine the frame-level feature alignment loss and temporal contrast alignment loss between the predicted features of each video frame and the temporal dynamic features of the corresponding frame.
[0063] Specifically, the frame-level feature alignment loss and temporal contrast alignment loss between the predicted features of each video frame and the temporal dynamic features of the corresponding frame are determined, including: Determine the cosine similarity between the predicted features of each video frame and the temporal dynamic features of the corresponding frame.
[0064] The frame-level feature alignment loss and temporal comparison alignment loss are determined based on cosine similarity.
[0065] In this embodiment of the invention, frame-level feature alignment loss Maximize the cosine similarity between each predicted feature and the video feature at the corresponding time index, specifically as follows: (6) in, The number of visual tokens; This represents the total number of frames in the video segment. For the first The visual token corresponds to the first Video prediction features of frames; The first extracted from the video base model Frame number Temporal dynamic characteristics of each spatial location; The cosine similarity function is used. This loss constraint aligns the predicted features of each spatial token with the true features of the corresponding frame in the semantic space.
[0066] Timing Comparison Alignment Loss Using InfoNCE-style contrastive loss, for each spatial token Build The similarity matrix forces inter-frame temporal discriminability. Temporal contrast alignment loss. Specifically as follows: (7) in, This refers to temperature hyperparameters. The time index variable in the denominator represents the predicted first time. Frame features and all The sum of similarities between real features of frames.
[0067] In practical applications, using frame-level alignment loss or temporal contrast loss alone outperforms the baseline, while the combination of both improves the video prediction features of the VLM model. And video basic model Temporal dynamic features extracted by layer Supervision enables further improvements across all subtasks and resolves the temporal averaging problem caused by the asymmetry between single-frame VLA and multi-frame video representation.
[0068] S4. Jointly train the VLA model and dynamic decoder using frame-level feature alignment loss and temporal contrast alignment loss to obtain an optimized VLA model. This loss prevents the aligned representation from degenerating into temporally averaged static features, ensuring that features at different time steps remain distinguishable.
[0069] Specifically, the VLA model and dynamic decoder are jointly trained using frame-level feature alignment loss and temporal contrast alignment loss, including: Determine the weighted sum of frame-level feature alignment loss and temporal contrast alignment loss.
[0070] The total loss is determined based on the weighted sum and the loss due to actions.
[0071] The parameters of the VLA model and the dynamic decoder are jointly trained using the total loss.
[0072] In this embodiment of the invention, the total loss The standard action loss and the two alignment losses are combined, specifically: (8) in, The action loss is used to constrain the difference between the action blocks predicted by the VLA model and the actual actions; To control the frame-level alignment strength, the empirically optimal value of 5×10 is preferred. -1 ; To control the timing contrast strength, the empirically optimal value of 10 is preferred. -1 .
[0073] During the joint training process described above, the video base model remains frozen to ensure that the extracted temporal dynamic features serve as stable alignment signals. Without modifying the architecture and inference process of the base VLA model, only the VLA model parameters and dynamic decoder parameters are updated, allowing direct integration into the optimized OpenVLA model (OpenVLA with Optimized Fine-Tuning, OpenVLA-OFT). , It achieves consistent gains in both simulation and real-world tasks, and is plug-and-play.
[0074] By constraining the frame-level feature alignment loss to align predicted features with real video features frame-by-frame in the semantic space, and by forcing temporal discriminative alignment loss to maintain temporal distinguishability between different frames, this effectively solves the temporal averaging problem caused by the asymmetry between single-frame VLA and multi-frame video representations. This allows the VLA model to effectively transfer temporal dynamic priors from the video base model, thereby significantly improving its success rate in various manipulation tasks. For example, on the LIBERO simulation benchmark, the technical solution of this invention increases the average success rate of OpenVLA-OFT from 97.1% to 98.2%. The average success rate increased from 94.2% to 97.4%. The average success rate increased from 96.9% to 98.3%; in the RoboTwin 2.0 dual-arm simulation platform, The average success rate increased from 45.0% to 53.0%; in real robot experiments, the average success rate increased from 47.1% to 65.0%, an increase of about 18 percentage points.
[0075] The temporal dynamic features obtained in the above manner possess inherent anti-disturbance properties. Experimental verification shows that, under seven types of disturbance dimensions (no obstructions, illumination changes, background changes, camera viewpoint changes, etc.) in the LIBERO-Plus benchmark, the technical solution of this invention effectively... The model's average success rate increased from 70.9% to 72.4%, maintaining consistent gain under various out-of-distribution perturbations, and achieving the aforementioned robustness without additional training on perturbation data.
[0076] In this embodiment of the invention, after joint training is completed, the dynamic decoder and the video base model no longer participate in the inference process. During inference, only the trained VLA model is retained and runs according to its original inference flow: Input the current observation frame With language instructions .
[0077] Action blocks are generated using the visual encoder, multi-layer attention backbone, and action head of the VLA model.
[0078] The robot executes action blocks to complete the control task.
[0079] This invention implicitly transfers the temporal dynamic priors inherent in the video base model to the VLA model through representation alignment during training. Since the inference phase does not involve the dynamic decoder and the video base model, the inference latency of the VLA model is completely consistent with that of the base VLA model. Compared to the Video-Action Model (VAM), which requires iterative generation of future frames during the inference phase and has an inference latency of 403ms, this invention's embodiment has an inference latency of only 61ms, making it more than 6 times faster. The actual deployment cost is extremely low, making it suitable for robot control deployment scenarios with high real-time requirements.
[0080] A second aspect of this invention provides an optimization system based on the above-described robot VLA model optimization method, comprising: The temporal feature determination module is used to utilize the video base model to determine... Feature extraction is performed on the video segments to obtain... Frame temporal dynamic characteristics; The video frame contains the currently observed frame.
[0081] The video feature determination module is used to obtain the first multi-layer attention backbone of the VLA model. The visual token features output by the layer are transformed using a dynamic decoder. Frame-based video prediction features; Frame video prediction features and The frame-by-frame dynamic features correspond to the time sequence.
[0082] The loss determination module is used to determine the frame-level feature alignment loss and temporal contrast alignment loss between the predicted features of each video frame and the temporal dynamic features of the corresponding frame.
[0083] The joint training module is used to jointly train the VLA model and the dynamic decoder using frame-level feature alignment loss and temporal contrast alignment loss to obtain an optimized VLA model.
[0084] The robot VLA model optimization method based on video representation alignment provided in this invention first utilizes the video base model to optimize the VLA model including the current observation frame. Feature extraction is performed on the video segments to obtain... Frame-time dynamic features serve as alignment signals, providing a supervisory target that includes prior knowledge of object motion patterns and physical interactions for subsequent alignment representations. Secondly, the VLA model's multi-layer attention backbone is obtained... The visual token features output by the layer are then converted using a dynamic decoder into features similar to those described above. Frame-by-frame corresponding temporal dynamic features The system predicts video features from the VLA model feature space to the video base model feature space. Based on this, it determines frame-level feature alignment loss and temporal contrast alignment loss between the predicted features of each frame and the corresponding frame's temporal dynamic features. The frame-level feature alignment loss constrains the similarity between each frame's predicted features and the corresponding frame's ground truth features in the semantic space, while the temporal contrast alignment loss forces temporal discriminability between different frames. Together, they effectively address the problem of alignment collapse into temporally averaged static features caused by the asymmetry between single-frame VLA model and multi-frame video representations. Finally, the VLA model and dynamic decoder are jointly trained using these two losses, implicitly transferring the temporal dynamic priors inherent in the video base model to the VLA model through representation alignment. After training, the dynamic decoder and video base model are discarded; the inference phase uses only the original VLA model inference process without introducing any additional parameters or computational overhead. This approach maintains the efficient inference advantage of the VLA model while endowing it with temporal dynamic understanding capabilities, without modifying the basic VLA model architecture, and exhibits good plug-in compatibility. This optimization method can solve the problems of existing VLA models not involving temporal dynamic priors and being unable to utilize context.
[0085] The optimization system based on the above-described robot VLA model optimization method provided in this invention utilizes a temporal feature determination module, a video feature determination module, a loss determination module, and a joint training module to collaboratively optimize the representation alignment of the VLA model using the temporal dynamic priors inherent in the video base model. This enables the VLA model to possess temporal dynamic understanding capabilities after training. This system does not modify the architecture and inference process of the base VLA model and can be directly integrated into various types of VLA models, exhibiting excellent plug-and-play characteristics. This system addresses the technical problem of existing VLA models, which rely solely on single-frame input, resulting in limited performance in long-term manipulation and dynamic scenes.
[0086] The above description is merely a few embodiments of this application and is not intended to limit this application in any way. Although this application discloses preferred embodiments as described above, it is not intended to limit this application. Any changes or modifications made by those skilled in the art without departing from the scope of the technical solution of this application using the disclosed technical content are equivalent to equivalent implementation cases and fall within the scope of the technical solution.
Claims
1. A robot VLA model optimization method based on video representation alignment, characterized in that, include: S1. Using video basic models to... Feature extraction is performed on the video segments to obtain... Frame temporal dynamic characteristics; The A video clip contains the currently observed frame; S2, Obtain the VLA model's multi-layer attention backbone. The visual token features output by the layer are converted using a dynamic decoder. Frame video prediction features; the Frame video prediction features and the Frame-by-frame correspondence of temporal dynamic features; S3. Determine the frame-level feature alignment loss and temporal contrast alignment loss between the predicted features of each video frame and the temporal dynamic features of the corresponding frame; S4. The VLA model and the dynamic decoder are jointly trained using the frame-level feature alignment loss and the temporal contrast alignment loss to obtain the optimized VLA model.
2. The robot VLA model optimization method based on video representation alignment according to claim 1, characterized in that, Determine the frame-level feature alignment loss and temporal contrast alignment loss between the predicted features of each video frame and the temporal dynamic features of the corresponding frame, including: Determine the cosine similarity between the predicted features of each video frame and the temporal dynamic features of the corresponding frame; The frame-level feature alignment loss and the temporal comparison alignment loss are determined based on the cosine similarity.
3. The robot VLA model optimization method based on video representation alignment according to claim 2, characterized in that, The VLA model and the dynamic decoder are jointly trained using the frame-level feature alignment loss and the temporal contrast alignment loss, including: Determine the weighted sum of the frame-level feature alignment loss and the temporal contrast alignment loss; The total loss is determined based on the weighted sum and the action loss. The parameters of the VLA model and the parameters of the dynamic decoder are jointly trained using the total loss.
4. The robot VLA model optimization method based on video representation alignment according to claim 1, characterized in that, The VLA model includes a visual encoder, a text segmenter, and a multi-layer attention backbone. The VLA model's multi-layer attention backbone is then obtained. The visual token features output by the layer include: The current observation frame is encoded using a visual encoder to generate a visual token; Use a text segmenter to segment language instructions into language tokens; The visual token and the language token are used to perform cross-modal interaction using a multi-layer attention backbone, and the first multi-layer attention backbone is obtained. Visual token features output by the layer.
5. The robot VLA model optimization method based on video representation alignment according to claim 1, characterized in that, The visual token features are converted using a dynamic decoder. Frame video prediction features include: The hidden dimension of the visual token features is amplified using a dynamic decoder. times, the ; Magnify the hidden dimension Projecting the visual token features of multiples onto the video base model In the multiple dimensions, the token features of the video model are obtained; Reshape the video model token features into the video base model. 1 eigenvector, to obtain Frame video prediction features.
6. The robot VLA model optimization method based on video representation alignment according to claim 1, characterized in that, S1 includes: S11, regarding the above Noise perturbation is applied to frame video segments to obtain noise latent features; S12. Input the language instructions and the noise latent features into the diffusion Transformer of the video base model to extract the first... Intermediate features of the layer; S13, regarding the first Spatial average pooling is performed on the intermediate features of the layer to obtain Frame-time dynamic characteristics.
7. The robot VLA model optimization method based on video representation alignment according to claim 6, characterized in that, Regarding the Noise perturbation is applied to frame video segments to obtain noise latent features, including: S111, the above Frame video segments are compressed into clean latent variables using a video variational autoencoder based on the video fundamental model. S112. Apply noise perturbation to the clean latent variable to obtain the noise latent feature.
8. An optimization system based on the video representation alignment-based robot VLA model optimization method according to any one of claims 1-7, characterized in that, include: The temporal feature determination module is used to utilize the video base model to determine... Feature extraction is performed on the video segments to obtain... Frame temporal dynamic characteristics; The A video clip contains the currently observed frame; The video feature determination module is used to obtain the first multi-layer attention backbone of the VLA model. The visual token features output by the layer are converted using a dynamic decoder. Frame video prediction features; the Frame video prediction features and the Frame-by-frame correspondence of temporal dynamic features; The loss determination module is used to determine the frame-level feature alignment loss and temporal contrast alignment loss between the predicted features of each video frame and the temporal dynamic features of the corresponding frame. The joint training module is used to jointly train the VLA model and the dynamic decoder using the frame-level feature alignment loss and the temporal contrast alignment loss to obtain an optimized VLA model.