Robot operation method, system, device and medium based on vla model

CN122473720BActive Publication Date: 2026-09-08UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610968516.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-01
Publication Date
2026-09-08
Estimated Expiration
2046-07-01

AI Technical Summary

Technical Problem

[0006]现有方案的缺陷在于:(1)推理时仍需输出显式定位信息:模型每一步均需先输出二维点、边界框、路径,再据此生成动作;这部分输出会显著增加推理延迟,并占用额外的解码资源;(2)定位错误会传导至动作:当模型预测的定位标签出现偏移或漂移时,后续动作生成将基于错误的空间先验执行,整体鲁棒性下降;(3)部署链路冗长:依赖外部检测器或显式定位输出的方案,在实际部署时需要维护额外的模型与数据通路,工程复杂度高

Benefits of technology

[0014]As can be seen from the technical solution provided by the present invention: injecting bounding boxes as visual cues visible during training and removed during inference into the image itself enables the VLA model to learn implicit spatial localization capabilities during training; during inference, no localization labels need to be output and no external detector is required, fundamentally eliminating the serial dependency between localization and action; furthermore, by first training the visual language model separately and then adding the action head for joint training, and gradually reducing the proportion of data with given bounding boxes in the later stages of training, a smooth transition in the distribution of training and inference inputs is achieved, so that the VLA model has seen a large number of samples without localization input by the end of training; thanks to the above improvements, the VLA model obtains explicit localization cues during the training phase and works independently during the inference phase, thereby injecting spatial localization capabilities into the model without changing the existing VLA model computation graph or increasing any inference overhead. Compared with the existing schemes that output and display localization information during inference, the present invention saves the localization decoding overhead in each step of inference; compared with the schemes that rely on external detectors, the present invention no longer relies on external detection or segmentation modules on the deployment side, and the end-to-end inference chain is consistent with the existing VLA model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122473720B_ABST
    Figure CN122473720B_ABST
Patent Text Reader

Abstract

The application discloses a robot operation method, system, device and medium based on a VLA model, which are corresponding solutions, in which a bounding box is taken as a visual prompt which is visible during training and is removed during reasoning and is injected into an image itself, so that the VLA model learns an implicit spatial positioning capability during training; positioning information does not need to be output during reasoning, and an external detector is not relied on, so that serial dependence of positioning and action is eliminated; and through separate training of a visual language model first, joint training by adding an action head, and then gradual reduction of a data proportion of the bounding box, smooth transition of training and reasoning input distribution is realized, so that the VLA model has seen a large number of samples without positioning at the end of training; based on this, the VLA model obtains an explicit positioning prompt during the training stage and works independently during the reasoning stage, so that spatial positioning capability is injected into the model under the premise that a calculation graph of the existing VLA model is not changed and any reasoning overhead is not increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot manipulation technology, and in particular to a robot manipulation method, system, device and medium based on a VLA model. Background Technology

[0002] The concept of deep learning originated from the research on artificial neural networks. In recent years, deep learning has been widely used in robotic embodied intelligence tasks, enabling robots to directly translate natural language instructions into low-level actions.

[0003] Vision-Language-Action (VLA) models are end-to-end policy models that use a Visual Language Model (VLM) as their backbone and output robot actions (referred to as output actions) through an action head. They receive images and natural language commands as input and output low-level control actions of the robot (such as end effector displacement, gripper opening and closing, etc.), directly driving the robot to complete tasks such as grasping, placing, and assembling.

[0004] In robot manipulation, the model needs to identify and locate the object or operating area indicated by the command. Existing methods typically require additional explicit localization information (points, bounding boxes, paths, etc.) to be output during the inference phase, or rely on external detectors to provide this information during inference. This means that the model needs to "point out" the operating area before generating the action at each step, which not only increases inference latency, but also allows errors in localization to propagate to action generation.

[0005] Furthermore, existing solutions propose action reasoning models that break down the workflow of the robot's basic model into three stages: perception, planning, and control. The VLM model backbone explicitly outputs depth perception tokens and two-dimensional visual path tokens on the image plane, and then generates actions based on these explicit localization information. Both the training and reasoning stages require the model to explicitly output the aforementioned localization information. For existing solutions in this regard, please refer to the literature: Jason LEE et al., MolmoAct: Action Reasoning Models that can Reason in Space, arXiv:2508.07917, 2025.9.18.

[0006] The shortcomings of the existing solutions are: (1) explicit location information still needs to be output during inference: the model needs to output two-dimensional points, bounding boxes and paths for each step, and then generate actions based on these; this output will significantly increase the inference latency and occupy additional decoding resources; (2) location errors will be transmitted to actions: when the location labels predicted by the model are offset or drifted, the subsequent action generation will be based on the incorrect spatial prior, and the overall robustness will decrease; (3) the deployment chain is lengthy: the solution that relies on external detectors or explicit location output needs to maintain additional models and data paths during actual deployment, which is highly complex.

[0007] In view of this, the present invention is hereby proposed. Summary of the Invention

[0008] The purpose of this invention is to provide a robot operation method, system, device, and medium based on a VLA model. Through training, the VLA model can acquire spatial positioning capabilities. During inference, it does not output any explicit positioning tags or rely on any external detectors, thereby accurately completing robot operation tasks without increasing inference overhead.

[0009] The objective of this invention is achieved through the following technical solution:

[0010] A robot manipulation method based on a VLA model includes: The demonstration video is divided into video segments corresponding to each sub-task. For each video segment, the task interaction region bounding box of each original RGB image is extracted and then superimposed on the corresponding original RGB image to obtain the superimposed image. The original prompts are added with prompts containing the focus bounding box region to obtain focus prompts. The superimposed image and focus prompts are used to construct the first type of training data, and the original RGB image and original prompts are used to construct the second type of training data. The VLA model, a visual language action model, is trained using two types of training data. The VLA model comprises a visual language model and an action head. The training process is as follows: Stage 1: Input the first type of training data and train the visual language model based on the two-dimensional pixel trajectory words output by the visual language model. Stage 2: Input the first type of training data and jointly train the visual language model and the action head based on the two-dimensional pixel trajectory words output by the visual language model and the output actions of the action head. Stage 3: Input both types of training data, gradually decrease the proportion of the first type of training data while gradually increasing the proportion of the second type of training data along a linear curve, and train using the same method as in Stage 2. After training, the original RGB image and the original prompt words are input into the VLA model to obtain the output action.

[0011] A robot operating system based on the VLA model includes: The training data construction unit is used to divide the demonstration video into video segments corresponding to each sub-task. For each video segment, the task interaction region bounding box of each original RGB image is extracted and then superimposed on the corresponding original RGB image to obtain the superimposed image. The original prompts are added with prompts containing the focus bounding box region to obtain focus prompts. The superimposed image and focus prompts are used to construct the first type of training data, and the original RGB image and original prompts are used to construct the second type of training data. The model training unit is used to train the VLA model using two types of training data. The VLA model is a visual language action model, which includes a visual language model and an action head. The training process is as follows: Stage 1: Input the first type of training data and train the visual language model based on the two-dimensional pixel trajectory words output by the visual language model; Stage 2: Input the first type of training data and jointly train the visual language model and the action head based on the two-dimensional pixel trajectory words output by the visual language model and the output actions of the action head; Stage 3: Input the two types of training data, and gradually reduce the proportion of the first type of training data while gradually increasing the proportion of the second type of training data according to the linear curve of the training steps, and train in the same way as in Stage 2. The action prediction unit is used to input the original RGB image and the original prompt words into the VLA model after training to obtain the output action.

[0012] A processing device includes: one or more processors; and a memory for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method.

[0013] A readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned method.

[0014] As can be seen from the technical solution provided by the present invention: injecting bounding boxes as visual cues visible during training and removed during inference into the image itself enables the VLA model to learn implicit spatial localization capabilities during training; during inference, no localization labels need to be output and no external detector is required, fundamentally eliminating the serial dependency between localization and action; furthermore, by first training the visual language model separately and then adding the action head for joint training, and gradually reducing the proportion of data with given bounding boxes in the later stages of training, a smooth transition in the distribution of training and inference inputs is achieved, so that the VLA model has seen a large number of samples without localization input by the end of training; thanks to the above improvements, the VLA model obtains explicit localization cues during the training phase and works independently during the inference phase, thereby injecting spatial localization capabilities into the model without changing the existing VLA model computation graph or increasing any inference overhead. Compared with the existing schemes that output and display localization information during inference, the present invention saves the localization decoding overhead in each step of inference; compared with the schemes that rely on external detectors, the present invention no longer relies on external detection or segmentation modules on the deployment side, and the end-to-end inference chain is consistent with the existing VLA model. Attached Figure Description

[0015] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 A flowchart illustrating a robot operation method based on a VLA model, provided as an embodiment of the present invention.

[0017] Figure 2 This is a schematic diagram of the overall process of the training phase and the inference phase provided in an embodiment of the present invention; wherein (a) is the training phase and (b) is the inference phase.

[0018] Figure 3 This is a schematic diagram of the three-stage training process of the VLA model provided in an embodiment of the present invention.

[0019] Figure 4 This is a schematic diagram illustrating the proportion of the first type of training data provided in an embodiment of the present invention.

[0020] Figure 5 This is a schematic diagram illustrating the application scenario of the VLA model of the present invention as provided in an embodiment of the present invention.

[0021] Figure 6 This is a schematic diagram of a robot operating system based on the VLA model provided in an embodiment of the present invention.

[0022] Figure 7This is a schematic diagram of a processing device provided in an embodiment of the present invention. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.

[0024] First, the following explanations are provided for the terms that may be used in this article: The terms "comprising," "including," "containing," "having," or other similar semantic descriptions should be interpreted as non-exclusive inclusion. For example, including a technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction conditions, processing conditions, parameter, algorithm, signal, data, product or article of manufacture, etc.) should be interpreted as including not only the expressly listed technical feature element, but also other technical feature elements that are not expressly listed and are well-known in the art.

[0025] The term "composed of" excludes any technical features not expressly listed. When used in a claim, it closes the claim to exclude all technical features other than those expressly listed, except for associated conventional impurities. If the term appears only in a clause of a claim, it limits the claim to the elements expressly listed in that clause; elements recited in other clauses are not excluded from the overall claim.

[0026] The following provides a detailed description of a robot operation method, system, device, and medium based on a VLA model provided by this invention. Contents not described in detail in the embodiments of this invention are prior art known to those skilled in the art. Where specific conditions are not specified in the embodiments of this invention, they are performed according to conventional conditions in the art or conditions recommended by the manufacturer. Instruments used in the embodiments of this invention, unless otherwise specified by the manufacturer, are all commercially available conventional products.

[0027] Example 1 like Figure 1 The flowchart shown is a robot operation method based on a VLA model provided by an embodiment of the present invention, which mainly includes the following steps: Step 1: Construct training data.

[0028] In this embodiment of the invention, each demonstration video is divided into several video segments, and each video segment corresponds to a subtask. For each video segment, the corresponding subtask is assigned semantics. Using the semantics as a query, the task interaction region bounding box of each original RGB image in the video segment is obtained and then superimposed on the original RGB image to obtain a superimposed image. A prompt word containing the focus bounding box region is added to the original prompt word to obtain a focus prompt word. The superimposed image and the focus prompt word are used to construct the first type of training data, and the original RGB image and the original prompt word are used to construct the second type of training data.

[0029] In this embodiment of the invention, a preferred method for obtaining the task interaction region bounding box is as follows: Each demonstration video is divided into several video segments based on gripper state transitions or with the aid of an external visual language model; semantics are assigned to the sub-tasks corresponding to each video segment; for each sub-task, the corresponding interactive target object is extracted based on the sub-task semantics; using the interactive target object as the query, an open-word pointing model is used to output a representative single point on the keyframe of the sub-task, where the representative single point is the two-dimensional coordinate position of the interactive target object in the keyframe; a point-hint segmentation model is used to expand the representative single point into a segmentation mask and extract the circumscribed rectangle to obtain the task interaction region bounding box of the corresponding keyframe; the task interaction region bounding boxes of the remaining intermediate frames are obtained by linear interpolation of adjacent keyframes, where each frame here is the original RGB image in the video segment.

[0030] Those skilled in the art will understand that dividing the demonstration video into several video segments is primarily based on the task process performed by the robot, so that each video segment corresponds to a separate sub-task. For example, in the task of placing a can of cola into a basket, the gripper approaching and grabbing the cola can be considered a sub-task, and the gripper placing the cola into the basket can be considered another sub-task. The segmentation process can be implemented based on gripper state transitions or with the help of an external visual language model. Subsequently, keyframes can be determined based on semantic changes in the video segments, and specific details can be found using conventional techniques, which will not be elaborated upon in this invention.

[0031] In this embodiment of the invention, the superimposition onto the original RGB image includes: superimposing the task interaction area bounding box with set color parameters, line type parameters and line width parameters onto the original RGB image.

[0032] Step 2: Train the VLA model.

[0033] In this embodiment of the invention, the VLA model (Visual Language Action Model) mainly includes a Visual Language Model (VLM) and an action head. Its training process comprises three stages: Stage 1: Inputting the first type of training data into the VLA model and supervising its training based on the two-dimensional pixel trajectory words output by the Visual Language Model in the VLA model; Stage 2: Inputting the first type of training data into the VLA model to obtain the two-dimensional pixel trajectory words output by the Visual Language Model and the output action of the action head, and jointly training the Visual Language Model and the action head using the two-dimensional pixel trajectory words and the output action; Stage 3: Inputting both types of training data into the VLA model, gradually reducing the proportion of the first type of training data along a linear curve as the time step progresses, while gradually increasing the proportion of the second type of training data, and jointly training the Visual Language Model and the action head using the same method as in Stage 2. That is, the Visual Language Model predicts the two-dimensional pixel trajectory words based on the input training data (first or second type of training data), and the action head predicts the output action, thereby jointly training the Visual Language Model and the action head.

[0034] In this embodiment of the invention, supervised training of the visual language model based on the two-dimensional pixel trajectory words output by the visual language model in the VLA model includes: inputting a first type of training data, and having the visual language model output two-dimensional pixel trajectory words that plan the gripper's path from the current position to the next subtask path point, wherein the current position refers to the position of the task interaction region bounding box in the first type of training data; and combining the two-dimensional pixel trajectory words output by the visual language model with the pre-acquired ground truth values ​​of the subtask path points to calculate the word-by-word cross-entropy loss for autoregressive supervision of the visual language model.

[0035] In this embodiment of the invention, the joint training of the visual language model and the action head based on the two-dimensional pixel trajectory words output by the visual language model and the output action of the action head includes: inputting a first type of training data; the visual language model outputting two-dimensional pixel trajectory words from the gripper to the next sub-task path point, where the current position refers to the position of the task interaction region bounding box in the first type of training data; the action head iteratively denoising and generating the output action based on the hidden state of the specified words in the last layer of the visual language model; combining the two-dimensional pixel trajectory words output by the visual language model and the pre-acquired ground truth values ​​of the sub-task path points, calculating the word-by-word cross-entropy loss and the flow matching loss of the output action, and combining the two losses to jointly train the visual language model and the action head.

[0036] In this embodiment of the invention, the true value of the subtask path point is obtained in the following way: for each video segment, the position of the gripper in the original RGB image is located by gripper feature matching, and then its two-dimensional pixel coordinates are obtained by segmentation, which are used as the path point of the subtask. The pixel sequence formed by connecting adjacent path points constitutes the true value of the subtask path point.

[0037] In this embodiment of the invention, the method of gradually reducing the proportion of the first type of training data and gradually increasing the proportion of the second type of training data along a linear curve during training steps includes: defining stage three as starting from training steps... Begin, training steps End. Input the first type of training data with probability p(t) and the second type of training data with probability 1-p(t), then: Current training step At that time, p(t) = 1.0; The current training step t satisfies hour, .

[0038] Step 3: VLA model inference.

[0039] In this embodiment of the invention, after training is completed, the original RGB image and the original prompt words are input into the VLA model to obtain the output action.

[0040] In this embodiment of the invention, an RGB image and a prompt word are input. The hidden state of a specified word is obtained through the forward processing of a visual language model. Then, an action head uses a DiT (Diffusion Transformer)-based flow matching module for iterative denoising to obtain the output action of the VLA model. The output action is processed in the same way during training and inference. Therefore, the RGB image and prompt word here can refer to the overlaid image and focused prompt word, or it can refer to the original RGB image and original prompt word.

[0041] The above-described solution of this invention enables the VLA model to receive explicit localization cues during the training phase and operate independently during the inference phase. This injects spatial localization capabilities into the VLA model while maintaining its computational graph and without increasing inference overhead. All modifications are concentrated on the training side, making it widely integrateable as a training-time enhancement plugin for existing VLA models. Compared to existing solutions that require outputting localization labels during inference, this invention saves the localization decoding overhead at each inference step. Compared to solutions that rely on external detectors, this invention no longer depends on external detection or segmentation modules on the deployment side, and the end-to-end inference chain is consistent with existing VLA models.

[0042] To more clearly demonstrate the technical solution and its effects provided by the present invention, the method provided by the embodiments of the present invention will be described in detail below with reference to specific examples.

[0043] I. Overall Overview of the Plan

[0044] 1. Training phase.

[0045] First, for each demonstration video, the video is segmented into several sub-tasks and assigned semantics based on the gripper state transition or with the aid of a large vision-language model. Then, using the sub-task semantics as the query, an open-word pointing model and a point-based prompting segmentation model are used to automatically obtain the bounding boxes of the task interaction regions for each frame. Next, the bounding boxes of the task interaction regions are directly drawn onto the RGB image as a superposition of bounding boxes with visible lines (e.g., solid red lines, 2-3 pixels wide), and the prompt "Focus on the region within the bounding box" is appended after the original prompt. This forms the first type of training data. The RGB image without bounding boxes and the original prompt form the second type of training data.

[0046] Training is divided into three phases: Phase 1 trains only the visual language model, with the supervision signal consisting of the two-dimensional pixel trajectory of the subtask text and the path point of the next subtask; Phase 2 adds an action head for joint fine-tuning with the visual language model, while supervising low-order actions (i.e., output actions); Phase 3 gradually reduces the proportion of the first type of training data from 100% to a certain percentage (e.g., 20%) along a linear curve, allowing the model to smoothly adapt to the "bounded box input" working mode.

[0047] 2. During the inference phase, no bounding boxes are drawn on the input image, and the text prompts do not contain any focus descriptions. The action head can be directly driven to output actions based solely on the original RGB image and the prompt words, without outputting any explicit localization labels or calling any external detectors or segmentation models. The above method enables the VLA model to have implicit spatial localization capabilities during inference, and the inference path is consistent with the VLA model without the modifications of this invention.

[0048] Based on the above description, the present invention mainly solves the following technical problems: (1) Existing VLA models still need to output explicit localization information (points / boundary boxes / paths) during inference, or still need to receive localization information provided by external detectors. This invention can solve the problem of how to enable the model to have spatial localization capabilities, but not output any explicit localization labels or rely on any external detectors during the inference stage, so as to complete the robot operation task without increasing the inference overhead.

[0049] (2) When the explicit positioning information prediction is offset or drifted, the error will propagate to the action generation. The present invention can solve the problem of how to internalize the positioning capability in the model representation in an implicit way, so as to fundamentally avoid this cascading error of positioning first and then action.

[0050] (3) Relying on the method of providing localization supervision during training and providing localization input during inference makes it difficult to break free from the constraint of consistent training-inference input distribution. This invention can solve the problem of how to smoothly transition the model to a working mode with no localization input during inference under the asymmetric setting of "providing localization prompts during training and no localization prompts during inference".

[0051] To visually demonstrate the two core technical aspects of this invention, the following explanation will focus on the differences between this invention and existing technologies.

[0052] (1) In the prior art, it is necessary to continuously output explicit localization labels or receive external localization input during inference. The serial link of localization first and then action causes inference overhead and cascaded error. In this invention, the bounding box is introduced on the training side and removed on the inference side. Specifically, the bounding box is injected into the image itself as a visual cue that is visible during training and removed during inference, so that the model learns implicit spatial localization ability during training; during inference, there is no need to output any localization labels or rely on any external detector, which fundamentally eliminates the serial dependency between localization and action and reduces inference overhead.

[0053] (2) In the prior art, both training and inference involve or do not involve location information input, and there is no mechanism to smoothly transition the model from "depending on location input" to "no location input". This invention achieves a smooth transition in the distribution of training and inference inputs by first training the visual language model separately, then adding the action head for joint training, and gradually reducing the proportion of the first type of training data in the later stage of training. This allows the model to have seen a large number of samples without location input (i.e., the second type of training data) by the end of training, thereby solving the problem of input distribution offset during training and inference, and ensuring that the performance of the model will not collapse due to input differences when deployed.

[0054] II. Detailed introduction of the plan.

[0055] 1. Obtaining the bounding box of the task interaction area.

[0056] First, the demonstration video is divided into several sub-tasks and given semantic descriptions based on the gripper state transition or with the help of an external visual language model (e.g., Google's Gemini model). Then, the interactive target objects corresponding to the sub-tasks are extracted from the semantics of the sub-tasks. Next, an open-word pointing model (e.g., the MolmoPoint model) is used with the target object as the query to output a representative single point on the keyframe of the sub-task. Finally, a point-based segmentation model (e.g., the SAM 3.1 model) expands this single point into a segmentation mask and takes the bounding rectangle to obtain the bounding box of the task interaction region of the frame. The bounding boxes of intermediate frames are obtained by linear interpolation of adjacent keyframes. The entire process does not require manual annotation. The specific sub-task segmentation granularity, model selection, call frequency, interpolation method, etc., are conventional engineering implementations and are not within the scope of protection of this invention.

[0057] The MolmoPoint model mentioned above is an open word pointing model launched by the Allen Institute for Artificial Intelligence, and the SAM 3.1 model is version 3.1 of the Segment Everything Model (SAM).

[0058] 2. Assembling the input bounding box and focus prompt.

[0059] The automatically obtained bounding boxes are drawn as superimposed bounding boxes with red solid lines and a line width of 2-3 pixels onto the corresponding coordinates of the original RGB image I. Above, we obtain the superimposed image I', where... The coordinates of the two diagonal vertices of the bounding box are given; at the same time, a prompt word containing the area to focus on the bounding box is constructed (for example, please focus on the area inside the bounding box). The prompt word no longer contains coordinates, and is concatenated with the original prompt word to form the focus prompt word.

[0060] 3. Construct training data.

[0061] This invention mainly involves two types of training data: the first type of training data is constructed using superimposed images and focus prompts, and the second type of training data is constructed using original RGB images and original prompts.

[0062] 4. Supervision information.

[0063] The VLA model includes a visual language model and an action head, with different supervision information set during training.

[0064] (1) The 2D supervision path of the visual language model, together with the task interaction area bounding box and focus prompt words mentioned above, constitutes a complete processing link.

[0065] The first type of training data is input into the visual language model within the VLA model, guiding the VLA model to constrain its attention to the corresponding task interaction region. Under the condition of inputting the first type of training data, the visual language model outputs a two-dimensional pixel trajectory of the gripper from its current position to the next subtask path point. This trajectory is represented in the form of a discrete token sequence and is directly supervised by the language model (LLM) backbone within the visual language model using per-token cross-entropy comparison with the ground truth of the path point, without adding a separate trajectory decoding head. Thus, the bounding box and focus cue words provide spatial priors for "where to operate," while the two-dimensional pixel trajectory provides planning supervision for "how to move." The two are coupled in the same forward process, jointly constraining the model's implicit spatial localization capability.

[0066] The supervision labels (i.e., the ground truth values ​​of path points for each subtask) of the above two-dimensional pixel trajectories are obtained offline as follows: For grasping-placement tasks, the boundaries of subtasks are directly identified by the transition frames of the gripper's opening and closing states; for other types of operation tasks, where it is difficult to segment subtasks using gripper states, the boundary frames and semantic descriptions of subtasks are automatically labeled by an external visual language model (e.g., Gemini). The above work has already been completed in the previous task interaction area boundary box acquisition stage.

[0067] Within each subtask interval (i.e., a single video segment), DINOv3 is first used for gripper feature matching to locate the gripper's position in the original RGB image. Then, SAM is used for fine segmentation to obtain its two-dimensional pixel coordinates, which serve as the path points for that subtask. The pixel sequence formed by connecting adjacent path points constitutes the supervised ground truth for the trajectory. Here, DINOv3 is an abbreviation for self-DIstillation with NO labels, and "v3" means the third generation / third version. It is a self-supervised visual foundation model (visual feature extractor) proposed by Meta, which can learn general image features without manual annotation.

[0068] For example, the visual language model in the VLA model can be the Qwen2.5-VL-3B model, which is a visual language model launched by Alibaba. Qwen is the model name, 2.5 is the version number, VL refers to visual language, and 3B refers to the number of parameters of 3 billion.

[0069] (2) Action head and low-order action supervision: The action head adopts a DiT-based flow matching module, which receives the hidden state of the specified token in the last layer of the visual language model as a condition, and generates low-order actions (e.g., 6-DOF end effector pose increment and 1-dimensional gripper opening and closing) through iterative denoising.

[0070] 5. Training process.

[0071] like Figure 2 As shown, this illustrates the overall process of the training and inference phases.

[0072] In this embodiment of the invention, the training can be divided into three stages, such as... Figure 3 As shown.

[0073] Phase One (Training Steps) ): Input the first type of training data, a 2D supervised path for the visual language model, and supervise the training of the visual language model. No action head is added at this stage. That is, input the first type of training data, and the visual language model outputs the two-dimensional pixel trajectory words of the gripper from the current position to the next sub-task path point. The current position here refers to the two-dimensional pixel position of the gripper predicted by the visual language model from the first frame of the corresponding video segment. Combine the two-dimensional pixel trajectory words output by the visual language model with the pre-acquired ground truth values ​​of the sub-task path points, calculate the word-by-word cross-entropy loss to perform autoregressive supervision on the visual language model. Only the visual language model is trained at this stage, and the purpose is to first establish spatial localization and trajectory planning capabilities. The proportion of the first type of training data is 100% at this stage.

[0074] Phase Two (Training Steps) ): Input is the same as in Stage 1; the visual language model and action head are jointly trained using a 2D supervised path and low-order actions (flow matching loss); in this stage, the action head iteratively denoises and generates low-order actions (output actions) based on the hidden states of specified lexical units in the last layer of the visual language model, calculates the flow matching loss of the low-order actions, and the word-by-word cross-entropy loss introduced in Stage 1, thereby jointly training the visual language model and action head, that is, while preserving spatial localization supervision, allowing the hidden states to learn to drive the generation of specific actions. The proportion of the first type of training data in this stage is 100%.

[0075] Phase Three (Training Steps) (Linear decay of the first type of training data): Input the first type of training data with probability p(t), and input the second type of training data with probability 1-p(t). Jointly train the visual language model and the action head using a 2D supervised path of the visual language model and low-order actions (flow matching loss). The scheduling (linear decay) is: p(t) = 1.0. ; , Because a second type of training data is introduced in this stage, when calculating the current position using the 2D supervised path of the visual language model, if the input is the first type of training data, the first frame is the first frame containing the bounding box of the task interaction region (that is, the overlay image corresponding to the first frame); if it is the second type of training data, the first frame is the original RGB image.

[0076] like Figure 4 The figure shows the curves of the proportion of the first type of training data in three stages. The 20% proportion decay in stage three is just an example provided here. In actual applications, the proportion can be set according to the actual situation.

[0077] 6. Reasoning process.

[0078] (A1) Input: Original RGB image (without any bounding boxes drawn in the image) and original prompt words (without prompt words that focus on the bounding box region). Of course, the actual input also includes the robot state, but considering that this part of the process is all conventional technology, it will not be described in detail.

[0079] (A2) Obtain the hidden state of the specified token in the last layer through the forward process of the visual language model.

[0080] (A3) Based on the hidden state of the specified token, the action is obtained by iteratively denoising through the forward process of the action head.

[0081] (A4) Execute the action; the loop returns to step (A1) until the task is completed or the maximum number of steps is reached.

[0082] During inference, the image retains its original RGB color and no bounding box is drawn; no external detectors or segmentation models are called; no explicit localization labels are generated; the overall inference path is completely consistent with VLA without the modifications of this invention.

[0083] III. Application Scenarios

[0084] like Figure 5 As shown, the above-described solutions provided in the embodiments of the present invention are applicable to various scenarios, such as: 1. General robot operation: When single-arm / dual-arm robots perform tasks such as grasping, placing, pressing, and assembling, this invention can provide spatial positioning capabilities without deploying external detectors, reducing the complexity of system integration.

[0085] 2. Dual-arm collaborative operation: When performing tasks requiring coordination of both hands on a dual-arm robot platform, bounding boxes can be provided separately for the interaction areas of the left and right hands during training, and neither arm relies on external positioning input during inference.

[0086] 3. Service robots: This invention can be applied to service robots when performing object grasping and placement tasks using open language commands.

[0087] 4. Industrial Automation: Applying this invention in industrial picking and assembly scenarios can reduce reliance on external detectors.

[0088] 5. Remote teleoperation / shared autonomy: As an automated sub-task module for remote operation, it completes local tasks without increasing the inference burden on the end side.

[0089] Through the above description of the embodiments, those skilled in the art can clearly understand that the above embodiments can be implemented by software, or by using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, mobile hard drive, etc.), including several instructions to cause a computer device (such as a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0090] Example 2 This invention also provides a robot operating system based on a VLA model, which is mainly used to implement the methods provided in the foregoing embodiments, such as... Figure 6 As shown, the system mainly includes: The training data construction unit is used to divide the demonstration video into video segments corresponding to each sub-task. For each video segment, the task interaction region bounding box of each original RGB image is extracted and then superimposed on the corresponding original RGB image to obtain the superimposed image. The original prompts are added with prompts containing the focus bounding box region to obtain focus prompts. The superimposed image and focus prompts are used to construct the first type of training data, and the original RGB image and original prompts are used to construct the second type of training data. The model training unit is used to train the VLA model using two types of training data. The VLA model is a visual language action model, which includes a visual language model and an action head. The training process is as follows: Stage 1: Input the first type of training data and train the visual language model based on the two-dimensional pixel trajectory words output by the visual language model; Stage 2: Input the first type of training data and jointly train the visual language model and the action head based on the two-dimensional pixel trajectory words output by the visual language model and the output actions of the action head; Stage 3: Input the two types of training data, and gradually reduce the proportion of the first type of training data while gradually increasing the proportion of the second type of training data according to the linear curve of the training steps, and train in the same way as in Stage 2. The action prediction unit is used to input the original RGB image and the original prompt words into the VLA model after training to obtain the output action.

[0091] Since the main technical details of this system have been described in detail in previous embodiments, they will not be repeated here.

[0092] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above.

[0093] Example 3 The present invention also provides a processing device, such as Figure 7 As shown, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the foregoing embodiments.

[0094] Furthermore, the processing device also includes at least one input device and at least one output device; in the processing device, the processor, memory, input device, and output device are connected via a bus.

[0095] In this embodiment of the invention, the specific types of the memory, input device, and output device are not limited; for example: Input devices can be touchscreens, image acquisition devices, physical buttons, or mice, etc. The output device can be a display terminal; The memory can be random access memory (RAM) or non-volatile memory, such as disk storage.

[0096] Example 4 The present invention also provides a readable storage medium storing a computer program that, when executed by a processor, implements the method provided in the foregoing embodiments.

[0097] In this embodiment of the invention, the readable storage medium is a computer-readable storage medium and can be disposed in the aforementioned processing device, for example, as a memory in the processing device. Furthermore, the readable storage medium can also be any medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.

[0098] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims. The information disclosed in the background section is intended only to enhance the understanding of the overall background technology of the present invention and should not be construed as an admission or implication in any way that such information constitutes prior art known to those skilled in the art.

Claims

1. A robot manipulation method based on a VLA model, characterized in that, include: The demonstration video is divided into video segments corresponding to each subtask. For each video segment, the task interaction region bounding box of each original RGB image is extracted and then superimposed on the corresponding original RGB image to obtain the superimposed image. Add a focus tooltip containing the focus bounding box region to the original tooltip; The first type of training data is constructed using overlaid images and focus cues, and the second type of training data is constructed using original RGB images and original cues. The VLA model, a visual language action model, is trained using two types of training data. The VLA model comprises a visual language model and an action head. The training process is as follows: Stage 1: Input the first type of training data and train the visual language model based on the two-dimensional pixel trajectory words output by the visual language model. Stage 2: Input the first type of training data and jointly train the visual language model and the action head based on the two-dimensional pixel trajectory words output by the visual language model and the output actions of the action head. Stage 3: Input both types of training data, gradually decrease the proportion of the first type of training data while gradually increasing the proportion of the second type of training data along a linear curve, and train using the same method as in Stage 2. After training, the original RGB image and the original prompt words are input into the VLA model to obtain the output action.

2. The robot operation method based on a VLA model according to claim 1, characterized in that, The process of dividing the demonstration video into video segments corresponding to each sub-task, and extracting the task interaction region bounding box of each original RGB image for each video segment, includes: Based on the gripper state transition or with the help of an external visual language model, the demonstration video is divided into several video segments, each video segment corresponding to a sub-task; For each video segment, a semantic meaning is assigned to the corresponding subtask. Based on the subtask semantics, the interactive target object corresponding to the subtask is extracted. Using the interactive target object as the query, an open-word pointing model is used to output a representative single point on the keyframe of the subtask. The representative single point is the two-dimensional coordinate position of the interactive target object in the keyframe. A point-hint segmentation model is used to expand the representative single point into a segmentation mask and extract the bounding rectangle to obtain the task interaction region bounding box of the corresponding keyframe. The task interaction region bounding boxes of the remaining intermediate frames are obtained by linear interpolation of adjacent keyframes. Here, each frame is the original RGB image in the video segment.

3. The robot operation method based on a VLA model according to claim 1, characterized in that, The superimposition onto the corresponding original RGB image includes: superimposing the task interaction area bounding box with set color parameters, line type parameters and line width parameters onto the corresponding original RGB image.

4. The robot operation method based on a VLA model according to claim 1, characterized in that, The step of training the visual language model based on the two-dimensional pixel trajectory words output by the visual language model includes: Input the first type of training data, and the visual language model outputs the two-dimensional pixel trajectory words of the gripper from the current position to the path point of the next subtask. The current position is the two-dimensional pixel position of the gripper predicted by the visual language model from the first frame of the corresponding video segment. By combining the two-dimensional pixel trajectory words output by the visual language model with the pre-acquired ground truth values ​​of subtask path points, the word-by-word cross-entropy loss is calculated to provide autoregressive supervision for the visual language model.

5. The robot operation method based on a VLA model according to claim 1, characterized in that, The joint training of the visual language model and the action head based on the two-dimensional pixel trajectory words output by the visual language model and the output action of the action head includes: Input the first type of training data, and the visual language model outputs the two-dimensional pixel trajectory word of the gripper from the current position to the path point of the next subtask. The current position is the two-dimensional pixel position of the gripper predicted by the visual language model from the first frame of the corresponding video segment. The action head iteratively denoises and generates the output action based on the hidden state of the specified word in the last layer of the visual language model. By combining the two-dimensional pixel trajectory words output by the visual language model with the pre-acquired ground truth values ​​of subtask path points, word-by-word cross-entropy loss and flow matching loss of the output action are calculated. The visual language model and action head are jointly trained by combining the two losses.

6. A robot manipulation method based on a VLA model according to claim 4 or 5, characterized in that, The truth value of the subtask path point is obtained in the following way: For each video segment, the position of the gripper in the original RGB image is located by gripper feature matching, and then its two-dimensional pixel coordinates are obtained by segmentation, which are used as path points of the subtask. The pixel sequence formed by connecting adjacent path points constitutes the ground truth value of the subtask path point.

7. The robot operation method based on a VLA model according to claim 1, characterized in that, The training steps involve gradually reducing the proportion of the first type of training data along a linear curve, while simultaneously gradually increasing the proportion of the second type of training data, including: Phase 3 is defined as starting from the training step Begin, training steps End. Input the first type of training data with probability p(t) and the second type of training data with probability 1-p(t), then: Current training step At that time, p(t) = 1.0; The current training step t satisfies hour, .

8. A robot operating system based on the VLA model, characterized in that, To implement the method according to any one of claims 1 to 7, comprising: The training data construction unit is used to divide the demonstration video into video segments corresponding to each sub-task. For each video segment, the task interaction region bounding box of each original RGB image is extracted and then superimposed on the corresponding original RGB image to obtain the superimposed image. The original prompts are added with prompts containing the focus bounding box region to obtain focus prompts. The superimposed image and focus prompts are used to construct the first type of training data, and the original RGB image and original prompts are used to construct the second type of training data. The model training unit is used to train the VLA model using two types of training data. The VLA model is a visual language action model, which includes a visual language model and an action head. The training process is as follows: Stage 1: Input the first type of training data and train the visual language model based on the two-dimensional pixel trajectory words output by the visual language model; Stage 2: Input the first type of training data and jointly train the visual language model and the action head based on the two-dimensional pixel trajectory words output by the visual language model and the output actions of the action head; Stage 3: Input the two types of training data, and gradually reduce the proportion of the first type of training data while gradually increasing the proportion of the second type of training data according to the linear curve of the training steps, and train in the same way as in Stage 2. The action prediction unit is used to input the original RGB image and the original prompt words into the VLA model after training to obtain the output action.

9. A processing device, characterized in that, include: One or more processors; Memory, used to store one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method as described in any one of claims 1 to 7.

10. A readable storage medium storing a computer program, characterized in that, When a computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-granularity visual reasoning model construction method and device based on reinforcement learning

    CN121121289A

  • Target detection training method and system based on visual language model semantic scoring

    CN121415185A