Robot action generation method oriented to fine tasks

By employing a hierarchical intent alignment approach and utilizing cross-attention and availability prediction models, the referential and execution uncertainties of VLA models in complex scenarios are addressed, achieving high precision and reliability in robot operations while reducing computational resources and training costs.

CN121670633APending Publication Date: 2026-03-17NORTHEASTERN UNIV CHINA +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511794212.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-02
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing VLA models struggle to achieve deterministic recognition and precise operation of robot actions in complex scenarios, especially in multi-object scenarios where there are issues of referential and execution uncertainty. Current technologies have failed to effectively address the deterministic requirement for intent alignment.

Method used

By introducing multimodal data processing, a hierarchical alignment mechanism for the semantic and physical layers is designed. By utilizing the cross-attention mechanism and the availability prediction model, deterministic binding of instructions and targets and real-time mapping of local geometric features are achieved, generating smooth robot joint motion sequences.

Benefits of technology

It significantly improves the success rate and accuracy of robot operations, reduces computing resources and training costs, meets the requirements of high-precision tasks, and enhances the accuracy and reliability of robot operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121670633A_ABST
    Figure CN121670633A_ABST
Patent Text Reader

Abstract

The robot action generation method for the fine task comprises the steps that a VLA model receives multi-modal data and extracts global visual features and text feature vectors; performing semantic layer alignment and instruction remodeling to obtain an aligned instruction representation; based on instruction representation output by semantic alignment, analyzing action semantics and generating an accurate six-degree-of-freedom guide pose; performing hierarchical fusion on the global visual features, the instruction characterization and the 6-degree-of-freedom guide poses to generate a smooth future joint action sequence of the robot; and the robot controller drives the actuator to output the action sequence according to the future joint action sequence of the robot. According to the method, by establishing an end-to-end process from multi-modal input to action execution, deterministic alignment of human intention and robot operation is achieved. The core of the method is that a complex operation task is decomposed into a step sequence with clear logic, all steps are linked with one another, and the accuracy of instruction understanding and the accuracy of action generation are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of robot control technology, and relates to a method for generating robot motions for fine tasks. Background Technology

[0002] With the development of artificial intelligence technology, VLA models have achieved the mapping from natural language instructions to robot actions through end-to-end learning. However, their probabilistic reasoning paradigm still has significant limitations in complex scenarios. Existing technologies mainly improve the performance of VLA models through hierarchical architecture optimization and world representation enhancement, but they fail to fundamentally solve the deterministic requirement for intent alignment.

[0003] Layered VLA design based on diffusion models has become a mainstream approach. Its core principle is to decouple high-level semantic planning from low-level action generation to balance inference efficiency and control precision. A representative work is the RDT-1B model (RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation[C] / / The Thirteenth International Conference on Learning Representations. 2025), which constructs a unified diffusion foundation framework for bimanual tasks. RDT-1B expands the parameter scale of the diffusion model to 1.2B through a Transformer architecture and designs a physically interpretable action space, achieving cross-task and cross-platform generalization capabilities. Its function is to decompose complex manipulative tasks into a high-level planning module and a low-level diffusion controller: the high-level module parses language instructions and generates abstract sub-goals, while the low-level module generates smooth, high-dimensional continuous action sequences based on iterative denoising of the diffusion process. In implementation, the model is trained using large-scale demonstration data, searches for high-probability action regions through Langevin dynamics, supports multimodal output, and significantly improves the robustness and adaptability of bimanual coordination.

[0004] In enhancing world representation, these methods aim to compensate for the limitations of VLA models in physical understanding. Their function is to improve the model's perception accuracy of the operational scene by introducing 3D geometric information and visual region references. The process involves acquiring scene point cloud data using a depth camera, extracting local geometric features of objects through a point cloud segmentation network, and simultaneously employing pixel-level masking or bounding box techniques to strongly bind linguistic commands to specific regions in the image to clarify the operational target. For example, when handling fine-grained tasks, the model first converts the point cloud data into multi-view depth maps or directly extracts 3D features, then fuses them with a visual encoder to enhance the understanding of spatial relationships and object properties. These representation enhancement techniques are typically integrated into the VLA framework as preprocessing steps to improve the model's grounding capabilities.

[0005] The aforementioned solutions have significant shortcomings in applications requiring precise robotic manipulation, stemming from inherent limitations in their technical design. For the hierarchical architecture of the RDT-1B model, its high-level planning module relies entirely on implicit attention mechanisms for intent reasoning. When instructions are ambiguous or the scene contains multiple similar objects, it cannot guarantee deterministic identification of the target, easily leading to sub-target errors due to probabilistic guessing. While common methods for enhancing world representation, such as point cloud segmentation and masking, provide 3D geometric information, these representations are only used as preprocessing inputs and are disconnected from the action generation module. This results in physical features not being utilized in real-time during critical decision-making stages, making it difficult to meet the precision requirements of geometric matching in precise manipulation. Summary of the Invention

[0006] To address the aforementioned technical problems, the purpose of this invention is to provide a method for generating robot motions for fine-grained tasks.

[0007] The present invention provides a method for generating robot motions for fine-grained tasks, comprising:

[0008] Step 1: The VLA model receives multimodal data and extracts global visual features and text feature vectors;

[0009] Step 2: Perform semantic layer alignment and instruction reshaping to obtain aligned instruction representations;

[0010] Step 3: Based on the semantically aligned instruction representation, parse the action semantics and generate a precise 6-DOF guided pose;

[0011] Step 4: The global visual features, instruction representations, and 6-DOF guided poses are fused in layers to jointly guide the diffusion strategy to generate a smooth sequence of future joint movements of the robot;

[0012] Step 5: The robot controller drives the actuators to output a sequence of motions based on the robot's future joint motion sequence.

[0013] The robot motion generation method for fine-grained tasks of the present invention has the following beneficial effects:

[0014] Compared with existing technologies, this invention achieves significant technological improvements in robot precision manipulation tasks through an innovative hierarchical intent alignment architecture. This method effectively solves the core problems of referential and execution uncertainties through a dual alignment mechanism at the semantic and physical layers, greatly improving the robot's success rate and significantly reducing the misidentification rate in complex scenarios. Experimental verification shows that the pose accuracy generated by this invention reaches industrial-grade standards, and the operation trajectory is smooth and reliable, meeting the stringent requirements of high-precision assembly and medical assistance tasks.

[0015] This invention significantly reduces computing resources and training costs while maintaining high performance, making technology deployment simpler and more economical. Overall, this invention not only improves the accuracy and reliability of robot operations but also provides an efficient and economical solution for practical applications, possessing significant practical value. Attached Figure Description

[0016] Figure 1 This is a flowchart of a robot motion generation method for fine tasks according to the present invention. Detailed Implementation

[0017] This invention presents a robot motion generation method for fine-grained tasks, addressing the uncertainty issues of existing VLA models in fine-grained manipulation tasks. First, by introducing a user-provided first-frame mask as a visual constraint, referential uncertainty in multi-object scenes is completely eliminated. Second, a real-time mapping relationship between action semantics and local geometric features is established, enabling the robot to understand and execute complex operations requiring precise geometric alignment, such as insertion, removal, and tightening. Finally, through a hierarchically aligned interpretable framework, the system's adaptability and reliability in dynamic environments are improved, providing a reliable algorithmic solution for high-precision tasks (such as industrial assembly simulation) and demonstrating the model's generalization ability.

[0018] like Figure 1 As shown, the robot motion generation method for fine tasks according to the present invention specifically includes the following steps:

[0019] Step 1: Multimodal Data Input and Feature Encoding. The VLA model receives multimodal data and extracts global visual features and text feature vectors, specifically:

[0020] Step 1.1: The VLA model receives real-time RGB images, natural language instructions, and the user-provided first-frame target mask M∈{0,1}. H×W and scene point cloud data P∈R acquired by depth camera N×3 .

[0021] Step 1.2: Input the real-time RGB image into the dual-stream visual encoder to extract the global visual features F. v The DINOv2 branch of the dual-stream visual encoder is responsible for capturing fine geometric structures and object texture information in the scene, while the SigLIP branch focuses on extracting high-level semantic features that are highly aligned with the language semantics.

[0022] Step 1.3: Natural language instructions are first processed through word segmentation, and then input into a pre-trained text encoder to transform discrete word sequences into continuous text feature vectors F rich in deep semantic information. L .

[0023] Step 2: Perform semantic layer alignment and instruction reshaping to obtain aligned instruction representations, specifically:

[0024] Step 2.1: In order to provide high-quality, unambiguous visual evidence for subsequent intent alignment, a decoupled mask encoder containing semantic and spatial pathways is designed. This decoupled mask encoder processes the target’s intrinsic visual attributes and extrinsic spatial state through two independent pathways.

[0025] Step 2.2: Semantic Path: Element-wise filtering is performed on the first frame RGB image using the target mask M provided by the first frame to obtain the foreground image; the foreground image is then fed into a lightweight CNN encoder, where its multi-scale features are extracted and fused, followed by global average pooling to finally generate a highly condensed semantic token. , used to describe the appearance and inherent texture properties of an object.

[0026] Step 2.3: Spatial Path: After resetting the target mask M of the first frame to a resolution of 224×224 and flattening it, it is encoded into a compact spatial token through a linear projection layer. It is used to accurately characterize the position and geometric contour of a target in an image.

[0027] Step 2.4: The semantic token and spatial token are concatenated and fused through a linear layer to form a unified mask feature vector. Copy the vector Next, match the text feature vector F. L Sequence length. This decoupled design ensures that subsequent attention mechanisms can simultaneously utilize the semantic features and positional information of the target for precise intent alignment.

[0028] Step 2.5: Design an instruction reshaping module consisting of N identical decoder layers stacked together to achieve intent disambiguation. Each decoder layer contains two sub-modules: a multi-head cross-attention mechanism and a feedforward network. The sub-modules are connected through residual connections and layer normalization connections. Multiple decoders are connected serially, with the output of the previous layer serving as the input of the next layer. For the nth decoder layer, its input text feature vector FL As a query, the globally shared mask feature vector As Key and Value; through a multi-head cross-attention mechanism, each lexical element in the instruction is forced to actively query and absorb visual evidence. For the i-th head, the calculation is as follows:

[0029]

[0030] The attention function is calculated using the standard scaled dot product. , , It is a learnable projection matrix specific to each head. d model The main feature dimension of the VLA model is represented by the concatenation of the outputs of all h heads along this feature dimension, which is then fused by a single output linear transformation matrix.

[0031]

[0032] in, The output projection matrix is ​​denoted as x, which is the intermediate feature obtained after residual connection and layer normalization of the output of the above attention mechanism. This intermediate feature is then fed into the positional feedforward network FFN.

[0033]

[0034] in, , This is the weight matrix of the fully connected layer. , For the corresponding bias term, through the aforementioned N-layer stacking process, the instruction reshaping module reshapes an open description into an internal instruction representation rich in precise visual and spatial information, pointing to a unique physical entity. This provides a solid and unambiguous input foundation for all subsequent decisions. In specific implementation, this module contains N=2 decoder layers, with each cross-attention using h=8 heads. The final output is an aligned instruction representation. This reshapes an open description into an internal representation that is explicitly bound to specific entities in the scene.

[0035] Step 3: Physical Layer Alignment and Pose Prediction. Based on the instruction representation output by semantic alignment, the action semantics are parsed and a precise 6-DOF guided pose is generated, specifically as follows:

[0036] Step 3.1: First, start with the aligned instruction representation. The middle part extracts action primitives representing core actions through a small network and embeds them. To encode user intent.

[0037] Step 3.2: Use the target mask M of the first frame to segment the point cloud subset of the target object from the scene point cloud data P, and then extract the local geometric features through the PointNet++ network.

[0038] Step 3.3: Embedding with Action Primitives For the query, local geometric features are used as the key and value. The attention weight of each point is calculated through the attention mechanism. After the attention weight is normalized by Softmax, it is mapped to the point cloud coordinate space to generate an availability heatmap. The high-scoring area represents the best operation point.

[0039] Step 3.4: The attention weights from Step 3.3 are used to perform weighted aggregation on the original local geometric features, thereby suppressing interference from irrelevant regions and highlighting the geometric details of high-response regions, generating enhanced geometric features focused on the optimal operation point. These enhanced geometric features are concatenated with action primitive embeddings and input into a pose regression network to predict the 6-DOF guided pose T required to perform the action.

[0040] Step 4: Multimodal Conditional Fusion and Action Sequence Generation. Global visual features, command representations, and 6-DOF guided poses are fused hierarchically to jointly guide a diffusion strategy and generate a smooth sequence of future robot joint movements. Specifically:

[0041] Step 4.1: High-level semantic feature fusion. First, global visual features are concatenated with aligned instruction representations, and a cognitive token is appended to the end of the sequence to construct a unified multimodal input sequence.

[0042] Step 4.2: Feed the multimodal input sequence into the LLaMA-2 large language model, utilize its self-attention mechanism for cross-modal interaction, deeply fuse the visual context of the scene with the user's operational intent, and extract the output vector corresponding to the cognitive token as the high-level semantic conditional feature C. out .

[0043] Step 4.3: Decoding the pose-guided diffusion action. The 6-DOF guided pose T predicted in Step 3 is used as a strong geometric constraint, along with the high-level semantic conditional feature C. out Together, they serve as the denoising network in the conditional input diffusion model; during the back diffusion process, the diffusion model starts with Gaussian noise and iteratively denoises under the dual guidance of semantic intent and geometric pose; finally, the decoder outputs a smooth sequence of future robot joint movements that conforms to kinematic constraints, which is used to drive the robot to execute.

[0044] Step 5: The robot controller drives the actuators to output a sequence of movements based on the robot's future joint motion sequence, specifically:

[0045] Step 5.1: The robot's future joint motion sequence is sent to the robot controller, which drives the actuator to output the motion sequence.

[0046] Step 5.2: The system monitors the execution status through a real-time visual sensor. If a deviation is detected, a replanning mechanism is triggered: return to step 1 to update the input data, and re-execute steps 2 to 4.

[0047] This invention achieves a deterministic breakthrough in precise robot operation through a hierarchical intent alignment framework. It innovatively decomposes complex robot operation tasks into two logical levels: semantic layer alignment and physical layer alignment, forming a deterministic decision-making process of "identification before execution". The semantic layer specifically solves the referential uncertainty problem of "which object to operate", while the physical layer specifically solves the execution uncertainty problem of "how to operate precisely".

[0048] At the semantic alignment level, this invention designs an innovative cross-attention mechanism to achieve deterministic binding between instructions and targets. Using textual instruction features as query vectors and visual features of masked regions as key-value pairs, multi-head attention computation forces each lexical unit in the language instruction to actively query visual evidence from the masked regions, thereby reshaping the abstract instruction description into an internal representation explicitly bound to a specific physical entity.

[0049] At the level of physical alignment, this invention proposes an availability prediction model based on an attention mechanism. By using an availability attention network to quantify the compatibility of local geometry with specific actions, it automatically generates an availability heatmap covering the surface of an object, enabling the model to understand the physical constraint relationship between the local geometric features of the object and the actions.

[0050] The above description is only a preferred embodiment of the present invention and is not intended to limit the ideas of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A fine task-oriented robot motion generation method, characterized by, The method comprises the following steps: Step 1: the VLA model receives multimodal data and extracts global visual features and text feature vectors; Step 2: semantic layer alignment and instruction remodeling are performed to obtain aligned instruction representations; Step 3: based on the instruction representations output by semantic alignment, action semantics are parsed and accurate 6-DOF guide poses are generated; Step 4: global visual features, instruction representations and 6-DOF guide poses are hierarchically fused to jointly guide a diffusion strategy to generate a smooth robot future joint action sequence; Step 5: a robot controller drives an effector to output the action sequence according to the robot future joint action sequence.

2. The fine task-oriented robot motion generation method according to claim 1, characterized by, The step 1 is specifically: Step 1.1: VLA model receives real-time RGB image, natural language instruction, user-provided first frame target mask M∈{0,1} H×W and scene point cloud data P∈R N×3 collected by depth camera Step 1.2: Input real-time RGB image into dual-stream visual encoder to extract global visual features F v The DINOv2 branch of the dual-stream visual encoder is responsible for capturing fine geometric structure and object texture information in the scene, while the SigLIP branch focuses on extracting high-level semantic features that are highly aligned with linguistic semantics; Step 1.3: The natural language instruction is first processed by word segmentation, and then input into a pre-trained text encoder to convert the discrete word sequence into a continuous text feature vector F rich in deep semantic information L .

3. The fine task-oriented robot motion generation method according to claim 2, characterized by, The step 2 is specifically: Step 2.1: a decoupled mask encoder containing semantic channels and spatial channels is designed, which processes the internal visual attributes and external spatial states of the target through two independent channels respectively; Step 2.2: semantic channel: the first frame target mask M provided by the first frame is used to filter the first frame RGB image at the element level to obtain a foreground image; the foreground image is sent to a lightweight CNN encoder to extract and fuse multi-scale features thereof, After global average pooling, a highly condensed semantic token is finally generated to describe the appearance and texture inherent properties of the object; Step 2.3: Spatial pass: Reset the first frame target mask M to a resolution of 224x224 and flatten it, encoding it into a compact spatial token through a linear projection layer for accurate characterization of the position and geometric contour of the target in the image; Step 2.4: Semantic tokens are concatenated with spatial tokens and fused through a linear layer to form a unified mask feature vector The vector is copied Next to match the text feature vector F L Sequence length; Step 2.5: an instruction remodeling module stacked by N identical decoder layers is designed to realize intention disambiguation, each decoder layer contains a multi-head cross-attention mechanism and a feedforward network two sub-modules, and the sub-modules are connected through residual connection and layer normalization connection; The multiple decoders are connected in series, the output of the previous layer is taken as the input of the next layer, and for the nth decoder layer, the input text feature vector F L As Query, the mask feature vector shared globally As Key and Value; through multi-head cross-attention mechanism, force each token in the instruction to actively query and absorb visual evidence, for the ith head, calculate as follows: where the attention function is a standard scaled dot-product computation, , , is a per-head learnable projection matrix, , d model representing the backbone feature dimension of the VLA model, the outputs of all h heads are concatenated in the feature dimension and fused by an output linear transformation matrix: wherein, is an output projection matrix, and let x be the intermediate feature after the output of the attention mechanism is passed through a residual connection and layer normalization. The intermediate feature x is then fed into a positional feed-forward network (FFN): wherein, , is a full connection layer weight matrix, , is a corresponding bias term, through the above-mentioned N-layer stacking processing, the finally outputted aligned instruction representation .

4. The fine task-oriented robot motion generation method according to claim 3, characterized by, The step 3 is specifically: Step 3.1: First characterizing the aligned instruction table action primitives representing core actions are extracted from a small network to encode user intent; Step 3.2: the target object point cloud subset is segmented from the scene point cloud data P by using the first frame target mask M, and then local geometric features are extracted through a PointNet++ network; Step 3.3: Embedding with action primitives For Query, take local geometric features as Key and Value, calculate the attention weight of each point through attention mechanism, and map the normalized attention weight to the point cloud coordinate space to generate the availability heat map. The high-score area represents the best operation point. Step 3.4: the original local geometric features are weighted and aggregated by using the attention weights of step 3.3, so as to suppress the interference of irrelevant regions and highlight the geometric details of high response regions, and an enhanced geometric feature focused on the best operation point is generated; the enhanced geometric feature is spliced with the action primitive embedding and input into a pose regression network to predict a 6-DOF guide pose T required for executing the action.

5. The fine task-oriented robot motion generation method according to claim 4, characterized by, The step 4 is specifically: Step 4.1: the global visual features and the aligned instruction representations are spliced, a cognitive token is appended at the end of the sequence, and a unified multimodal input sequence is constructed; Step 4.2: The multi-modal input sequence is fed into the LLaMA-2 large language model, which uses its self-attention mechanism for cross-modal interaction, deeply fuses the visual context of the scene and the user's operation intention, and extracts the output vector corresponding to the cognitive token as the high-level semantic conditional feature C out ; Step 4.3: Use the predicted 6-DOF guidance pose T from Step 3 as a strong geometric constraint with the high-level semantic condition feature C out together as a conditional input to diffuse the denoising network in the model; In the reverse diffusion process, the diffusion model starts from Gaussian noise and iteratively denoises under the dual guidance of semantic intent and geometric pose; Finally, the decoder outputs a smooth robot future joint action sequence conforming to kinematic constraints, which is used to drive the robot to execute.

6. The fine task-oriented robot motion generation method of claim 1, wherein, The step 5 is specifically: Step 5.1: the robot future joint action sequence is sent to a robot controller to drive an effector to output the action sequence; Step 5.2: the system monitors the execution state through a real-time vision sensor, and if a deviation is detected, a re-planning mechanism is triggered: the input data is updated and steps 2 to 4 are re-executed.