Robot control method and device, electronic device, storage medium, program product
Patent Information
- Application Number
- CN202611311876.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-27
- Publication Date
- 2026-09-25
AI Technical Summary
视觉-语言-动作模型虽然能根据视觉观测与语言指令直接回归机器人本体动作,但缺少对动作执行后环境变化的显式建模,导致视觉-语言-动作模型无法预判动作后果,泛化性受限,在复杂操作中容易生成物理上不可行或导致任务失败的无效动作
[0020]在本申请的实施例所提供的技术方案中,一方面,根据机器人的多源观测数据构建基于统一三维坐标系的点云数据,并将观测图像关联的三维位置信息加入到视觉编码特征中,实现了二维语义特征与三维几何特征在同一空间坐标系下的精确对齐,为动作预测提供了高精度的空间上下文基础;另一方面,所得到的多模态融合特征同时蕴含了语言指令语义、二维视觉纹理、三维几何结构以及三维空间位置,能够显著提升对复杂操作场景的理解能力;再一方面,通过对基于统一三维坐标系的历史轨迹进行编码处理,所得到的轨迹编码特征能够为动作预测提供可靠的时序动态上下文。由此,本申请的实施例基于多模态融合特征和轨迹编码特征生成动作点流,该动作点流能够直接反映物理世界中的动作意图,并据此对机器人进行精确控制,有效避免了在复杂操作场景中生成物理上不可行或导致任务失败的无效动作的问题,从而提升了机器人从感知到动作的端到端控制可靠性。
Smart Images

Figure CN122807959A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robot control technology, specifically to a robot control method and device, electronic equipment, computer-readable storage medium, and computer program product. Background Technology
[0002] Currently, natural language command control for embodied robots typically employs a Vision-Language-Action (VLA) model architecture, directly outputting robot execution commands by fusing visual observations and verbal instructions. While the VLA model can directly regress the robot's own actions based on visual observations and verbal commands, it lacks explicit modeling of environmental changes after action execution. This results in the VLA model's inability to predict the consequences of actions, limiting its generalization capabilities and making it prone to generating physically infeasible or invalid actions that lead to task failure in complex operations. Therefore, improving the reliability of natural language command control for embodied robots is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0003] To address the aforementioned technical problems, embodiments of this application provide a robot control method, a robot control device, an electronic device, a computer-readable storage medium, and a computer program product. Based on a vision-language-action model architecture, embodiments of this application predict a four-dimensional motion point flow based on a unified three-dimensional coordinate system and perform robot control based on the predicted four-dimensional motion point flow, significantly improving the reliability of natural language command control for embodied robots.
[0004] One aspect of this application provides a robot control method, the method comprising: constructing point cloud data based on a unified three-dimensional coordinate system based on multi-source observation data of a robot, the point cloud data including three-dimensional position information associated with observation images of the robot; jointly encoding natural language commands and the observation images to obtain semantically aligned visual encoding features and language encoding features, and adding the three-dimensional position information associated with the observation images to the visual encoding features; extracting geometric encoding features from the point cloud data, and performing feature fusion processing on the geometric encoding features, the language encoding features, and the visual encoding features after adding the three-dimensional position information to obtain multimodal fusion features; encoding historical trajectories based on the unified three-dimensional coordinate system to obtain trajectory encoding features; generating an action point flow based on the multimodal fusion features and the trajectory encoding features, and controlling the robot based on the action point flow.
[0005] In another exemplary embodiment, the method further includes: acquiring new observation data of the robot after responding to control, and generating an actual trajectory based on the unified three-dimensional coordinate system based on the new observation data; saving the actual trajectory as a new historical trajectory so as to use the new historical trajectory for the next round of control.
[0006] In another exemplary embodiment, the action point stream generated based on the multimodal fusion features and the trajectory encoding features includes multiple sets; the method further includes: generating a corresponding scene point stream based on each set of action point streams, wherein the scene point stream represents the scene consequence corresponding to the action point stream; the step of controlling the robot based on the action point streams includes: selecting a target action point stream from the multiple sets of action point streams according to the scene point streams corresponding to each set of action point streams, and controlling the robot based on the target action point stream.
[0007] In another exemplary embodiment, the step of encoding the historical trajectory based on the unified three-dimensional coordinate system to obtain trajectory encoding features includes: acquiring trajectory sampling information of each spatial anchor point at multiple consecutive sampling times, the trajectory sampling information including three-dimensional position information; encoding the trajectory sampling information of the same spatial anchor point at the multiple consecutive sampling times into a feature vector of fixed dimension to obtain the trajectory encoding features.
[0008] In another exemplary embodiment, encoding the trajectory sampling information of the same spatial anchor point at the consecutive sampling times into a fixed-dimensional feature vector includes: encoding the timestamps corresponding to the same spatial anchor point at the consecutive sampling times into a consecutive set of negative values.
[0009] In another exemplary embodiment, the step of performing feature fusion processing on the geometric coding features, the language coding features, and the visual coding features after incorporating the three-dimensional position information to obtain multimodal fusion features includes: calculating the cross-attention between the language coding features and the visual coding features after incorporating the three-dimensional position information, and generating a two-dimensional weight mask based on the cross-attention; mapping the two-dimensional weight mask to a three-dimensional weight mask based on the unified three-dimensional coordinate system, and performing spatial gating modulation on the visual coding features after incorporating the three-dimensional position information through the three-dimensional weight mask to obtain modulated visual coding features; and fusing the modulated visual coding features and the geometric coding features to obtain the multimodal fusion features.
[0010] In another exemplary embodiment, the method is executed through a trained inference model, which includes a visual-language encoding module, a point cloud encoding module, a fusion module, a trajectory encoding module, an action flow prediction module, and a scene flow prediction module; wherein the scene flow prediction module is pruned in real-time inference mode and retained in scene deduction mode.
[0011] In another exemplary embodiment, the training process of the inference model includes the following steps: obtaining a training sample set containing multiple training samples, and inputting the multiple training samples into the inference model to obtain the action point flow output by the action flow prediction module and the scene point flow output by the scene flow prediction module; calculating the training loss value based on the action point flow output by the action flow prediction module and the scene point flow output by the scene flow prediction module, and optimizing the parameters of the inference model based on the training loss value.
[0012] In another exemplary embodiment, the method further includes: selectively inputting the ground truth values of the action point flow contained in the training samples and the action point flow output by the action flow prediction module into the scene flow prediction module based on preset constraints, so as to obtain the scene point flow output by the scene flow prediction module.
[0013] In another exemplary embodiment, the training process of the inference model further includes: initializing corresponding ontology adaptation modules for multiple ontology executors, wherein the ontology adaptation modules are used to convert action point flows into control instructions adapted to the ontology executors; freezing the parameters of the pre-trained inference model and training each ontology adaptation module based on the inference model with frozen parameters; unfreezing the parameters of the pre-trained inference model and performing joint fine-tuning on the inference model with unfrozen parameters and each pre-trained ontology adaptation module.
[0014] In another exemplary embodiment, the training process of the inference model further includes: determining a unified tool center point coordinate system corresponding to the plurality of ontology executors; mapping the action point flow output by the parameter-frozen inference model to an action point flow based on the unified tool center point coordinate system, so as to perform control command conversion on the mapped action point flow.
[0015] In another exemplary embodiment, the X-axis of the unified tool center point coordinate system points from the connecting rod to the gripper, the Y-axis is parallel to the opening and closing motion direction of the gripper, and the Z-axis is determined based on the right-hand rule.
[0016] In another aspect of this application, a robot control device is provided, comprising: a point cloud construction module configured to construct point cloud data based on a unified three-dimensional coordinate system according to multi-source observation data of the robot, wherein the point cloud data includes three-dimensional position information associated with the robot's observation images; a semantic encoding module configured to jointly encode natural language commands and the observation images to obtain semantically aligned visual encoding features and language encoding features, and to add the three-dimensional position information associated with the observation images to the visual encoding features; a multimodal fusion module configured to extract geometric encoding features from the point cloud data, and to perform feature fusion processing on the geometric encoding features, the language encoding features, and the visual encoding features after adding the three-dimensional position information to obtain multimodal fusion features; a trajectory encoding module configured to encode historical trajectories based on the unified three-dimensional coordinate system to obtain trajectory encoding features; and a generation and control module configured to generate an action point flow based on the multimodal fusion features and the trajectory encoding features, and to control the robot based on the action point flow.
[0017] Another aspect of this application provides an electronic device, including: one or more processors; and a memory for storing one or more computer programs, which, when executed by the one or more processors, cause the electronic device to implement the robot control method as described above.
[0018] Another aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor of an electronic device, causes the electronic device to perform the robot control method as described above.
[0019] Another aspect of this application provides a computer program product, including a computer program that, when executed by a processor of an electronic device, implements the robot control method described above.
[0020] In the technical solution provided by the embodiments of this application, on the one hand, point cloud data based on a unified three-dimensional coordinate system is constructed based on the robot's multi-source observation data, and the three-dimensional position information associated with the observation images is added to the visual coding features. This achieves precise alignment of two-dimensional semantic features and three-dimensional geometric features in the same spatial coordinate system, providing a high-precision spatial context foundation for action prediction. On the other hand, the obtained multimodal fusion features simultaneously contain language instruction semantics, two-dimensional visual texture, three-dimensional geometric structure, and three-dimensional spatial position, which can significantly improve the understanding of complex operation scenarios. Furthermore, by encoding the historical trajectory based on the unified three-dimensional coordinate system, the obtained trajectory coding features can provide a reliable temporal dynamic context for action prediction. Thus, the embodiments of this application generate action point streams based on multimodal fusion features and trajectory coding features. These action point streams can directly reflect the action intentions in the physical world and are used to precisely control the robot. This effectively avoids the problem of generating physically infeasible or invalid actions that lead to task failure in complex operation scenarios, thereby improving the end-to-end control reliability of the robot from perception to action.
[0021] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description
[0022] Figure 1 This is a schematic diagram of an implementation environment involved in this application.
[0023] Figure 2 A schematic diagram of an exemplary inference model architecture is shown.
[0024] Figure 3 A flowchart of an exemplary robot control method is shown.
[0025] Figure 4 An exemplary real-time inference architecture diagram for robot control is shown.
[0026] Figure 5 Another exemplary scenario simulation architecture diagram for robot control is shown.
[0027] Figure 6 A schematic diagram of the training process corresponding to an exemplary inference model is shown.
[0028] Figure 7 It shows the relationship with Figure 6 The diagram shows the training architecture corresponding to the training process shown.
[0029] Figure 8 A schematic diagram of the training process corresponding to another exemplary inference model is shown.
[0030] Figure 9 It shows the relationship with Figure 8 The diagram shows the training architecture corresponding to the training process shown.
[0031] Figure 10 A block diagram of an exemplary robot control device is shown.
[0032] Figure 11 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown. Detailed Implementation
[0033] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0034] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0035] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0036] In this application, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0037] The terms "first," "second," "third," and "fourth," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. The terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0038] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0039] An embodied system is an intelligent agent system that possesses a physical entity and can operate in a closed loop of "seeing-thinking-moving" within a real three-dimensional environment. The entity can be a robotic arm, a dual-arm robot, a gripper, or a human hand with a MANO hand mold, etc. Unlike a simple large model that only answers questions in symbolic or pixel space, an embodied system transforms perceived physical observations into executable actions of the entity, which are then applied to the environment, and new observations are obtained and fed back.
[0040] Current embodied robot intelligence solutions typically focus on "how the robot should move," and to address this issue, a Vision-Language-Action (VLA) model architecture has been proposed.
[0041] Vision-language-action (VLA-MA) models are end-to-end policy models that use a multimodal large-scale model as their backbone, simultaneously receiving visual observations and natural language commands, and directly outputting robot action signals. Their core characteristic is the unification of visual encoding, language understanding, and action generation within the same network architecture, achieving an integrated mapping from perception to execution through large-scale pre-training and imitation learning. While current VLA-MA architectures can directly regress and generate robot actions based on visual observations and language commands, they lack explicit modeling of environmental changes after action execution. This results in the VLA-MA model's inability to predict the consequences of actions, limiting its generalization ability. This directly leads to the generation of physically infeasible or invalid actions that cause task failure in complex operations.
[0042] To improve the robustness and reliability of robots executing natural language commands in complex open environments, the inventors of this application discovered that current vision-language-action models, when parsing high-level semantic commands, often rely solely on static visual observations at the current moment to directly generate actions. This "only looking at the present, ignoring the consequences" reasoning paradigm has inherent flaws. The inventors thus realized that if dynamic information related to the expected changes in the scene after the robot performs an action could be explicitly introduced during the action prediction process, constructing an internal logic from prediction to execution, the robot's execution deviations caused by complex operations could be fundamentally resolved, thereby improving the reliability of natural language command control for embodied robots.
[0043] Therefore, embodiments of this application propose a robot control scheme. Based on the current vision-language-action model architecture, it predicts a four-dimensional motion point flow in a unified three-dimensional coordinate system by introducing point cloud encoding, multimodal feature fusion, trajectory encoding, and other methods. This is equivalent to explicitly introducing dynamic information related to the expected scene changes after the robot performs the action during the action prediction process, so that the predicted motion point flow can directly reflect the action intention in the physical world. Based on this four-dimensional motion point flow, precise control of the robot can be achieved, effectively avoiding the problem of generating physically infeasible or invalid actions that lead to task failure in complex operation scenarios. Therefore, it can significantly improve the reliability of end-to-end robot control.
[0044] The technical solutions provided by the embodiments of this application will be described in detail below:
[0045] Please refer to the following first. Figure 1 , Figure 1 This is a schematic diagram of an implementation environment involved in this application. Specifically, the implementation environment is a robot control system, which includes a terminal 110 and a robot 120, and the terminal 110 and the robot 120 can communicate wirelessly.
[0046] Terminal 110 is used to run the robot control client, where users can trigger relevant operations, such as inputting natural language commands, through the user interface. Terminal 110 sends the user-input natural language commands to the controller on the robot 120 via the network, serving as task description conditions for subsequent long-range control processes. Terminal 110 can be any electronic device, such as a smartphone, tablet, laptop, desktop computer, smart TV, smartwatch, vehicle terminal, or aircraft, and is not limited thereto.
[0047] Robot 120 is an autonomous operating platform with visual perception and joint execution capabilities, including but not limited to a mobile robotic arm, a humanoid robot, or a desktop dual-arm operating robot. Robot 120 is equipped with at least one visual sensor (such as a wrist camera, head camera, depth camera, multi-view camera, or panoramic camera) facing the operating area to acquire observational images. Robot 120 is also equipped with body actuators, which are drive devices integrated within the physical body of Robot 120 responsible for translating control commands into physical motion. Examples include end-effectors (EEFs) and joint drive units. The end-effector is mounted at the end of the robotic arm and is responsible for performing specific operational tasks, such as grasping and placing. The joint drive units are responsible for the movement of the robot arm and mobile chassis to achieve macroscopic displacement and posture adjustment. Robot 120 has a built-in controller, which may be an embedded computing unit equipped with a GPU (Graphics Processing Unit) or a TPU (Tensor Processing Unit), without limitation.
[0048] The controller built into Robot 120 contains a trained inference model. In real-time inference mode, this model predicts future action point flows based on the robot's multi-source observation data, natural language commands, and historical trajectories. The controller then sends specific control commands to the robot's actuators based on these predicted action point flows, controlling the actuators to perform corresponding operations. In scenario simulation mode, after predicting future action point flows based on the robot's multi-source observation data, natural language commands, and historical trajectories, the inference model further predicts scene point flows based on these action point flows. The controller then sends specific control commands to the robot's actuators based on both the predicted action point flows and scene point flows. For example, the controller selects actions from the action point flows based on the predicted scene point flows, and according to task cost, safety constraints, or target distance constraints, controls the actuators to execute the selected actions.
[0049] Figure 2 A schematic diagram of an exemplary inference model architecture is shown, such as... Figure 2 As shown, the inference model includes a visual-language encoding module, a point cloud encoding module, a fusion module, a trajectory encoding module, an action flow prediction module, and a scene flow prediction module.
[0050] The visual-language encoding module is used to jointly encode natural language commands and the current observation images of robot 120, outputting semantically aligned visual and language encoding features. The point cloud encoding module is used to encode the constructed point cloud data based on a unified three-dimensional coordinate system, outputting geometric encoding features. This point cloud data is constructed based on multi-source observation data of robot 120. The fusion module is used to perform feature fusion processing on the geometric encoding features output by the point cloud encoding module, the language encoding features output by the visual-language encoding module, and the visual encoding features after adding three-dimensional position information associated with the observation images, outputting multimodal fusion features. The trajectory encoding module is used to encode historical trajectories based on a unified three-dimensional coordinate system, outputting trajectory encoding features. The action flow prediction module is used to predict the action point flow corresponding to future time steps based on the multimodal fusion features output by the fusion module and the trajectory encoding features output by the trajectory encoding module. The scene flow prediction module is used to predict the scene flow data after robot 120 executes the predicted action based on the action point flow output by the action flow prediction module.
[0051] It should be noted that, Figure 2 The illustrated scene flow prediction module can be cut off in real-time inference mode and retained in scene deduction mode. It is understood that the real-time inference mode and scene deduction mode mentioned in the embodiments of this application refer to two different inference modes corresponding to the inference model. The real-time inference mode is suitable for high-frequency control scenarios with stringent response latency requirements. The model receives multi-source observation data, natural language commands, and historical trajectories. After multimodal fusion and spatiotemporal encoding, the action flow prediction module directly outputs future action point flow data. The real-time inference mode significantly reduces inference latency by reducing redundant physical deduction calculations, meeting the millisecond-level response requirements for tasks such as real-time obstacle avoidance and high-speed grasping. The scene deduction mode is suitable for complex operation scenarios with stringent safety and reliability requirements. The model first generates candidate action point flow data, and then uses this action point flow as a condition to drive the scene flow prediction module to predict the corresponding future scene changes. The controller combines the action point flow and scene point flow, performs multi-step look-ahead evaluation and candidate action screening based on conditions such as task cost, safety constraints, and target distance, and then issues the optimal action for execution. The scenario simulation mode realizes the model predictive control paradigm through internal physical simulation, which can effectively avoid potential collisions and execution errors in high-risk tasks.
[0052] Therefore, the controller built into Robot 120, based on the current vision-language-action model architecture, further introduces point cloud encoding, multimodal feature fusion, trajectory encoding and other methods to predict four-dimensional motion point flow in a unified three-dimensional coordinate system. This explicitly introduces dynamic information related to the expected scene changes after Robot 120 performs actions during the motion prediction process, so that the predicted motion point flow can directly reflect the action intention in the physical world. Based on this four-dimensional motion point flow, precise control of Robot 120 can be achieved, significantly improving the reliability of end-to-end robot control.
[0053] Please continue reading. Figure 3 , Figure 3 A flowchart illustrating an exemplary robot control method is shown. This method can be applied to... Figure 1 The implementation environment shown, and by Figure 1 The method is executed by robot 120 in the illustrated implementation environment, specifically by the controller built into robot 120. This method can also be applied to other implementation environments and executed by the controller built into the robot in those other environments to achieve motion control of the robot; the embodiments of this application do not limit this.
[0054] like Figure 3 As shown, in an exemplary embodiment, the robot control method includes steps S310-S350, which are described in detail below: S310, Based on the robot's multi-source observation data, construct point cloud data based on a unified three-dimensional coordinate system. The point cloud data contains three-dimensional position information associated with the robot's observation images.
[0055] First, it's important to clarify that multi-source observation data for a robot doesn't refer solely to images captured by visual sensors. Instead, it encompasses all sensor feedback data available to the robot at any given moment for decision-making. For example, multi-source observation data includes environmental observation data and body observation data. Environmental observation data might include images captured by one or more visual sensors, such as RGB images and RGB-D (Deep) images captured by one or more RGB cameras. Body observation data might include key points of the robot's body (e.g., EEF, gripper, hand), as well as the states of joints or TCP (Tool Center Point), without limitation.
[0056] Point cloud data refers to the conversion of environmental observation data into a set of discrete 3D points in a unified 3D coordinate system through depth backprojection and multi-source calibration. Each point carries at least 3D position coordinates in the unified 3D coordinate system and typically also carries attributes such as RGB, normal, depth, source camera identifier, and timestamp. It is used to discretize the geometry of object surfaces or scene structures. For example, depth pixels can be backprojected into camera coordinate points using camera intrinsic and extrinsic parameters, and then transformed to a unified 3D coordinate system using extrinsic parameters to obtain each observed pixel. In the case of multiple views, point clouds observed by multiple cameras can be fused and then voxel downsampled while retaining the camera source identifier. Key points such as robot EEF, grippers, and hands, as well as object surface points and scene points, also need to be transformed to a unified 3D coordinate system.
[0057] Point cloud data containing 3D positional information associated with the robot's observed images means that, in order to adapt to the vision-language-action model architecture, it is necessary to associate the 3D positional information corresponding to the image patches with visual features (tokens). Current vision-language-action models typically divide observed images into patches, and each patch is converted into a visual feature as input. These visual features originally only contain 2D positional information and do not have a 3D concept. The embodiments of this application need to associate these features with 3D positional coordinates. For example, the 3D positional coordinates corresponding to each image patch can be obtained by aggregating the 3D coordinates corresponding to all pixels within the image patch. Aggregation methods include, but are not limited to, calculating the depth mean, median, and taking the depth of the center pixel.
[0058] It should also be noted that the unified 3D coordinate system can be any of the world coordinate system, robot base coordinate system, camera coordinate system, object coordinate system, or TCP coordinate system, and the embodiments of this application do not limit this. When performing motion control, it is necessary to transform from the unified 3D coordinate system to the corresponding TCP coordinate system or robot base coordinate system.
[0059] S320 performs joint encoding on natural language instructions and observed images to obtain semantically aligned visual and linguistic encoding features, and adds the three-dimensional location information associated with the observed images to the visual encoding features.
[0060] Joint encoding of natural language commands and robot-observed images refers to the process of extracting joint visual-language features from the natural language commands and observed images using a visual-language encoding module. The output of the visual-language encoding module consists of semantically aligned visual and linguistic encoded features, but with separate sequence channels; these can also be called joint visual-language features. The visual-language encoding module can be the encoding part of a visual-language encoding model or an independent visual-language model (VLM); there is no restriction on this.
[0061] Visual coding features are essentially feature sequences composed of visual coding features of multiple image blocks arranged in a certain order. Since the visual-language coding module divides the observed image into several blocks and converts each block into visual features as input during joint encoding, and then performs encoding on each visual feature, the visual coding features output by the visual-language coding module only contain two-dimensional positional information for each image block, lacking a three-dimensional concept. The embodiments of this application superimpose the three-dimensional positional information associated with each image block in S310 onto the corresponding visual coding features in the feature sequence. Thus, the visual coding features with added three-dimensional positional information carry the three-dimensional geometric coordinates of the corresponding image block in a unified three-dimensional coordinate system, achieving precise alignment of two-dimensional semantic features and three-dimensional geometric features in the same spatial coordinate system, providing a high-precision spatial context foundation for subsequent action prediction.
[0062] S330 extracts the geometric coding features of point cloud data and performs feature fusion processing on the geometric coding features, language coding features, and visual coding features after adding 3D position information to obtain multimodal fusion features.
[0063] The embodiments of this application can use a point cloud encoding module to extract the native three-dimensional geometric features of point cloud data to obtain geometric encoding features. These geometric encoding features may include information such as local geometry, object shape, and scene structure. It should be noted that these geometric encoding features and the visual encoding features obtained in S320 after incorporating three-dimensional position information use the same coordinate system.
[0064] The embodiments of this application perform feature fusion processing on geometric coding features, language coding features, and visual coding features with added three-dimensional position information. This results in multimodal fusion features that simultaneously contain language instruction semantics, two-dimensional visual texture, three-dimensional geometric structure, and three-dimensional spatial position, which can significantly improve the ability to understand complex operation scenarios.
[0065] In an exemplary embodiment, geometric coding features, language coding features, and visual coding features incorporating 3D positional information can be preprocessed separately, and the preprocessed coding features can be concatenated to obtain multimodal fusion features. The preprocessing process can sequentially include linear projection, normalization, and dimension alignment to eliminate representational biases between heterogeneous modal features, map semantic and geometric information to the same latent space, thereby reducing the computational complexity of cross-modal feature interaction and improving the efficiency and stability of feature fusion.
[0066] In another exemplary embodiment, the feature fusion process may include the following steps (not shown in the accompanying drawings): S330-1 calculates the cross-attention between language coding features and visual coding features after incorporating 3D positional information, and generates a 2D weight mask based on the cross-attention. S330-2 maps a two-dimensional weighted mask to a three-dimensional weighted mask based on a unified three-dimensional coordinate system, and uses the three-dimensional weighted mask to perform spatial gating modulation on the visual coding features after adding three-dimensional position information to obtain the modulated visual coding features. S330-3 fuses the modulated visual coding features and geometric coding features to obtain multimodal fusion features.
[0067] It should be noted that the embodiments of this application first calculate the cross-attention between the language encoding features as the query vector (Query) and the visual encoding features after adding 3D position information as the key vector (Key), generating an attention map of dimension H×W×1. This attention map is then normalized to obtain a two-dimensional weight mask. Subsequently, combining the depth image and camera intrinsic and extrinsic parameters, the two-dimensional weight mask is mapped to a 3D space under a unified 3D coordinate system to obtain a 3D weight mask. Finally, based on the 3D weight mask, spatial gating modulation is performed on the visual encoding features after adding 3D position information, i.e., weighting is applied to the visual encoding features. This amplifies the feature responses of target-related regions and suppresses the feature responses of background-irrelevant regions, thereby guiding subsequent action generation to focus on the physical entity referred to by the language instruction.
[0068] Therefore, the embodiments of this application further propose an instruction condition space prior construction mechanism. This mechanism combines cross-modal attention guided by language instructions with three-dimensional geometric mapping, forcing a focus on the physical entity referred to by the natural language instructions at the feature level. This blocks the path of decision-making based on background texture, thereby eliminating the background shortcut phenomenon and effectively overcoming the background shortcut problem that is common in existing vision-language coding architectures in robot control tasks. Moreover, this also provides accurate geometric targeting for subsequent action point flow prediction. For example, it ensures that the generated gripper trajectory can accurately point to the graspable part of the target object, avoiding grasping failure or environmental collision caused by target positioning deviation, and significantly improving the success rate of action execution.
[0069] S340 encodes historical trajectories based on a unified three-dimensional coordinate system to obtain trajectory coding features.
[0070] The historical trajectory mentioned in the embodiments of this application refers to the sequence data of the three-dimensional spatial position and state (such as speed, posture, etc.) of the robot's body actuator at multiple consecutive historical sampling moments. For example, it is the trajectory data generated by the body actuator in response to the previous round of control, which is used to characterize the motion inertia and recent state of the body actuator.
[0071] Historical trajectories based on a unified 3D coordinate system refer to historical trajectories that also correspond to a unified 3D coordinate system. This ensures that the trajectory encoding features and multimodal fusion features obtained through encoding are both anchored to the same unified 3D coordinate system. Their vector spaces are continuous and aligned, thus eliminating the spatial semantic gap between different modal features. This allows for the subsequent generation of action point flows based on multimodal fusion features and trajectory encoding features. Moreover, trajectory encoding features have a temporal dimension, and multimodal fusion features have a spatial dimension. By unifying the 3D coordinate system, the temporal and spatial dimensions are combined, achieving a spatiotemporal joint representation based on a unified coordinate system. This enables the effective capture of the dynamic coupling relationship between actions and the environment when generating action point flows based on multimodal fusion features and trajectory encoding features.
[0072] In an exemplary embodiment, trajectory sampling information for each spatial anchor point included in the historical trajectory can be obtained at multiple consecutive sampling times. This trajectory sampling information includes three-dimensional position information. The trajectory sampling information of the same spatial anchor point at multiple consecutive sampling times is encoded into a fixed-dimensional feature vector to obtain trajectory encoding features. The spatial anchor point may include key points of the ontology actuator and scene points, and the sampling time refers to the actual timestamp carried by the spatial anchor point.
[0073] The embodiments of this application treat the trajectory sampling information of the same spatial anchor point at multiple consecutive sampling times as a single trajectory, encoding it into a fixed-dimensional feature vector. This can be understood as compressing an entire trajectory into a fixed-length D-dimensional vector. It should also be noted that the dimension of the multimodal fusion features should be consistent with the dimension of the trajectory encoding features, and the feature dimensions of the subsequently predicted action point stream and scene point stream should also be consistent with the dimension of the trajectory encoding features. Thus, all features share the same feature space and support cross-modal attention interaction.
[0074] In another exemplary embodiment, the trajectory sampling information of each spatial anchor point at multiple consecutive sampling times may include not only three-dimensional position information, but also dynamic information such as velocity, visibility, and category, without limitation.
[0075] In another exemplary embodiment, the timestamps corresponding to the same spatial anchor point at multiple consecutive sampling times are encoded as multiple consecutive negative values. This can be understood as the embodiment of this application using continuous time encoding for timestamps, with historical trajectories corresponding to negative timestamps and future query trajectories corresponding to positive timestamps. For example, normalized continuous time, sine / cosine encoding, time embedding, etc., can be used to convert real timestamps into vector values, and there are no limitations here. For ease of understanding, for example, if we assume that the negative timestamps are [-10, -9, ..., -1], representing "10 steps forward from the current moment to 1 step forward", the corresponding positive timestamps [1, 2, ..., 10] represent "1 step backward from the current moment to 10 steps backward". Thus, based on the encoding method of positive and negative timestamps, the subsequent action generation can infer the motion trend from the speed changes of the historical trajectory. For example, when the speed of the historical trajectory decreases near t=0, it can be predicted that the actuator will stop or turn; when the speed direction is stable at the end of the historical trajectory, it can be predicted that the trajectory will extend along the inertial direction.
[0076] Therefore, the embodiments of this application can encode historical trajectories based on a unified three-dimensional coordinate system through a trajectory encoding module. This can effectively capture the motion state and change trend of the robot's body actuator in the past time series, providing a reliable temporal dynamic context for predicting future actions and improving the coherence and rationality of action generation.
[0077] The S350 generates motion point streams based on multimodal fusion features and trajectory coding features, and controls the robot based on the motion point streams.
[0078] Embodiments of this application can generate an action point flow based on multimodal fusion features and trajectory encoding features through an action flow prediction module. This action flow prediction module is configured to receive multimodal fusion features and trajectory encoding features, and generate the action point flow based on a query token mechanism. The number and dimensions of the query tokens correspond to the physical structure of the robot's body actuators. For example, query tokens spatially correspond to the point cloud of the actuator surface, key skeleton points, or abstract geometric points of the end effector; no limitation is imposed here. Through a multi-layer Transformer decoder architecture, a cross-attention mechanism can be established between the query tokens, the multimodal fusion features, and the trajectory encoding features. In the cross-attention calculation, the query tokens, as query vectors, are used to actively retrieve feature information related to the body actuator's actions. The multimodal fusion features and trajectory encoding features, as key vectors and value vectors, provide the static geometric layout of the current scene, the semantics of language commands, and the historical motion inertia of the actuator itself, thereby obtaining a predefined sequence of future time steps. The future time step is coupled with the feature vector of the query token as a position encoding. This allows for the parallel prediction of the three-dimensional displacement or absolute position coordinates of each query token at each future time step based on the current and historical states, ultimately obtaining the action point flow.
[0079] The embodiments of this application generate motion point flows based on multimodal fusion features and trajectory encoding features. These flows precisely depict the refined motion trajectory of the robot's actuators over a continuous future time period, directly reflecting the intended action in the physical world; essentially, they are four-dimensional point flows. Furthermore, due to the correspondence between query tokens and the physical points of the actuators, this motion point flow naturally possesses the characteristic of decoupling from the specific robot body, providing a universal geometric intermediate representation for the cross-body adaptation disclosed in subsequent embodiments. Therefore, after generating the motion point flow, by converting it into control commands adapted to the robot's actuators, specific control can be performed on the actuators based on these control commands.
[0080] Compared to the current vision-language-action model architecture, which can only directly regress and generate robot body actions based on visual observations and language commands, the embodiments of this application further introduce point cloud geometric coding and historical trajectory coding. Point cloud geometric coding features provide accurate three-dimensional structural priors, which can solve the problems of inaccurate grasping pose and environmental collision caused by the lack of geometric constraints in the current vision-language-action model. Historical trajectory coding features introduce physical inertia and temporal continuity, overcoming the defects of action jumps and dynamic infeasibility caused by open-loop prediction in the current vision-language-action model. Furthermore, the embodiments of this application also add three-dimensional position information to the visual coding features, realizing the conversion of visual semantic features from two-dimensional semantics to three-dimensional semantics, which can accurately locate specific physical entities based on language commands, significantly improving the accuracy of command following. Based on these differences, it can be concluded that the embodiments of this application completely reconstruct the underlying logic of action generation. The action point flow generated by the embodiments of this application is a high-level spatial representation that is independent of the ontology. It only depicts the motion trajectory of the ontology executor in the physical world and can directly reflect the action intention in the physical world. In this way, the decoupling of "perception decision" and "execution control" is achieved, and it has the generalization across ontology executors. The current vision-language-action model architecture outputs low-level control instructions that are bound to the ontology executor. The two action generation logics are substantially different.
[0081] Therefore, the embodiments of this application control the robot based on action point flow, which effectively avoids the problem of generating physically infeasible or invalid actions that lead to task failure in complex operation scenarios, thereby effectively improving the end-to-end control reliability of the robot from perception to action.
[0082] In another exemplary embodiment, in Figure 3 Based on the illustrated embodiment, after performing control of the robot based on the motion point flow, the robot control method further includes the following steps (not shown in the accompanying drawings): Acquire new observation data of the robot after responding to control, and generate the actual trajectory based on the new observation data using a unified three-dimensional coordinate system; The actual trajectory is saved as a new historical trajectory so that the new historical trajectory can be used for the next round of control.
[0083] It should be noted that in the above process, the new observation data after the robot responds to control refers to the newly acquired multi-source observation data after the predicted motion point flow is converted into control commands adapted to the robot's actuators and the actuators are controlled, following the prediction of the motion point flow in the previous round. This can be understood as the process of obtaining physical execution feedback. Generating the actual trajectory based on the new observation data in a unified 3D coordinate system means back-projecting the depth information of the new observation image onto the unified 3D coordinate system, then parsing the position sequence of the key points of the actuators in this unified 3D coordinate system, and adding timestamp information according to the actual acquisition time, thereby obtaining the actual trajectory. Saving the actual trajectory as a new historical trajectory allows it to be used in the next round of control, that is, providing the historical trajectory as input for the next round.
[0084] Figure 4 An exemplary real-time inference architecture diagram for robot control is shown, such as... Figure 4 As shown, this exemplary real-time inference architecture is a closed loop, with each closed-loop inference corresponding to one round of robot control. Each closed loop sequentially includes the processes of input, action point flow generation, ontology actuator adaptation, robot execution, and the entry of a new observation into the next closed loop. Figure 4 The dashed line points from "new observations entering the next closed loop" to "input," indicating that the new observation data obtained by the robot after executing this round of control will be used as the historical trajectory input for the next round of control, thus forming a closed-loop architecture.
[0085] Therefore, the closed-loop control process disclosed in the embodiments of this application achieves dynamic updates of physical memory by writing the actual trajectory after each round of robot execution back into the trajectory encoding features of the next round. Compared with the "prediction-execution" paradigm of the current vision-language-action model architecture, this closed-loop control process effectively eliminates the cumulative effect of robot execution deviations, enabling action point flow prediction to respond to the dynamic changes of real physical interactions. Moreover, compared with the traditional closed-loop following the "pixel-level deviation correction" paradigm, this closed-loop control process achieves trajectory-level updates based on a unified three-dimensional coordinate system, which can avoid the spatiotemporal drift problem of long-sequence tasks and adapt to the general control requirements of different robot body actuators.
[0086] In yet another exemplary embodiment, in Figure 3 Based on the embodiment shown, S340 can generate multiple sets of action point flows as candidates based on multimodal fusion features and trajectory coding features. After generating multiple sets of candidate action point flows, the robot control method also generates corresponding scene point flows based on each set of action point flows. Thus, based on the scene point flows corresponding to each set of action point flows, the target action point flow is selected from multiple sets of candidate action point flows, and the robot is controlled based on the target action point flow.
[0087] It should be noted that scene point flow refers to the three-dimensional displacement trajectory of spatial anchor points on the surface of non-actuator objects within a scene at consecutive future moments, driven by the predicted action point flow. This scene point flow, conditioned on the action point flow, constructs an explicit mapping relationship between actions and scene changes to deduce the physical interaction consequences. Therefore, this scene point flow is used to characterize the scene consequences corresponding to the action point flow.
[0088] In this embodiment, a target action point flow is selected from multiple action point flows based on the scene point flow corresponding to each group of action point flows. For example, this can be achieved by evaluating one or more of the following for each group of action point flows: task cost, safety constraint compliance, and target distance deviation. A comprehensive score is calculated based on the evaluation results, and the action point flow with the highest comprehensive score is selected as the target action point flow. The evaluation of safety constraint compliance can include calculating the minimum Euclidean distance between the actuator point flow and the scene point flow, and calculating the displacement of scene objects. If the minimum Euclidean distance is less than a preset safety threshold or the displacement of the scene object exceeds a preset deformation threshold, it is determined that the safety constraint is not met. The target distance deviation can be obtained by calculating the Euclidean distance between the action endpoint and the preset target task point. The task cost can be determined by calculating parameters such as the path length of the action, the joint rotation angle, and the estimated energy consumption.
[0089] Figure 5 Another exemplary scenario inference architecture for robot control is shown. It can be seen that under the scenario inference architecture, multiple sets of action point flows are first predicted as candidates, and then scene point flows are predicted based on each set of action point flows as conditions. The target action point flows are then selected based on the predicted scene point flows to enter the subsequent body actuator adaptation process. Figure 5 The dashed line still points from "new observations entering the next closed loop" to "input," indicating that the new observation data obtained by the robot after executing this round of control will be used as the historical trajectory input for the next round of control, thus forming a closed-loop architecture. This can also be understood as the scenario inference architecture still being a closed loop, with each closed-loop inference corresponding to one round of robot control.
[0090] Therefore, based on the prediction of action point flows, the embodiments of this application further propose a scene point flow generation mechanism conditioned by action point flows, establishing an explicit physical causal chain of action-environment interaction. The controller predicts the scene consequences corresponding to candidate action point flows and uses scene point flows for "pre-rehearsal," selecting the best action point flow as the target action point flow from multiple candidate action point flows. This enables the controller to actively screen the best candidate actions, transforming passive safety protection into active consequence prediction, greatly improving the safety of human-machine collaboration and precision operation tasks.
[0091] It should also be noted that the robot control method disclosed in the above embodiments can be executed by a trained inference model, which can be found in [reference needed]. Figure 2 The model architecture shown includes a vision-language encoding module, a point cloud encoding module, a fusion module, a trajectory encoding module, an action flow prediction module, and a scene flow prediction module.
[0092] As mentioned earlier, this inference model can correspond to two different inference modes: real-time inference mode and scene deduction mode. The scene flow prediction module is pruned in real-time inference mode but retained in scene deduction mode. However, it should be noted that the training of this inference model is for the inference model as a whole. In other words, the scene flow prediction module also needs to participate in the training process and should not be pruned during the training phase.
[0093] The training process of this inference model will be described in detail below: Please see Figure 6 , Figure 6 A schematic diagram of the training process corresponding to an exemplary inference model is shown. For example... Figure 6 As shown, in an exemplary embodiment, the training process of the inference model includes S610-S620, which are described in detail below: S610: Obtain a training sample set containing multiple training samples, and input the multiple training samples into the inference model to obtain the action point flow output by the action flow prediction module and the scene point flow output by the scene flow prediction module. S620, calculate the first training loss value based on the action point flow output by the action flow prediction module and the scene point flow output by the scene flow prediction module, and optimize the parameters of the inference model based on the first training loss value.
[0094] In the above training process, each training sample includes natural language command samples, observed image samples, point cloud samples, historical trajectory samples, action point flow ground truth, and scene point flow ground truth. It may also include camera calibration parameters and other information, which are not limited here. It should be noted that the point cloud samples are still based on a unified three-dimensional coordinate system, and the process of constructing point cloud samples is the same as the process of constructing point cloud data based on a unified three-dimensional coordinate system during the inference stage described in the aforementioned embodiments, and will not be repeated here. The process of inputting multiple training samples into the inference model to obtain the corresponding action point flow output by the action flow prediction module and the corresponding scene point flow output by the scene flow prediction module in this embodiment can also be found in the action point flow prediction and scene point flow prediction processes described in the aforementioned embodiments, and will not be repeated here either.
[0095] In this embodiment, a first training loss value is calculated based on the action point flow output by the action flow prediction module and the scene point flow output by the scene flow prediction module. The parameters of the inference model are then optimized based on this first training loss value. This can be understood as the inference model simultaneously receiving two supervision signals: one is the ground truth value of the action point flow from the action flow prediction module, and the other is the ground truth value of the scene point flow from the scene flow prediction module. The loss values corresponding to these two supervision signals are calculated, and a weighted sum of the two loss values is obtained to obtain a total loss value. This total loss value is then used as the first training loss value to update the parameters of each module of the inference model. Specifically, the loss value can be calculated using L1 distance loss or mean squared error loss, or a combination of both, or other loss forms; no limitation is imposed here.
[0096] Figure 7 It shows the relationship with Figure 6 The diagram illustrates the training architecture corresponding to the training process shown. (Example) Figure 7 As shown, multi-source samples are used to calculate action loss and scene loss separately, and joint optimization is performed based on the action loss and scene loss. The multi-source samples can be understood as multiple training samples in the training sample set corresponding to different sources. For example, different training samples may correspond to different sampling rates (Frames Per Second, FPS) or different prediction window lengths (Horizon). Furthermore, multiple training samples can be collected by either a dual-arm robot or a single-arm robot; no restrictions are imposed here. In the trajectory encoding module, historical trajectory samples from different sources can all be mapped to the same spatiotemporal domain, without requiring all sources to be forcibly pruned to the exact same discrete-time index.
[0097] Therefore, the embodiments of this application do not separate "how the robot moves" and "how the world changes after the action" into two separate problems. Instead, they completely change the traditional "phased and fragmented" training paradigm of robot learning. During model training, all modules in this inference model (from the lowest-level vision-language encoding to the highest-level action / scene prediction) share the same set of backpropagation gradients, achieving joint optimization. This joint optimization mechanism not only ensures the executability of the predicted actions but also guarantees the consistency between action execution and scene changes (causal consistency). This allows scene flow prediction to act as a physical constraint regularization term, and the gradient information it generates can be backpropagated to the action flow prediction module and the low-level feature extraction module. This not only forces the model to generate physically reasonable actions that conform to geometric and dynamic constraints but also promotes the deep fusion of multimodal features in a unified semantic space, significantly improving the model's understanding of the essence of physical interaction and its cross-scene generalization ability.
[0098] The inference model trained through the embodiments of this application no longer merely fits motion trajectories at the pixel level, but understands the purposefulness of actions at the semantic level, thereby predicting whether the current action will lead to the achievement of the target state. This ability to perceive future states enables the model to effectively avoid actions that, while conforming to current image features, would lead to incorrect future states when generating action point flows during the inference phase. This significantly improves the robot's understanding of natural language commands and control accuracy, achieving a qualitative leap from "blind execution" to "goal-oriented" operation.
[0099] In one exemplary embodiment, during the training of the inference model, a portion of the pre-trained backbone network can be frozen, and only the adaptation layer is trained. In other words, one or more of the visual-language encoding module, point cloud encoding module, fusion module, and trajectory encoding module included in the inference model can be pre-trained modules. During the training of the inference model, the parameters of these pre-trained modules are frozen, and only the parameters of the remaining modules that are not frozen are updated. This achieves efficient transfer learning and training acceleration, effectively suppresses the risk of catastrophic forgetting and overfitting, and reduces computational resource consumption.
[0100] In another exemplary embodiment, during the training of the inference model, based on preset constraints, the ground truth values of action point flows contained in the training samples and the action point flows output by the action flow prediction module are selectively input into the scene flow prediction module to obtain the scene point flows output by the scene flow prediction module. For example, the ground truth values of action point flows contained in the training samples can be forced as the input to the scene flow prediction module, instead of the action point flows output by the action flow prediction module. This allows the scene flow prediction module to focus on learning physical laws, but it is prone to input dependency. Alternatively, the ground truth values of action point flows contained in the training samples and the action point flows output by the action flow prediction module can be used together as the input to the scene flow prediction module, which can greatly enhance the robustness of the model. For example, during the training process, as the number of rounds increases, the proportion of using the action point flows output by the action flow prediction module as the input to the scene flow prediction module can be gradually increased. This simulates the real scenario during inference and effectively alleviates the gap between the training and inference phases. Alternatively, within the same training batch, the ground truth values of the action point flow contained in the training samples and the action point flow output by the action flow prediction module can be randomly mixed as input to the scene flow prediction module, without any restrictions.
[0101] In another exemplary embodiment, the calculation of the first training loss value may further incorporate one or more of the following, in addition to the aforementioned action loss value and scene loss value: action-scene alignment loss value, contact area weighted loss value, future geometry reconstruction loss value, and cycle consistency loss value. Among them, the action-scene alignment loss value is used to constrain the spatiotemporal consistency between action point flow and scene point flow, ensuring that the actuator action and the scene object response conform to the physical causal relationship, and avoiding the logical disconnect between action prediction and scene inference; the contact area weighted loss applies higher penalty weights to the features and displacement prediction of the contact area, guiding the model to focus on the key physical contact points of the interaction between the actuator and the environment, thereby significantly improving the stability and reliability of grasping and manipulation tasks; the future geometry reconstruction loss value reconstructs the three-dimensional geometry of future moments based on the predicted scene point flow and constrains its consistency with the real geometry, thereby strengthening the model's understanding of rigid body dynamics and geometric invariance, and effectively suppressing non-physical deformation and penetration phenomena; the cyclic consistency loss introduces time reversal constraints, requiring the model to cyclically and consistently reconstruct the initial state from the predicted future state, thereby deepening the model's ability to model state transition probabilities and the reversibility of physical processes, and improving the rationality of long-term inference.
[0102] Please continue reading. Figure 8 and Figure 9 ,in Figure 8 A schematic diagram of the training process corresponding to another exemplary inference model is shown. Figure 9 It shows the relationship with Figure 8 The diagram illustrates the training architecture corresponding to the training process shown. (Example) Figure 8 As shown, in an exemplary embodiment, the training process of the inference model includes S810-S830, which are described in detail below: S810 initializes the corresponding body adaptation module for each of the multiple body actuators. The body adaptation module is used to convert the action point flow into control instructions adapted to the body actuator. S820 freezes the parameters of the pre-trained inference model and trains each ontology adaptation module based on the inference model with frozen parameters. S830 unfreezes the parameters of the pre-trained inference model and performs joint fine-tuning on the unfrozen inference model and the various pre-trained ontology adaptation modules.
[0103] In the embodiments of this application, the training dataset needs to be oriented towards proprio actuators from different sources, such as manifolds, single-arm robots, dual-arm robots, and different grippers, etc., without any limitation. The training samples also need to include the control ground truth of the proprio actuators, such as the joint angles of manifolds, the joint angles of single-arm robots, the poses of the left / right end effectors of dual-arm robots, and the gripper opening and closing degrees.
[0104] During training, it is necessary to determine the unified tool center point (TCP) coordinate system corresponding to multiple ontology actuators. Then, the action point flow output by the parameter-frozen inference model is mapped to an action point flow based on the TCP coordinate system, so that each ontology adaptation module can convert control commands to the mapped action point flow. For example, the X-axis of the TCP coordinate system points from the link to the gripper, the Y-axis is parallel to the opening and closing direction of the gripper, and the Z-axis is determined based on the right-hand rule.
[0105] It is understandable that the above is based on Figure 6 The process of training the inference model in the illustrated embodiment can be regarded as a pre-training process for the inference model, and the resulting inference model is the inference model that has completed pre-training.
[0106] After initializing the corresponding ontology adaptation modules for multiple ontology executors, the parameters of the pre-trained inference model are frozen. Based on the parameter-frozen inference model, each ontology adaptation module is trained separately. Adapters from different ontology types are trained in parallel without interference. Training samples from multiple ontology executors can be mixed within the same batch of data. Each training sample is only fed into the ontology adapter of its corresponding ontology executor to calculate the training loss. Thus, after the parameters are frozen, the pre-trained inference model serves as a shared backbone among the multiple ontology adaptation modules. The action point flow output by the shared backbone contains common action semantics. The multiple ontology adaptation modules are only responsible for translating this semantics into control instructions for their respective ontology executors. This avoids modifying the general rules of the shared backbone and prevents interference from mapping differences between different ontology executors, thereby completely confining the differences between ontology executors within their respective ontology adapters.
[0107] The training loss value can be obtained by calculating the loss between the control quantity output by the ontology adapter and the corresponding control truth value, such as the L1 distance loss or L2 distance loss between the two; it can also be further calculated by back-projecting the action point flow output by the ontology adaptation module back to the coordinate system of the unified tool center point, and calculating the loss between the action point flow obtained by back-projection and the action point flow output by the shared backbone, and the sum or weighted sum of the two losses is used as the final adaptation loss.
[0108] When multiple body actuators are deployed on the same robot, such as a dual-arm robot with left and right end effectors, the left and right end effectors can be mapped to two independent unified tool center point coordinate systems to avoid execution conflicts.
[0109] In S830, the parameters of the pre-trained inference model are unfrozen, and joint fine-tuning is performed on the inference model after parameter unfreezing and each pre-trained ontology adaptation module. During this stage, action loss, scene loss, and adaptation loss can be calculated, and the three losses are weighted and summed to obtain the total loss. This total loss is then used to optimize the parameters of the entire inference model and adapter, thereby solving the problem that the shared backbone is not fully adapted to multi-aspect scenarios. At the same time, the consistency of general point flow generation and exclusive mapping is optimized to ensure that the control quantity output by the ontology adaptation module conforms to the physical constraints of the ontology actuator and does not violate the general physical laws learned by the shared backbone.
[0110] During the joint fine-tuning process, the weight of adaptation loss can be gradually reduced while the weight of action loss and scene loss can be increased to ensure that the physical laws of the shared backbone are not interfered with by the exclusive mapping of the ontology adaptation module.
[0111] Therefore, in the embodiments of this application, a unified tool center point coordinate system is used throughout the training process, completely resolving the optimization conflict between action point flow and the ontology executor. Furthermore, the shared backbone learns general physical laws, independent of any ontology executor's joint dimensions, supporting joint training of human hands, single arms, dual arms, and different grippers, achieving cross-ontology generalization of the inference model. The ontology adaptation module freezes the shared backbone during the pre-training phase, and multiple ontology adaptation modules can be trained in parallel, significantly reducing training costs. When adding a new ontology executor, only minor adjustments to the adapter are needed, making training highly efficient.
[0112] Figure 10 A block diagram of an exemplary robot control device is shown. This device can be applied to... Figure 1 The implementation environment shown is configured in Figure 1 The device is specifically configured on the controller built into the robot 120 in the illustrated implementation environment. This device can also be applied to other implementation environments and specifically configured on the controller built into the robot in other implementation environments to achieve motion control of the robot; the embodiments of this application do not limit this.
[0113] like Figure 10 As shown, in an exemplary embodiment, the robot control device includes a point cloud construction module 1010, a semantic encoding module 1020, a multimodal fusion module 1030, a trajectory encoding module 1040, and a generation and control module 1050.
[0114] The point cloud construction module 1010 is configured to construct point cloud data based on a unified three-dimensional coordinate system based on the robot's multi-source observation data. The point cloud data includes the three-dimensional position information associated with the robot's observation images. The semantic encoding module 1020 is configured to jointly encode natural language commands and observation images to obtain semantically aligned visual encoding features and language encoding features, and add the three-dimensional position information associated with the observation images to the visual encoding features. The multimodal fusion module 1030 is configured to extract the geometric encoding features of the point cloud data, and perform feature fusion processing on the geometric encoding features, language encoding features, and visual encoding features after adding three-dimensional position information to obtain multimodal fusion features. The trajectory encoding module 1040 is configured to encode historical trajectories based on a unified three-dimensional coordinate system to obtain trajectory encoding features. The generation and control module 1050 is configured to generate an action point flow based on the multimodal fusion features and trajectory encoding features, and control the robot based on the action point flow.
[0115] In another exemplary embodiment, based on the foregoing scheme, the device further includes an execution feedback module, which is configured to: acquire new observation data of the robot after responding to control, and generate an actual trajectory based on a unified three-dimensional coordinate system based on the new observation data; save the actual trajectory as a new historical trajectory so that the new historical trajectory can be used for the next round of control.
[0116] In another exemplary embodiment, based on the aforementioned scheme, the generation and control module 1050 generates multiple sets of action point flows based on multimodal fusion features and trajectory coding features; the generation and control module 1050 is further configured to: generate corresponding scene point flows based on each set of action point flows, the scene point flows characterizing the scene consequences corresponding to the action point flows, select a target action point flow from the multiple sets of action point flows according to the scene point flows corresponding to each set of action point flows, and control the robot based on the target action point flow.
[0117] In another exemplary embodiment, based on the aforementioned scheme, the trajectory encoding module 1040 is further configured to: acquire trajectory sampling information of each spatial anchor point at multiple consecutive sampling times, the trajectory sampling information including three-dimensional position information; and encode the trajectory sampling information of the same spatial anchor point at multiple consecutive sampling times into a feature vector of fixed dimension to obtain trajectory encoding features.
[0118] In another exemplary embodiment, based on the aforementioned scheme, the trajectory encoding module 1040 is further configured to: encode the timestamps corresponding to the same spatial anchor point at multiple consecutive sampling times into multiple consecutive negative values.
[0119] In another exemplary embodiment, based on the aforementioned scheme, the multimodal fusion module 1030 is further configured to: calculate the cross-attention between language coding features and visual coding features with added three-dimensional position information, and generate a two-dimensional weight mask based on the cross-attention; map the two-dimensional weight mask to a three-dimensional weight mask based on a unified three-dimensional coordinate system, and perform spatial gating modulation on the visual coding features with added three-dimensional position information through the three-dimensional weight mask to obtain modulated visual coding features; and fuse the modulated visual coding features and geometric coding features to obtain multimodal fusion features.
[0120] In another exemplary embodiment, based on the foregoing scheme, the device further includes a model training module, which is configured to: acquire a training sample set containing multiple training samples, input the multiple training samples into the inference model, obtain the action point flow output by the action flow prediction module and the scene point flow output by the scene flow prediction module; calculate the training loss value based on the action point flow output by the action flow prediction module and the scene point flow output by the scene flow prediction module, and optimize the parameters of the inference model based on the training loss value; wherein, the inference model includes a visual-language encoding module, a point cloud encoding module, a fusion module, a trajectory encoding module, an action flow prediction module, and a scene flow prediction module; the scene flow prediction module is pruned in the real-time inference mode and retained in the scene inference mode.
[0121] In another exemplary embodiment, based on the aforementioned scheme, the model training module is further configured to: selectively input the true values of the action point flow contained in the training samples and the action point flow output by the action flow prediction module into the scene flow prediction module based on preset constraints, so as to obtain the scene point flow output by the scene flow prediction module.
[0122] In another exemplary embodiment, based on the aforementioned scheme, the model training module is further configured to: initialize corresponding ontology adaptation modules for multiple ontology executors, wherein the ontology adaptation modules are used to convert action point flows into control instructions adapted to the ontology executors; freeze the parameters of the pre-trained inference model, and train each ontology adaptation module based on the inference model with frozen parameters; unfreeze the parameters of the pre-trained inference model, and perform joint fine-tuning on the inference model with unfrozen parameters and each pre-trained ontology adaptation module.
[0123] In another exemplary embodiment, based on the aforementioned scheme, the model training module is further configured to: determine a unified tool center point coordinate system corresponding to multiple ontology executors; map the action point flow output by the parameter-frozen inference model to an action point flow based on the unified tool center point coordinate system, so as to perform control command conversion on the mapped action point flow.
[0124] It should be noted that the apparatus and method provided in the above embodiments belong to the same concept, and the specific ways in which each module and unit performs operations have been described in detail in the method embodiments, and will not be repeated here. In practical applications, the robot control device provided in the above embodiments can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above, and this is not a limitation here.
[0125] Embodiments of this application also provide an electronic device, including: one or more processors; and a memory for storing one or more computer programs, which, when executed by the one or more processors, cause the electronic device to implement the robot control methods provided in the above embodiments.
[0126] Figure 11 A schematic diagram of a computer system suitable for implementing an electronic device according to embodiments of this application is shown. It should be noted that the electronic device can be... Figure 1 The robot 120 in the illustrated implementation environment can also be a terminal or server in other implementation environments; no restrictions are imposed here. It should also be noted that... Figure 11 The computer system 1100 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0127] like Figure 11 As shown, the computer system 1100 includes a Central Processing Unit (CPU) 1101, which can perform various appropriate actions and processes based on a computer program stored in Read-Only Memory (ROM) 1102 or a computer program loaded from storage portion 1108 into Random Access Memory (RAM) 1103, such as performing the methods described in the above embodiments. Various computer programs and data required for system operation are also stored in RAM 1103. The CPU 1101, ROM 1102, and RAM 1103 are interconnected via bus 1104. An input / output (I / O) interface 1105 is also connected to bus 1104.
[0128] The following components are connected to I / O interface 1105: an input section 1106 including a keyboard, mouse, etc.; an output section 1107 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1108 including a hard disk, etc.; and a communication section 1109 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 1109 performs communication processing via a network such as the Internet. A drive 1110 is also connected to I / O interface 1105 as needed. Removable media 1111, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1110 as needed so that computer programs read from them can be installed into storage section 1108 as needed.
[0129] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1109, and / or installed from removable medium 1111. When the computer program is executed by central processing unit (CPU) 1101, it performs various functions defined in the system of this application.
[0130] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0131] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0132] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.
[0133] Another aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor of an electronic device, implements the robot control method as described above. This computer-readable storage medium may be included in the electronic device described in the above embodiments, or it may exist independently and not assembled into the electronic device.
[0134] Another aspect of this application provides a computer program product comprising a computer program stored in a computer-readable storage medium. A processor of an electronic device reads the computer program from the computer-readable storage medium and executes the computer program, causing the electronic device to perform the robot control methods provided in the various embodiments described above.
[0135] The above description is merely a preferred exemplary embodiment of this application and is not intended to limit the implementation of this application. Those skilled in the art can easily make corresponding modifications or alterations based on the main concept and spirit of this application. Therefore, the scope of protection of this application should be determined by the scope of protection claimed in the claims.
Claims
1. A robot control method, characterized in that, The method includes: Based on the robot's multi-source observation data, point cloud data based on a unified three-dimensional coordinate system is constructed. The point cloud data contains three-dimensional position information associated with the robot's observation images. The natural language instructions and the observed image are jointly encoded to obtain semantically aligned visual coding features and language coding features, and the three-dimensional position information associated with the observed image is added to the visual coding features. Extract the geometric coding features of the point cloud data, and perform feature fusion processing on the geometric coding features, the language coding features, and the visual coding features after adding the three-dimensional position information to obtain multimodal fusion features; The historical trajectories based on the unified three-dimensional coordinate system are encoded to obtain trajectory encoding features; An action point flow is generated based on the multimodal fusion features and the trajectory encoding features, and the robot is controlled based on the action point flow. Acquire new observation data of the robot after responding to control, and generate an actual trajectory based on the unified three-dimensional coordinate system based on the new observation data; The actual trajectory is saved as a new historical trajectory so that the new historical trajectory can be used in the next round of control.
2. The method according to claim 1, characterized in that, The action point stream generated based on the multimodal fusion features and the trajectory encoding features includes multiple sets; the method further includes: Based on each group of action point flows, a corresponding scene point flow is generated, and the scene point flow represents the scene consequence corresponding to the action point flow. The control of the robot based on the motion point flow includes: Based on the scene point flow corresponding to each group of action point flows, a target action point flow is selected from multiple groups of action point flows, and the robot is controlled based on the target action point flow.
3. The method according to claim 1, characterized in that, The process of encoding historical trajectories based on the unified three-dimensional coordinate system to obtain trajectory encoding features includes: For each spatial anchor point, its trajectory sampling information is obtained at multiple consecutive sampling times, and the trajectory sampling information includes three-dimensional position information; The trajectory sampling information of the same spatial anchor point at multiple consecutive sampling times is encoded into a fixed-dimensional feature vector to obtain the trajectory encoding feature.
4. The method according to claim 3, characterized in that, Encoding the trajectory sampling information of the same spatial anchor point into a fixed-dimensional feature vector at multiple consecutive sampling times includes: The timestamps corresponding to the same spatial anchor point at multiple consecutive sampling times are encoded as multiple consecutive negative values.
5. The method according to claim 1, characterized in that, The feature fusion processing of the geometric coding features, the language coding features, and the visual coding features after incorporating the three-dimensional position information to obtain multimodal fusion features includes: The cross attention between the language encoding features and the visual encoding features after incorporating the three-dimensional position information is calculated, and a two-dimensional weight mask is generated based on the cross attention. The two-dimensional weight mask is mapped to a three-dimensional weight mask based on the unified three-dimensional coordinate system, and the visual coding features after adding the three-dimensional position information are spatially gated and modulated using the three-dimensional weight mask to obtain the modulated visual coding features. The modulated visual coding features and the geometric coding features are fused to obtain the multimodal fusion features.
6. The method according to claim 1, characterized in that, The method is executed through a trained inference model, which includes a visual-language encoding module, a point cloud encoding module, a fusion module, a trajectory encoding module, an action flow prediction module, and a scene flow prediction module; wherein, the scene flow prediction module is pruned in real-time inference mode and retained in scene deduction mode.
7. The method according to claim 6, characterized in that, The training process of the inference model includes the following steps: Obtain a training sample set containing multiple training samples, and input the multiple training samples into the inference model to obtain the action point flow output by the action flow prediction module and the scene point flow output by the scene flow prediction module. The training loss value is calculated based on the action point flow output by the action flow prediction module and the scene point flow output by the scene flow prediction module, and the parameters of the inference model are optimized based on the training loss value.
8. The method according to claim 7, characterized in that, The method further includes: Based on preset constraints, the ground truth values of the action point flow contained in the training samples and the action point flow output by the action flow prediction module are selectively input into the scene flow prediction module to obtain the scene point flow output by the scene flow prediction module.
9. The method according to claim 7, characterized in that, The training process of the inference model also includes: For multiple body actuators, initialize their respective corresponding body adaptation modules. The body adaptation modules are used to convert the action point flow into control instructions adapted to the body actuators. The parameters of the pre-trained inference model are frozen, and each ontology adaptation module is trained based on the inference model with the parameters frozen. The parameters of the pre-trained inference model are unfrozen, and joint fine-tuning is performed on the inference model after parameter unfreezing and each pre-trained ontology adaptation module.
10. The method according to claim 9, characterized in that, The training process of the inference model also includes: Determine the unified tool center point coordinate system corresponding to the plurality of body actuators; The action point stream output by the inference model after the parameters are frozen is mapped to an action point stream based on the coordinate system of the unified tool center point, so as to convert control commands for the mapped action point stream.
11. The method according to claim 10, characterized in that, The X-axis of the unified tool center point coordinate system points from the connecting rod to the gripper, the Y-axis is parallel to the opening and closing direction of the gripper, and the Z-axis is determined based on the right-hand rule.
12. A robot control device, characterized in that, The device includes: The point cloud construction module is configured to construct point cloud data based on a unified three-dimensional coordinate system based on the robot's multi-source observation data. The point cloud data includes the three-dimensional position information associated with the robot's observation images. The semantic encoding module is configured to jointly encode the natural language command and the observed image to obtain semantically aligned visual encoding features and language encoding features, and to add the three-dimensional position information associated with the observed image to the visual encoding features; The multimodal fusion module is configured to extract the geometric coding features of the point cloud data, and perform feature fusion processing on the geometric coding features, the language coding features, and the visual coding features after adding the three-dimensional position information to obtain multimodal fusion features; The trajectory encoding module is configured to encode historical trajectories based on the unified three-dimensional coordinate system to obtain trajectory encoding features; The generation and control module is configured to generate a motion point flow based on the multimodal fusion features and the trajectory encoding features, and to control the robot based on the motion point flow; The execution feedback module is configured to acquire new observation data of the robot after responding to control, and generate an actual trajectory based on the unified three-dimensional coordinate system based on the new observation data; the actual trajectory is saved as a new historical trajectory so that the new historical trajectory can be used for the next round of control.
13. An electronic device, characterized in that, include: One or more processors; A memory for storing one or more computer programs that, when executed by one or more processors, cause the electronic device to perform the method as described in any one of claims 1-11.
14. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by the processor of the electronic device, causes the electronic device to perform the method of any one of claims 1-11.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor of the electronic device, it implements the method as described in any one of claims 1-11.