A robot control method and device, electronic equipment, storage medium and product

CN121989233BActive Publication Date: 2026-09-22XINGHAITU (SUZHOU) ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610058912.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-16
Publication Date
2026-09-22
Estimated Expiration
2046-01-16

AI Technical Summary

Technical Problem

[0004]本发明提供了一种机器人控制方法、装置、电子设备、存储介质及产品,以解决现有技术中模型结构仅适用于单一任务,且机器人动作序列的准确性较低的问题

Benefits of technology

[0021]本发明实施例的技术方案,通过获取目标机器人的操作场景的视觉数据、目标任务的文本指令数据,以及目标机器人的本体状态特征,分别确定视觉数据的视觉特征、文本指令数据的文本特征,根据视觉特征、文本特征、本体状态特征和预设扩散步确定多模态特征,从高斯分布采样初始动作噪声,根据初始动作噪声和多模态特征确定目标多模态特征,基于预设扩散模型确定目标多模态特征对应的机器人动作序列,基于机器人动作序列控制目标机器人执行目标任务,实现视觉与文本联合引导模型,能准确识别目标物体与任务顺序,实现复杂多对象操作,提高机器人动作序列的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121989233B_ABST
    Figure CN121989233B_ABST
Patent Text Reader

Abstract

The application discloses a robot control method and device, electronic equipment, storage medium and product, and relates to the technical field of artificial intelligence. The method comprises the following steps: acquiring visual data of an operation scene of a target robot, text instruction data of a target task, and ontology state features of the target robot, respectively determining visual features of the visual data and text features of the text instruction data; determining multi-modal features according to the visual features, the text features, the ontology state features and a preset diffusion step; sampling initial action noise from a Gaussian distribution, determining target multi-modal features according to the initial action noise and the multi-modal features, and determining a robot action sequence corresponding to the target multi-modal features based on a preset diffusion model; wherein the preset diffusion model is constructed by a noise prediction network and a denoising diffusion implicit model, and the preset noise prediction network is composed of a diffusion transformer; and the target robot is controlled to perform the target task based on the robot action sequence, so that the accurate determination of the robot action sequence is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a robot control method, device, electronic device, storage medium, and product. Background Technology

[0002] Existing robot control systems typically employ imitation learning methods such as behavior cloning, directly regressing control commands by inputting sensor observations (images, force feedback, position, etc.) into the model. While these single-task strategies can accomplish specific actions with limited demonstration data, they generally exhibit vulnerability to task variations, environmental disturbances, or complex, long-term actions. In recent years, researchers have proposed the concept of "Large Behavior Models (LBMs)," which improve generalization ability by pre-training on large-scale multi-task datasets and then fine-tuning them on specific tasks. Extensive experiments have demonstrated that, compared to single-task models trained from scratch, models pre-trained on multiple tasks require fewer examples to achieve similar or even better success rates on new tasks.

[0003] Within the framework of large-scale behavioral models, diffusion policies have attracted attention due to their ability to generate long-term, multimodal actions. Diffusion models, by progressively denoising Gaussian noise to obtain complex output structures, have made groundbreaking progress in image generation. Recent work has applied the diffusion concept to robot action generation, enabling models to not only output multimodal, continuous action sequences but also share model structure and parameters across different tasks. However, existing multi-task diffusion policies often rely solely on visual perception, lacking high-level natural language instructions, or their model structures are only applicable to a single task, failing to accommodate multi-object operations and resulting in poor accuracy in determining robot action sequences. Therefore, overcoming the difficulty in task differentiation caused by the lack of language instruction guidance in traditional multi-task policies and improving the accuracy of determining robot action sequences has become an urgent problem to be solved. Summary of the Invention

[0004] This invention provides a robot control method, device, electronic device, storage medium, and product to solve the problems in the prior art where the model structure is only applicable to a single task and the accuracy of the robot action sequence is low.

[0005] According to one aspect of the present invention, a robot control method is provided, wherein the method includes:

[0006] The visual data of the target robot's operating scene, the text command data of the target task, and the body state characteristics of the target robot are acquired, and the visual features of the visual data and the text features of the text command data are determined respectively.

[0007] Multimodal features are determined based on the visual features, text features, ontology state features, and a preset diffusion step;

[0008] Initial motion noise is sampled from a Gaussian distribution. Target multimodal features are determined based on the initial motion noise and the multimodal features. The robot motion sequence corresponding to the target multimodal features is determined based on a preset diffusion model. The preset diffusion model is constructed from a noise prediction network and a denoising diffusion implicit model. The preset noise prediction network is composed of a diffusion transformer.

[0009] The target robot is controlled to perform the target task based on the robot's action sequence.

[0010] According to another aspect of the present invention, a robot control device is provided, wherein the device comprises:

[0011] The feature acquisition module is used to acquire visual data of the target robot's operation scene, text command data of the target task, and the body state features of the target robot, and to determine the visual features of the visual data and the text features of the text command data, respectively.

[0012] The feature fusion module is used to determine multimodal features based on the visual features, the text features, the ontology state features, and a preset diffusion step;

[0013] The motion determination module is used to sample initial motion noise from a Gaussian distribution, determine target multimodal features based on the initial motion noise and the multimodal features, and determine the robot motion sequence corresponding to the target multimodal features based on a preset diffusion model; wherein, the preset diffusion model is constructed by a noise prediction network and a denoising diffusion implicit model, and the preset noise prediction network is composed of a diffusion transformer;

[0014] A robot control module is used to control the target robot to perform the target task based on the robot's action sequence.

[0015] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0016] At least one processor; and

[0017] A memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the robot control method according to any embodiment of the present invention.

[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the robot control method according to any embodiment of the present invention.

[0020] According to another aspect of the present invention, embodiments of the present invention also provide a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the robot control method of any embodiment of the present invention.

[0021] The technical solution of this invention acquires visual data of the target robot's operation scene, text command data of the target task, and the body state characteristics of the target robot. It then determines the visual features of the visual data and the text features of the text command data, respectively. Based on the visual features, text features, body state characteristics, and a preset diffusion step, it determines multimodal features. Initial motion noise is sampled from a Gaussian distribution. Based on the initial motion noise and multimodal features, it determines the target multimodal features. Based on a preset diffusion model, it determines the robot motion sequence corresponding to the target multimodal features. Based on the robot motion sequence, it controls the target robot to execute the target task, realizing a vision-text joint guidance model. This model can accurately identify target objects and task sequences, achieve complex multi-object operations, and improve the accuracy of robot motion sequences.

[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a flowchart of a robot control method provided according to Embodiment 1 of the present invention;

[0025] Figure 2 This is a flowchart of a robot control method provided according to Embodiment 2 of the present invention;

[0026] Figure 3 This is a flowchart of a robot control method provided according to Embodiment 3 of the present invention;

[0027] Figure 4 This is a schematic diagram of the structure of a robot control device according to Embodiment 4 of the present invention;

[0028] Figure 5 This is a schematic diagram of the structure of an electronic device that implements the robot control method of this invention. Detailed Implementation

[0029] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0030] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0031] Example 1

[0032] Figure 1 This is a flowchart of a robot control method according to Embodiment 1 of the present invention. This embodiment is applicable to situations where robot actions need to be predicted. The method can be executed by a robot control device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:

[0033] S110. Obtain visual data of the target robot's operation scene, text instruction data of the target task, and body state characteristics of the target robot, and determine the visual features of the visual data and the text features of the text instruction data respectively.

[0034] In this context, the target robot can be understood as a robotic arm or mobile robot that needs to perform a specified task; it is the main entity executing the action. In actual operation, the target robot can include, but is not limited to, industrial assembly robotic arms and service robot robotic arms, and the robotic arm can be single-arm or dual-arm. The operating scenario can be understood as the environment in which the target robot is currently operating. Visual data can be understood as data collected using visual sensors; generally, visual data can be continuous images. In actual operation, continuous images or video frames of the operating environment collected by the robot's wrist camera or a fixed scene camera can be used as visual data. The target task refers to the task to be performed; for example, the target task can include a kitchen scene task, an assembly task, etc. In practical applications, there can be one or more target tasks. Text instruction data can be understood as a natural language task description input by the user, such as "Place the three parts onto the tray in order" or "Put the banana into the plate on the right." Generally, text instruction data can be collected through voice input or text input. The body state characteristics refer to the numerical characteristics of the target robot's own state, including joint angles, joint velocities, etc., which are characteristic information of the robot's perception of its own actions. Visual features can be understood as high-dimensional vectors obtained after encoding visual data. For example, visual features can be classification token (CLS) feature vectors extracted by a pre-trained visual encoder (such as CLIP Vision Transformer), facilitating subsequent model understanding of the visual scene. Text features can be understood as high-dimensional vectors obtained after encoding text instruction data. These are feature vectors extracted by a pre-trained text encoder (such as CLIP text encoder) paired with the visual encoder, and then adjusted in dimension by a projection layer. In one embodiment, ontology state features can be numerical features.

[0035] In this embodiment, continuous image frames of the operational scene can be acquired in real time using a camera configured on the target robot or a fixed camera in the scene as visual data. The system receives text instructions for the target task input by the user, and then reads the body state data (joint angles, joint velocities, etc.) from the robot control system, converting the body state data into numerical body state features. The acquired visual data is input into a pre-trained visual encoder to obtain visual features, and the instruction data is then input into a pre-trained text encoder paired with the visual encoder to obtain the original feature vector. The feature dimensions are adjusted through a preset projection layer to make the text feature dimensions consistent with the visual feature dimensions. In actual operation, the body state data can undergo data preprocessing such as outlier cleaning and standardization before being converted into numerical features as body state features. The visual features, text features, and body state features have the same dimensions.

[0036] S120. Determine multimodal features based on visual features, text features, ontological state features, and preset diffusion steps.

[0037] The preset diffusion step is a stage identifier in the denoising iteration process of the preset diffusion model. It refers to the step number corresponding to each denoising operation in the iterative process from initial motion noise to the final effective motion sequence. The larger the preset diffusion step value, the higher the noise ratio in the motion vector; the smaller the value, the lower the noise ratio, and the closer it is to the real motion. Multimodal features refer to the joint features fused from visual features, text features, ontological state features, and preset diffusion step encoding. These features can simultaneously carry environmental information, task instructions, robot's own state, and diffusion stage information, so as to facilitate the subsequent generation of accurate motion sequences through the preset diffusion model.

[0038] In this embodiment, the diffusion step code corresponding to the preset diffusion step can be determined first. Then, the visual features, text features, and ontology state features are expanded according to time steps, and the diffusion time is embedded according to the diffusion step code to obtain multimodal features. In actual operation, sinusoidal position coding can be used, followed by two layers of fully connected networks to generate the diffusion step code. The dimensions of the visual features, text features, and ontology state features are adjusted to be consistent through a projection layer. The visual features, text features, ontology state features, and diffusion step code features are fused with the diffusion step code features by adding them element by element to obtain multimodal features.

[0039] S130. Sample initial motion noise from a Gaussian distribution, determine target multimodal features based on initial motion noise and multimodal features, and determine the robot motion sequence corresponding to the target multimodal features based on a preset diffusion model.

[0040] The preset diffusion model is constructed from a noise prediction network and a denoising diffusion implicit model. The preset noise prediction network is composed of a diffusion transformer (DiT).

[0041] The initial motion noise refers to a vector randomly sampled from a standard Gaussian distribution, with dimensions consistent with the target robot motion sequence. For example, when the required robot motion sequence has 16 time steps × 20 motion dimensions, the initial motion noise also has 16 time steps × 20 motion dimensions. The pre-defined diffusion model can be understood as the model used to generate the robot motion sequence. It consists of a noise prediction network and a denoising diffusion implicit model (DDIM). It can progressively denoise the initial motion noise through multimodal feature constraints to generate an effective motion sequence matching the task. The noise prediction network consists of multiple diffusion transformers, used to input target multimodal features and predict noise components in the initial motion noise, providing a basis for the denoising process. The diffusion transformer is an improved model based on the Transformer architecture. The denoising diffusion implicit model is an efficient diffusion denoising inference framework that does not require traversing all diffusion steps and can complete denoising in just k iterations. Its core is to use the noise estimation output by the noise prediction network and gradually update the action vector through the implicit sampling formula to remove noise and approximate the real action sequence.

[0042] In this embodiment, a vector with the same total dimension as the target robot motion sequence can be randomly sampled from a standard Gaussian distribution as initial motion noise. For example, the initial motion noise could be 16 time steps × 20 motion dimensions. The multimodal features of the initial motion noise are combined to obtain the target multimodal features. These target multimodal features are then input into a preset diffusion model, which iteratively denoises the motion according to a preset diffusion step to obtain the robot motion sequence. In actual operation, the noise estimate corresponding to the target multimodal features can be determined through the noise prediction network in the preset diffusion model. The noise estimate is then used to update the motion vector using a denoising diffusion implicit model based on the DDIM sampling formula, removing some noise to obtain a vector closer to the actual motion. This process continues until, after iterating the preset diffusion step, the initial motion noise is gradually transformed into a noise-free motion vector, which is then used as the robot motion sequence. In one embodiment, the robot action sequence may have the problem of being unrecognizable by the target robot. The final denoised action vector can be reshaped to transform it into a structured robot action sequence. For example, a 320-dimensional vector can be transformed into a [16,20] temporal action sequence, corresponding to 16 time steps and 20 action dimensions per time step, which facilitates the execution of the target robot.

[0043] S140. Control the target robot to perform the target task based on the robot action sequence.

[0044] In this embodiment, the robot action sequence can be sequentially sent to the robot control system in time step order to control the target robot to perform the target task. In actual operation, the first half of the time step actions can be executed first, reserving space for subsequent adjustments. For example, if the action sequence is of fixed length (e.g., 16 time steps), the first half of the time step actions (e.g., the first 8 time steps) can be executed first.

[0045] The robot action sequence refers to the structured action set output by the diffusion model, which is a time-series vector of fixed length (e.g., 16 time steps × 20 action dimensions). Each time step corresponds to a set of robot actions, such as joint pose increments, gripping commands, etc., which are action instructions that the robot can directly execute.

[0046] In this embodiment of the invention, visual data of the target robot's operation scene, text command data of the target task, and ontological state features of the target robot are acquired. Visual features of the visual data and text features of the text command data are determined respectively. Multimodal features are determined based on visual features, text features, ontological state features, and a preset diffusion step. Initial motion noise is sampled from a Gaussian distribution. Target multimodal features are determined based on the initial motion noise and multimodal features. The robot motion sequence corresponding to the target multimodal features is determined based on a preset diffusion model. The target robot is then controlled to execute the target task based on the robot motion sequence. This achieves a joint visual and textual guidance model, which can accurately identify target objects and task sequences, realize complex multi-object operations, and improve the accuracy of robot motion sequences.

[0047] In one embodiment, the training process of the noise prediction network includes:

[0048] Obtain the training dataset. Each historical data sample in the training dataset includes continuous visual images, historical task text instructions, historical robot body states, and corresponding real action sequences.

[0049] The visual features of continuous visual images are identified as historical visual features, and the text features of historical task text instructions are identified as historical text features.

[0050] Randomly sample a diffusion step from the preset total number of diffusion steps, and add target Gaussian noise to the real action sequence according to the noise intensity corresponding to the diffusion step to generate a noisy action sequence;

[0051] Sinusoidal position coding combined with a two-layer fully connected network is used to encode the diffusion step to obtain diffusion step coding features. Historical visual features, historical text features, historical robot body state and diffusion step coding features are fused to obtain training condition input.

[0052] The training conditions and noisy action sequences are input into the noise prediction network to generate prediction noise. The mean square error between the prediction noise and the target Gaussian noise is used as the loss function. The model parameters are updated by the gradient descent algorithm until the loss converges to the preset stable interval, thus completing the training.

[0053] The training dataset is a collection of samples used to train the noise prediction network, containing a large number of historical data samples, such as historical robot manipulation demonstration data. A historical data sample refers to a single data unit in the training dataset; each historical data sample includes continuous visual images, historical task text instructions, historical robot body states, and corresponding real motion sequences. Continuous visual images refer to consecutive image frames captured by the robot's wrist camera and a fixed-scene camera in the historical data samples, used to record environmental visual information during task execution. Historical task text instructions refer to the natural language descriptions of the corresponding tasks in the historical data samples. Historical robot body states refer to the robot's own state data during task execution in the historical data samples, including joint angles and joint velocities. Real motion sequences refer to the precise set of actions actually output by the robot during task execution in the historical data samples; these are fixed-length time-series vectors containing joint pose increments and control commands for each time step. Target Gaussian noise refers to Gaussian distributed noise added to the real motion sequences according to the noise intensity corresponding to a specific diffusion step. Noisy motion sequences refer to the motion vectors obtained by adding target Gaussian noise to the real motion sequences. Historical visual features refer to high-dimensional vectors obtained by encoding continuous visual images. Historical text features refer to high-dimensional vectors obtained by encoding historical task text instructions. Diffusion step encoding features refer to high-dimensional numerical vectors obtained by transforming discrete diffusion steps through sinusoidal position encoding and a two-layer fully connected network. The preset total number of diffusion steps refers to the maximum number of denoising iterations of the preset diffusion model (1000 steps), which can be set according to business needs. Diffusion step refers to the stage identifier of the denoising process of the preset diffusion model. Noise intensity is the magnitude of noise added for the corresponding diffusion step. Generally, in the preset total number of diffusion steps, the diffusion step value and noise intensity are positively correlated. For example, the noise intensity is highest at step 1000 and lowest at step 1. Predicted noise refers to the output of the noise prediction network, which is an estimate of the target Gaussian noise in the noisy action sequence. Mean squared error is a loss function used to calculate the difference between the predicted noise and the target Gaussian noise. The preset stability interval is the criterion for determining the convergence of model training. For example, the preset stability interval may include, but is not limited to, a loss fluctuation of ≤0.001 for 50 consecutive iterations.

[0054] In this embodiment, historical robot demonstration data can be acquired to construct a training dataset. Each historical data sample in the training dataset includes a continuous visual image, historical task text instructions, historical robot body state, and the corresponding real action sequence. The visual features of the continuous visual image are determined as historical visual features, and the text features of the historical task text instructions are determined as historical text features. A diffusion step is randomly sampled from a preset total number of diffusion steps, and the noise intensity corresponding to that diffusion step is determined. Target Gaussian noise is added to the real action sequence according to the noise intensity corresponding to that diffusion step to generate a noisy action sequence. A sinusoidal positional encoding combined with a two-layer fully connected network is used to encode the diffusion step, obtaining diffusion step encoding features. The historical visual features, historical text features, historical robot body state, and diffusion step encoding features are fused to obtain the training condition input. The training condition input and the noisy action sequence are input into a noise prediction network to generate prediction noise. The mean square error between the prediction noise and the target Gaussian noise is determined, and the model parameters are updated according to the gradient descent algorithm until the loss converges to a preset stable interval, completing the training.

[0055] Example 2

[0056] Figure 2 This is a flowchart of a robot control method according to Embodiment 2 of the present invention. This embodiment is a further optimization and extension based on the above embodiments, and can be combined with various optional technical solutions in the above embodiments. Figure 2 As shown, the method includes:

[0057] S201. Read the continuous frame images of the operation scene collected by the preset visual sensor as visual data, input the visual data into the pre-trained visual encoder, and extract the classification label output vector output by the visual encoder as visual features.

[0058] A visual encoder is a neural network model used to transform raw visual data (such as images and video frames) into high-dimensional semantic feature vectors. It can extract key information from visual data and encode it into a machine-understandable feature form. In one embodiment, the visual encoder may include, but is not limited to, CLIP Vision Transformer.

[0059] In this embodiment, continuous frame images from the operation scene collected by a preset visual sensor can be extracted as visual data, and the classification label output vector corresponding to the visual data can be determined by the visual encoder as visual features.

[0060] S202. Identify the text instruction data of the target task, input the text instruction data into a pre-trained text encoder, and extract the feature vector output by the text encoder as the text features of the text instruction data.

[0061] Among them, visual features and text features are the target feature dimensions.

[0062] The text encoder is a neural network model used to transform raw text data (such as natural language instructions or task descriptions) into high-dimensional semantic feature vectors. In this embodiment, the text encoder can be a text encoder paired with a visual encoder (such as the CLIP text encoder). The text encoder can extract language features and adjust the feature dimensions through a projection layer. During training, the language encoder remains frozen, and only the weights of the projection layer are adjusted.

[0063] In this embodiment, text instruction data of the target task can be extracted, and the text instruction data can be input into a pre-trained text encoder. The text encoder can then determine the feature vector corresponding to the text instruction data as text features.

[0064] S203. Collect the body sensor data of the target robot, map the body sensor data to the target feature dimension through a linear projection layer to obtain numerical features, and use the mapped numerical features as the body state features.

[0065] Among them, body sensor data can be understood as data collected through body sensors, such as joint angles and speeds.

[0066] In this embodiment, the body sensor data of the target robot can be extracted, the body sensor data can be mapped to the target feature dimension through a linear projection layer, and the mapped numerical features can be used as the body state features.

[0067] S204. Expand and align the visual features, text features, and ontology state features according to a preset time step as the initial multimodal features.

[0068] In this embodiment, visual features, text features, and ontology state features can be expanded according to a preset time step, and the visual features, text features, and ontology state features expanded according to the preset time step can be temporally aligned as the initial multimodal features.

[0069] S205. The preset diffusion step is mapped to a high-dimensional vector as the diffusion step encoding by sinusoidal position encoding and a fully connected network. The diffusion step encoding is then embedded into the initial multimodal features as multimodal features.

[0070] In the embodiment, the encoding dimension can be determined. Sine position encoding maps the preset diffusion step to a high-dimensional vector as the diffusion step encoding. The diffusion step encoding is embedded into the initial multimodal features as multimodal features by using an element-wise addition embedding method.

[0071] S206. Sample a random vector from a Gaussian distribution that is consistent with the multimodal feature dimension as the initial motion noise. Map the initial motion noise to the same feature dimension as the multimodal feature to obtain the projected motion noise.

[0072] In one embodiment, a high-dimensional random vector can be randomly sampled from a standard Gaussian distribution as initial motion noise. The initial motion noise can then be formally mapped to the same feature dimension as the multimodal features through a fully connected layer or a projection layer to obtain the projected motion noise.

[0073] S207. The projected motion noise and multimodal features are fused to obtain the target multimodal features, and the target multimodal features are input into the preset diffusion model.

[0074] In this embodiment, the projected motion noise can be directly added to the corresponding dimension value of the multimodal feature to obtain the target multimodal feature, and the target multimodal feature can be input into the preset diffusion model.

[0075] S208. Determine the noise estimation corresponding to the target multimodal features based on the noise prediction network in the preset diffusion model.

[0076] Among them, noise estimation can be understood as the estimated value of the noise component in the current projected motion noise.

[0077] In this embodiment, the multi-layer self-attention mechanism and feedforward network of the noise prediction network can be used to deeply encode the multi-modal features of the target, capturing the task constraint information, i.e., the multi-modal features, within the multi-modal features. The distribution pattern of the implicit initial action noise is identified to generate a noise estimate.

[0078] S209. Based on the denoising implicit model in the preset diffusion model, iterative denoising is performed according to noise estimation to obtain the robot action sequence.

[0079] In one embodiment, a preset diffusion step can be determined as the iteration step number. The diffusion steps are selected in descending order of noise intensity. Based on the DDIM implicit sampling formula, the denoised action vector is calculated using noise estimation to achieve accurate noise removal. The diffusion step is then updated, and the next iteration begins, continuing until the required number of iterations is completed to obtain the robot action sequence. In one embodiment, the robot action sequence may have issues with the target robot being unable to recognize it. The final denoised action vector can be reshaped to transform it into a structured robot action sequence. For example, a 320-dimensional vector can be transformed into a [16, 20] temporal action sequence, corresponding to 16 time steps and 20 action dimensions per time step, facilitating execution by the target robot.

[0080] S210. Determine the sequence segment to be executed in the robot action sequence, and control the target robot to perform the target task through the sequence segment to be executed.

[0081] The sequence segment to be executed can be understood as the sequence of actions performed by the target robot. In actual operation, the sequence segment to be executed can be the first half of the robot's action sequence or the entire robot action sequence.

[0082] In this embodiment, a sequence segment to be executed can be extracted from the robot's action sequence, and the target robot can be controlled to perform the target task according to the sequence segment to be executed.

[0083] In this embodiment of the invention, continuous frame images of the operation scene collected by a preset visual sensor are read as visual data. The visual data is input into a pre-trained visual encoder, and the classification label output vector output by the visual encoder is extracted as visual features. Text command data of the target task is identified, and the text command data is input into a pre-trained text encoder. The feature vector output by the text encoder is extracted as the text features of the text command data. Body sensor data of the target robot is collected, and the body sensor data is mapped to the target feature dimension through a linear projection layer to obtain numerical features. The mapped numerical features are used as body state features. Visual features, text features, and body state features are expanded and temporally aligned according to a preset time step as initial multimodal features, achieving multimodal feature fusion. A preset diffusion step is mapped into a high-dimensional vector as the diffusion step encoding through sinusoidal position encoding and a fully connected network. The initial multimodal features are embedded into the step-encoding mechanism. A random vector with the same dimension as the multimodal features is sampled from a Gaussian distribution as the initial motion noise. The initial motion noise is mapped to the same feature dimension as the multimodal features to obtain the projected motion noise. The projected motion noise is fused with the multimodal features to obtain the target multimodal features. The target multimodal features are input into a preset diffusion model. Based on the noise prediction network in the preset diffusion model, the noise estimate corresponding to the target multimodal features is determined. Based on the denoising diffusion implicit model in the preset diffusion model, iterative denoising is performed according to the noise estimate to obtain the robot motion sequence, thus achieving accurate determination of the robot motion sequence. By determining the sequence segments to be executed in the robot motion sequence, the target robot is controlled to perform the target task through the sequence segments to be executed, which facilitates real-time replanning of the robot motion sequence and improves the feasibility of the robot motion sequence.

[0084] Example 3

[0085] Figure 3 This is a flowchart of a robot control method according to Embodiment 3 of the present invention. This embodiment is a further explanation of a robot control method based on the above embodiments. Figure 3As shown, the method includes:

[0086] The robot collects images / videos of the operational scene and inputs them into a visual feature encoder. The user-issued task command text is input into a text feature encoder. The body perception is input into a body perception feature encoder. The encoder output features are fused and input into a motion diffusion transformer (preset diffusion model). At the same time, the input time vector encoding (diffusion step encoding) is input into the motion diffusion transformer. The output motion, i.e., the robot motion sequence, is obtained by DDIM denoising.

[0087] In one embodiment, taking the determination of robot action sequence using a multi-object diffusion strategy model as an example, a further explanation of a robot control method is provided. The multi-object diffusion strategy model includes a multi-modal feature extraction module and a diffusion strategy generation module.

[0088] The multimodal feature extraction module is used for visual feature extraction, text feature extraction, and robot ontology state determination. Specifically, it acquires continuous images from the robot's wrist camera and a fixed scene camera, and inputs these images into a pre-trained visual encoder (e.g., CLIP Vision Transformer). The output with CLS tags is used as visual features. This module can be fine-tuned during training to adapt to the robotic domain. User input of task descriptions or operation instructions is obtained, and linguistic features are extracted using a text encoder paired with the visual encoder (e.g., CLIP text encoder). The feature dimensions are adjusted through a projection layer. During training, the linguistic encoder remains frozen, and only the projection layer weights are adjusted. Ontology sensor data such as robot joint angles and velocities (i.e., "ontology perception") are collected as numerical features, i.e., the robot's ontology state.

[0089] The diffusion strategy generation module expands visual features, textual features, and robot ontology state by time steps, embedding the diffusion moments into the diffusion step encoding (using sinusoidal position encoding followed by two fully connected layers). A Diffusion Transformer (DiT) is used as the noise prediction network, with the model containing multiple DiT blocks (e.g., eight layers, with a self-attention dimension of 768). Adaptive Layer Normalization (AdaLN) is used to modulate the observed features at each diffusion step. During inference, a set of random action vectors is sampled from Gaussian noise, and denoising is performed k-steps using a Denoising Diffusion Implicit Model (DDIM), gradually converting the noise into an action sequence. At each step, the action is updated using the current diffusion step features and the noise estimate output by the noise prediction network, ensuring the final action sequence matches the observed visual, linguistic, and ontology information. The output action is a fixed-length action sequence (e.g., 16 time steps, each with 20 action dimensions, including the end-effector pose increments and commands for both robotic arms). During execution, the robot selects the first half of the time steps for control, and the remaining time steps are used for planning rolling, allowing for real-time replanning.

[0090] In one embodiment, the training and fine-tuning mechanism includes a pre-training phase and a fine-tuning phase. The pre-training phase uses a large-scale multi-task dataset containing both simulated and real robot demonstrations to pre-train the diffusion policy model. The dataset contains approximately 1700 hours of manipulation demonstrations, covering hundreds of different tasks and various robot configurations. During pre-training, a noise prediction network is trained using random noise weighting, enabling it to learn to recover realistic actions at different diffusion step sizes. Images are resized and color-dithered during training, and optimization employs a large-batch, low-learning-rate strategy. The fine-tuning phase selects a small amount of demonstration data to fine-tune the pre-trained model for a specific new task (i.e., a multi-object task). The fine-tuning process maintains the model architecture unchanged, updating parameters only for the demonstration data of that task. Typically, only a few hundred examples are needed to reach or exceed the baseline performance of a single-task model trained from scratch. The learning rate is low during fine-tuning, and the number of training steps is significantly reduced.

[0091] In one embodiment, a multi-task demonstration dataset (training dataset) can be constructed. Using a dual-arm robot platform, demonstration data for various tasks, including kitchen operations, assembly, and inventory management, is recorded via teleoperation. Multi-view videos from wrist and scene cameras, body sensor data, and corresponding control commands are collected. Images are uniformly scaled to a resolution suitable for neural network processing (e.g., 256×342) and randomly cropped to 224×224, with appropriate color perturbations added to enhance robustness. The text portion is provided by the operator or program, corresponding to the natural language description of the task (e.g., "move the cup from the shelf to the middle of the table"). During the pre-training phase, the multi-task data is mixed, and the robot's actions are conditionally generated and learned using a diffusion model. For each sample, a random number of diffusion steps are selected, adding noise to the real action to form a "noisy action." The model learns to predict the noise and recover the real action. Simultaneously, the parameters of the visual encoder, text projection layer, and diffusion transformer are optimized. The batch training size is large (e.g., 2560), the learning rate is small (e.g., 3e-4), and the visual encoder learning rate is one-tenth that of the main body. During the fine-tuning phase, for a specific new task, only a few dozen to a few hundred demos need to be collected. The pre-trained model is then fine-tuned using a small batch size (e.g., 320) and a low learning rate (e.g., 2e-5). The model structure remains unchanged, and the parameters are updated only based on the data for that specific task, thus quickly adapting to the new task.

[0092] The task-driven control flow involves the user describing the task they want the robot to perform via verbal commands during the runtime phase, such as "put the banana on the plate" or "install the bicycle brake disc." The system inputs the command text into a text encoder to extract linguistic features, simultaneously acquires the latest images from multiple cameras and extracts visual features, and reads the current state from the robot's sensors. These multimodal features are then input into a diffusion strategy generation module, which iteratively denoises and generates an action sequence. The robot executes the first half of this sequence and updates its perception information in real time to continue its planning. Because the model has been exposed to a large number of different tasks during the pre-training phase and has undergone minor optimizations for the target task during the fine-tuning phase, it can accurately identify target objects and their relationships in complex scenes and complete tasks according to the verbal commands.

[0093] In one embodiment, the multi-object task can be a kitchen scene task, such as "put the banana on a plate and the apple in a basket." The system identifies the two items (banana and apple) to be manipulated based on text instructions, locates their positions using visual features, and generates a sequence of actions including grasping, moving, and placing, thereby sequentially completing the handling of the two objects. Because the diffusion model has been exposed to similar operations during the training phase, it can quickly find the target and avoid interference in complex desktop scenes. Alternatively, it can be an assembly task, such as "installing a bicycle brake disc." The robot needs to complete multiple steps, including identifying the brake disc parts, aligning the mounting holes, and tightening the screws with appropriate tools. The language input clarifies the task objective and sequence, and the diffusion model uses image and ontology feedback to generate a coordinated sequence of actions for both arms, including grasping tools, locating the brake disc, aligning the holes, and tightening, to accurately complete the assembly task.

[0094] In this invention, by pre-training on large-scale multi-task data and then fine-tuning on the target task, the model can achieve or exceed the performance of a single-task baseline. In the experiments presented in the paper, the fine-tuned model significantly outperformed the single-task strategy on most simulation and real-world tasks, reducing the need for demonstration data, improving training efficiency, robustness, and generalization ability. Utilizing text instructions as conditions, the single diffusion strategy model can switch between different tasks without structural modifications; only the language description needs to be changed to control the execution of multi-object, multi-step tasks. The vision-text joint guidance model can accurately identify target objects and task sequences, enabling complex multi-object operations, multi-task sharing, and high-level instruction controllability. By combining fixed-length action outputs with a rolling execution strategy, it can provide global planning for long-sequential tasks and correct local deviations using real-time perception. This approach is beneficial for handling unknown interference in real-world environments, improving safety. Simultaneously, the model structure is unified, and the training and inference processes can be implemented on standard computing platforms without relying on specific hardware; the visual encoder and text encoder can be replaced with newer pre-trained models, improving system scalability.

[0095] Example 4

[0096] Figure 4 This is a schematic diagram of a robot control device according to Embodiment 4 of the present invention. Figure 4 As shown, the device includes: a feature acquisition module 41, a feature fusion module 42, a motion determination module 43, and a robot control module 44.

[0097] The feature acquisition module 41 is used to acquire visual data of the target robot's operation scene, text instruction data of the target task, and body state features of the target robot, and to determine the visual features of the visual data and the text features of the text instruction data, respectively.

[0098] The feature fusion module 42 is used to determine multimodal features based on visual features, text features, ontology state features and preset diffusion steps.

[0099] The motion determination module 43 is used to sample initial motion noise from a Gaussian distribution, determine target multimodal features based on the initial motion noise and multimodal features, and determine the robot motion sequence corresponding to the target multimodal features based on a preset diffusion model; wherein, the preset diffusion model is constructed by a noise prediction network and a denoising diffusion implicit model, and the preset noise prediction network is composed of a diffusion transformer.

[0100] The robot control module 44 is used to control the target robot to perform the target task based on the robot action sequence.

[0101] In this embodiment of the invention, a feature acquisition module acquires visual data of the target robot's operation scene, text command data of the target task, and the body state features of the target robot. The visual features of the visual data and the text features of the text command data are determined respectively. A feature fusion module determines multimodal features based on the visual features, text features, body state features, and a preset diffusion step. An action determination module samples initial action noise from a Gaussian distribution and determines the target multimodal features based on the initial action noise and multimodal features. Based on a preset diffusion model, a robot action sequence corresponding to the target multimodal features is determined. A robot control module controls the target robot to execute the target task based on the robot action sequence, realizing a joint visual and text guidance model. This model can accurately identify target objects and task sequences, achieve complex multi-object operations, and improve the accuracy of robot action sequences.

[0102] In one embodiment, the feature acquisition module 41 includes:

[0103] The visual feature determination unit is used to read continuous frame images of the operation scene collected by the preset visual sensor as visual data, input the visual data into the pre-trained visual encoder, and extract the classification label output vector output by the visual encoder as visual features.

[0104] The text feature determination unit is used to identify the text instruction data of the target task. It inputs the text instruction data into a pre-trained text encoder and extracts the feature vector output by the text encoder as the text features of the text instruction data. Among them, visual features and text features are the target feature dimensions.

[0105] The ontology feature determination unit is used to collect ontology sensor data of the target robot, map the ontology sensor data to the target feature dimension through a linear projection layer to obtain numerical features, and use the mapped numerical features as ontology state features.

[0106] In one embodiment, the feature fusion module 42 includes:

[0107] The initial feature determination unit is used to expand and align visual features, text features, and ontology state features according to a preset time step as initial multimodal features;

[0108] The feature fusion unit is used to map a preset diffusion step into a high-dimensional vector as a diffusion step encoding through sinusoidal position encoding and a fully connected network, and to embed the diffusion step encoding into the initial multimodal features as multimodal features.

[0109] In one embodiment, the action determination module 43 includes:

[0110] The noise determination unit is used to sample a random vector with the same dimension as the multimodal features from a Gaussian distribution as the initial action noise, and to map the initial action noise to the same feature dimension as the multimodal features to obtain the projected action noise;

[0111] The model input unit is used to fuse the projected motion noise with multimodal features to obtain the target multimodal features, and input the target multimodal features into the preset diffusion model;

[0112] The noise estimation determination unit is used to determine the noise estimate corresponding to the multimodal features of the target based on the noise prediction network in the preset diffusion model.

[0113] The motion determination unit is used to perform iterative denoising based on the noise estimation implicit model in the preset diffusion model to obtain the robot motion sequence.

[0114] In one embodiment, the robot control module 44 includes:

[0115] The robot control unit is used to determine the sequence segments to be executed in the robot's action sequence, and to control the target robot to perform the target task through the sequence segments to be executed.

[0116] In one embodiment, the robot control device further includes:

[0117] The training set determination module is used to obtain the training dataset. Each historical data sample in the training dataset includes continuous visual images, historical task text instructions, historical robot body states, and corresponding real action sequences.

[0118] The historical feature determination module is used to determine the visual features of continuous visual images as historical visual features and the text features of historical task text instructions as historical text features.

[0119] The noise addition module is used to randomly sample a diffusion step from the preset total number of diffusion steps, add target Gaussian noise to the real action sequence according to the noise intensity corresponding to the diffusion step, and generate a noisy action sequence.

[0120] The input determination module is used to encode the diffusion step using sinusoidal position coding combined with a two-layer fully connected network to obtain the diffusion step coding features. The historical visual features, historical text features, historical robot body state and diffusion step coding features are fused to obtain the training condition input.

[0121] The training module is used to input the training conditions and the noisy action sequence into the noise prediction network to generate prediction noise. The mean square error between the prediction noise and the target Gaussian noise is used as the loss function. The model parameters are updated through the gradient descent algorithm until the loss converges to the preset stable interval, thus completing the training.

[0122] The robot control device provided in the embodiments of the present invention can execute the robot control method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0123] Example 5

[0124] Figure 5 This is a schematic diagram of the structure of an electronic device implementing the robot control method of embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0125] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0126] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0127] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as robot control methods.

[0128] In some embodiments, the robot control method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the robot control method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to execute the robot control method by any other suitable means (e.g., by means of firmware).

[0129] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0130] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0131] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0132] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0133] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0134] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0135] In one embodiment, the present invention further includes a computer program product, which includes a computer program that, when executed by a processor, implements the robot control method of any embodiment of the present invention.

[0136] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0137] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0138] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A robot control method, characterized in that, include: The visual data of the target robot's operating scene, the text command data of the target task, and the body state characteristics of the target robot are acquired, and the visual features of the visual data and the text features of the text command data are determined respectively. Multimodal features are determined based on the visual features, text features, ontology state features, and a preset diffusion step; Initial motion noise is sampled from a Gaussian distribution. Target multimodal features are determined based on the initial motion noise and the multimodal features. The robot motion sequence corresponding to the target multimodal features is determined based on a preset diffusion model. The preset diffusion model is constructed from a noise prediction network and a denoising diffusion implicit model. The noise prediction network is composed of a diffusion transformer. Based on the robot's motion sequence, the target robot is controlled to perform the target task; The step of determining multimodal features based on the visual features, the text features, the ontology state features, and the preset diffusion step includes: Visual features, text features, and ontology state features are expanded and temporally aligned according to a preset time step to form the initial multimodal features; The preset diffusion step is mapped to a high-dimensional vector as the diffusion step encoding by sinusoidal position encoding and a fully connected network. The diffusion step encoding is then embedded into the initial multimodal features as multimodal features.

2. The method according to claim 1, characterized in that, The acquisition of visual data of the target robot's operating scene, text command data of the target task, and the body state characteristics of the target robot, and the determination of the visual features of the visual data and the text features of the text command data, respectively, includes: Read continuous frame images from the operation scenario collected by a preset visual sensor as visual data, input the visual data into a pre-trained visual encoder, and extract the classification label output vector output by the visual encoder as visual features. The text instruction data for the target task is identified, and the text instruction data is input into a pre-trained text encoder. The feature vector output by the text encoder is extracted as the text feature of the text instruction data; wherein, the visual feature and the text feature are the target feature dimensions. Collect the body sensor data of the target robot, map the body sensor data to the target feature dimension through a linear projection layer to obtain numerical features, and use the mapped numerical features as the body state features.

3. The method according to claim 1, characterized in that, The process of sampling initial motion noise from a Gaussian distribution, determining target multimodal features based on the initial motion noise and the multimodal features, and determining the robot motion sequence corresponding to the target multimodal features based on a preset diffusion model includes: A random vector with the same dimension as the multimodal features is sampled from a Gaussian distribution as initial motion noise. The initial motion noise is then mapped to the same feature dimension as the multimodal features to obtain the projected motion noise. The projected motion noise is fused with the multimodal features to obtain the target multimodal features, and the target multimodal features are input into a preset diffusion model. The noise estimate corresponding to the target multimodal features is determined based on the noise prediction network in the preset diffusion model; Based on the denoising diffusion implicit model in the preset diffusion model, the noise estimation is used to perform iterative denoising to obtain the robot action sequence.

4. The method according to claim 1, characterized in that, The step of controlling the target robot to perform the target task based on the robot action sequence includes: The robot action sequence is determined to be executed by a sequence segment to be executed, and the target robot is controlled to perform the target task by the sequence segment to be executed.

5. The method according to claim 1, characterized in that, The training process of the noise prediction network includes: Obtain a training dataset, in which each historical data sample includes continuous visual images, historical task text instructions, historical robot body states, and corresponding real action sequences. The visual features of the continuous visual images are determined as historical visual features, and the text features of the historical task text instructions are determined as historical text features. Randomly sample a diffusion step from the preset total number of diffusion steps, and add target Gaussian noise to the real action sequence according to the noise intensity corresponding to the diffusion step to generate a noisy action sequence; A sinusoidal position coding combined with a two-layer fully connected network is used to encode the diffusion step to obtain the diffusion step coding features. The historical visual features, historical text features, historical robot body state and diffusion step coding features are fused to obtain the training condition input. The training conditions are input into the noisy action sequence, and the noise prediction network generates predicted noise. The mean square error between the predicted noise and the target Gaussian noise is used as the loss function. The model parameters are updated by the gradient descent algorithm until the loss converges to a preset stable interval, thus completing the training.

6. A robot control device, characterized in that, include: The feature acquisition module is used to acquire visual data of the target robot's operation scene, text command data of the target task, and the body state features of the target robot, and to determine the visual features of the visual data and the text features of the text command data, respectively. The feature fusion module is used to determine multimodal features based on the visual features, the text features, the ontology state features, and a preset diffusion step; The motion determination module is used to sample initial motion noise from a Gaussian distribution, determine target multimodal features based on the initial motion noise and the multimodal features, and determine the robot motion sequence corresponding to the target multimodal features based on a preset diffusion model; wherein, the preset diffusion model is constructed by a noise prediction network and a denoising diffusion implicit model, and the noise prediction network is composed of a diffusion transformer; The robot control module is used to control the target robot to perform the target task based on the robot action sequence; The feature fusion module includes: The initial feature determination unit is used to expand and align visual features, text features, and ontology state features according to a preset time step as initial multimodal features; The feature fusion unit is used to map a preset diffusion step into a high-dimensional vector as a diffusion step encoding through sinusoidal position encoding and a fully connected network, and to embed the diffusion step encoding into the initial multimodal features as multimodal features.

7. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the robot control method according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the robot control method according to any one of claims 1-5.

9. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the robot control method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Data generation method and device, product, equipment and medium

    CN118695051A

  • Method and device for controlling robot, medium and electronic equipment

    CN119973990A