Double-arm robot operation control method and device, computer equipment and storage medium

Through the diffusion model, candidate actions are generated and target actions are selected. Combined with voice and image processing technology, the adaptability and real-time problems of operation control of two-arm robots are solved, the accuracy and coordination of robot operations are improved, and applications in the fields of medical health and logistics are supported.

CN120552052APending Publication Date: 2025-08-29PING AN TECH (SHENZHEN) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510686030.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The existing two-arm robot operation control technology has obvious shortcomings in action mode adaptability, data utilization efficiency and real-time control capabilities, and it is difficult to widely use in medical and health fields.

Method used

The diffusion model is used to combine multimodal information to generate candidate actions, and the target actions are selected through comprehensive scoring, integrating speech recognition, image processing and robot sensors to obtain environmental information, achieving efficient and real-time operation control.

Benefits of technology

It improves the accuracy, coordination and adaptability of the operation control of two-arm robots, and supports automation and intelligent applications in the fields of medical health and logistics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120552052A_ABST
    Figure CN120552052A_ABST
Patent Text Reader

Abstract

The invention discloses a double-arm robot operation control method and device, computer equipment and a storage medium, and the method comprises the steps: receiving an operation instruction, and obtaining environment information; inputting the operation instruction and the environment information into a diffusion model to generate a plurality of groups of candidate actions; sorting the multiple groups of candidate actions from high to low according to the comprehensive scores, wherein the group of candidate actions with the highest comprehensive score is the target action; and executing the target action. By implementing the method, the accuracy, coordination, adaptability, real-time control capability and data utilization efficiency of operation control of the double-arm robot are improved, and the technical scheme can be applied to the field of medical health.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robot control technology, and more specifically to a dual-arm robot operation control method, device, computer equipment and storage medium. Background Art

[0002] In fields like healthcare, dual-arm robotic control technology has become a key driver of automation and intelligent advancements. This technology aims to improve efficiency, reduce labor costs, and ensure operational safety by enabling two highly coordinated robotic arms to precisely perform complex tasks such as grasping, handling, and fine manipulation. However, despite significant progress in dual-arm robotic control, numerous challenges remain.

[0003] The core of dual-arm robotic operation lies in achieving precise coordination and cooperation between the two arms. However, in practice, even a simple grasping action can be implemented in multiple ways. This diversity of motion patterns poses a significant challenge to existing technologies. Traditional methods often rely on preset rules or fixed motion plans, which struggle to fully capture and adapt to this diversity. This leads to inaccurate motion predictions and uncoordinated coordination between the two arms, which in turn affects the efficiency and success rate of task execution.

[0004] In the field of dual-arm robot manipulation and control, pre-training with multiple robot datasets is an effective way to improve model generalization. However, due to significant differences in structure, motion space, and functional characteristics among different robots, directly pre-training with this heterogeneous data often faces numerous difficulties. Existing methods either discard large amounts of valuable data due to excessive data disparity or only extract a subset of common features for training. This not only wastes data diversity but also limits the effective transfer of physical knowledge, compromising the model's adaptability and robustness.

[0005] In practical applications, dual-arm robots often need to operate in dynamic environments, such as those with fluctuating lighting and random object positions. However, existing methods often suffer from instability in the face of these uncertainties, making it difficult to guarantee accurate and reliable task execution. Furthermore, the high computational demands of complex models pose a significant challenge to the robot's real-time control capabilities. This is particularly true on resource-constrained robotic platforms, where achieving efficient, real-time operational control is a major challenge.

[0006] In summary, current dual-arm robot manipulation and control technologies have significant shortcomings in terms of motion mode adaptability, data utilization efficiency, and real-time control capabilities, severely restricting their widespread application and further development in fields such as healthcare. Therefore, developing a dual-arm robot manipulation and control scheme that can efficiently capture motion diversity, fully utilize data resources, adapt to dynamic environments, and achieve real-time control has important theoretical and practical significance. Summary of the Invention

[0007] The purpose of the present invention is to overcome the defects of the prior art and provide a dual-arm robot operation control method, device, computer equipment and storage medium.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] The dual-arm robot operation control method includes:

[0010] Receive operation instructions and obtain environmental information;

[0011] Inputting operation instructions and environmental information into the diffusion model to generate multiple sets of candidate actions;

[0012] Sort multiple groups of candidate actions from high to low according to their comprehensive scores, and the group of candidate actions with the highest comprehensive score is the target action;

[0013] Perform the target action.

[0014] A further technical solution is: the receiving of the operation instruction and obtaining of the environmental information includes:

[0015] Receive operation instructions from users through voice recognition or text input interface;

[0016] The robot uses the camera or external vision system to obtain image information of the current environment to obtain environmental information.

[0017] A further technical solution is: the operation instructions and environmental information are input into the diffusion model to generate multiple groups of candidate actions, including:

[0018] The received operation instructions are converted into vector representations that can be understood by the robot through natural language processing technology to obtain instruction information;

[0019] Using convolutional neural network computer vision technology, the captured environmental information is encoded and task-related visual features are extracted to obtain feature information;

[0020] Obtain and encode the state information perceived by the robot body, and then convert it into a vector representation related to the task to obtain perception information;

[0021] The instruction information, feature information and perception information are input into the diffusion model to generate multiple sets of candidate actions.

[0022] A further technical solution is: the diffusion model is obtained by big data training, and the specific training steps are as follows:

[0023] Collect demonstration data, which includes the robot's proprioception information, action sequence, and corresponding control frequency when completing the task;

[0024] Preprocess the demonstration data to obtain multimodal input encoding information;

[0025] At the beginning of training, random noise is injected into the correct action sequence to generate a set of random action sequences;

[0026] De-noising the random action sequence to obtain denoised actions;

[0027] Based on the current action state, multimodal input encoding information and denoised action, predict the next state that is closer to the real action to obtain the predicted action sequence;

[0028] In each iteration, the loss between the predicted action sequence and the true action sequence is calculated to obtain the loss value;

[0029] According to the loss value, the parameters of the core network are updated through the back propagation algorithm to minimize the loss;

[0030] Repeat the above steps until the model reaches the preset performance index to obtain the diffusion model.

[0031] A further technical solution is: the plurality of candidate action groups are sorted from high to low according to the comprehensive scores, and the group of candidate actions with the highest comprehensive score is the target action, including:

[0032] Based on the task requirements, set the criteria for evaluating candidate actions;

[0033] For each set of candidate actions, calculate the corresponding evaluation index using the set evaluation criteria;

[0034] Based on the calculated evaluation indicators, a comprehensive score is calculated for each group of candidate actions, and all candidate actions are sorted from high to low according to the comprehensive score;

[0035] According to the ranking results, a group of candidate actions with the highest comprehensive scores are selected as target actions.

[0036] A further technical solution is as follows: the executing target action includes:

[0037] Parse the target action from a vector or encoding form into a specific action sequence;

[0038] Send action instructions to the robot one by one according to the action sequence;

[0039] The robot performs corresponding actions according to the action instructions.

[0040] A further technical solution thereof is: after executing the target action, the method further comprises:

[0041] Determine whether the operation instruction is completed; if the operation instruction is not completed, return to execute the received operation instruction and obtain environmental information.

[0042] The present invention also provides a dual-arm robot operation control device, comprising:

[0043] A receiving and obtaining unit, configured to receive an operation instruction and obtain environmental information;

[0044] An input generation unit, configured to input operation instructions and environmental information into the diffusion model to generate multiple sets of candidate actions;

[0045] A sorting unit is used to sort multiple groups of candidate actions from high to low according to their comprehensive scores. The group of candidate actions with the highest comprehensive score is the target action.

[0046] Execution unit, used to execute the target action.

[0047] The present invention further provides a computer device, comprising a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the above method when executing the computer program.

[0048] The present invention also provides a storage medium, wherein the storage medium stores a computer program, and the computer program implements the above method when executed by a processor.

[0049] Compared with the existing technology, the beneficial effects of the present invention are: by integrating a series of steps such as receiving operation instructions, obtaining environmental information, generating multiple groups of candidate actions using diffusion models, screening target actions and executing them, the accuracy, coordination and adaptability of the dual-arm robot's operation control, real-time control capabilities and data utilization efficiency are improved, providing strong technical support for automation and intelligent applications in the fields of medical health, logistics, etc.

[0050] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0052] Figure 1 A schematic diagram of an application scenario of the dual-arm robot operation control method provided by an embodiment of the present invention;

[0053] Figure 2 A schematic flow chart of a dual-arm robot operation control method provided by an embodiment of the present invention;

[0054] Figure 3 A schematic block diagram of a dual-arm robot operation control device provided by an embodiment of the present invention;

[0055] Figure 4 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0057] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0058] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0059] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0060] See also Figure 1 and Figure 2 , Figure 1Schematic diagram of an application scenario of the dual-arm robot operation control method provided in an embodiment of the present invention. Figure 2 This is a schematic flow chart of a dual-arm robot operation and control method provided by an embodiment of the present invention. This dual-arm robot operation and control method is applied to a server that interacts with a terminal for data exchange. By integrating a series of steps, including receiving operation instructions, acquiring environmental information, generating multiple sets of candidate actions using a diffusion model, screening and executing target actions, etc., it improves the accuracy, coordination, and adaptability of the dual-arm robot's operation and control, as well as its real-time control capabilities and data utilization efficiency. This provides strong technical support for automated and intelligent applications in fields such as healthcare and logistics.

[0061] Figure 2 FIG. 1 is a flow chart of a dual-arm robot operation control method according to an embodiment of the present invention. Figure 2 As shown, the method includes the following steps S110 to S140.

[0062] S110, receiving an operation instruction and obtaining environmental information;

[0063] Specifically, the robot receives operational instructions from the user or system through voice recognition, natural language processing, or a graphical user interface (GUI). These instructions may include specific task descriptions, such as "grab object A and place it at location B." The robot uses sensors onboard (such as cameras, lidar, and force sensors) to collect environmental information. This information includes the object's location, shape, color, and the layout of the surrounding environment. This environmental information is presented in the form of images, point clouds, depth maps, etc., and is optimized through preprocessing steps (such as denoising and normalization) for subsequent processing.

[0064] In other words, receiving operational instructions through multiple channels allows the robot to more accurately understand user intent and adapt to different application scenarios. Using multiple sensors to acquire environmental information provides the robot with comprehensive environmental perception capabilities, facilitating subsequent action generation and decision-making.

[0065] In one embodiment, receiving the operation instruction and obtaining the environment information includes:

[0066] Receive operation instructions from users through voice recognition or text input interface;

[0067] Specifically, a high-sensitivity microphone array is integrated into the robot's operating interface or body to support multi-directional and long-range voice signal acquisition, while filtering out ambient noise using noise reduction algorithms (such as spectral subtraction and deep learning noise reduction models). A cloud-based or locally deployed speech recognition engine (such as an ASR model) is used to convert real-time speech streams into text commands. Multilingual recognition and dialect-adaptive optimization are supported, and accuracy is improved by continuously updating acoustic and language models. Based on natural language processing (NLP) technology, semantic analysis is performed on the converted text to extract key information such as task type (such as grasping, carrying), target object name, and operation parameters (such as position and force), and then convert them into structured instructions that the robot can execute. Alternatively, a graphical user interface (GUI) or command-line interface (CLI) can be developed, providing controls such as buttons, drop-down menus, and text boxes to support user input of operation commands via a touchscreen, keyboard, or external device. Standardized command formats (such as JSON and XML) are defined, requiring users to enter task parameters according to predefined templates to ensure accurate command parsing. Provide instant verification prompts (such as syntax error highlighting and parameter range prompts) during user input, and generate an execution confirmation dialog box after the instruction is confirmed to avoid incorrect operations.

[0068] In other words, it supports dual-modal input of voice and text to meet the operational requirements in different scenarios (such as voice commands in noisy environments and text parameter input in precise tasks), lowering the user's operation threshold. In addition, through NLP parsing and standardized instruction formats, the impact of natural language ambiguity on task execution is reduced, ensuring that the robot accurately understands the user's intentions. In addition, after the voice recognition and text parsing processes are optimized, the instruction processing delay is less than 200ms, and dynamic task adjustment is supported (such as terminating the current task and switching operation modes).

[0069] The robot uses the camera or external vision system to obtain image information of the current environment to obtain environmental information.

[0070] Specifically, a high-resolution RGB-D camera (such as Intel RealSense or Azure Kinect) is installed on the robot's head or end-effector to support the simultaneous acquisition of color images, depth maps, and point cloud data. For complex scenes, an external fixed vision system (such as an industrial camera or a panoramic camera) can be connected for auxiliary assistance. Multi-camera calibration and spatial registration algorithms are used to unify image data from different perspectives into a common coordinate system, eliminating geometric distortion caused by perspective differences and generating a globally consistent environmental model. The captured raw images are then subjected to denoising (such as Gaussian filtering and non-local mean denoising), enhancement (such as histogram equalization and contrast stretching), and distortion correction (such as radial distortion correction) to improve image quality. Deep learning models (such as YOLO and Mask R-CNN) are used for object detection and semantic segmentation to identify the category, location, and bounding box of objects in the environment. Traditional computer vision techniques (such as SIFT and ORB) are combined to extract key point features for subsequent pose estimation. Based on the SLAM algorithm, image features and depth information are integrated to construct a three-dimensional environment map, and the obstacle distribution and passable areas are updated in real time to provide a basis for path planning.

[0071] In other words, multi-sensor fusion and 3D modeling technologies enable the robot to achieve millimeter-level spatial resolution of its environment, supporting safe obstacle avoidance in complex scenarios. Furthermore, a real-time environmental update mechanism (such as a point cloud refresh rate of 10 frames per second) enables the robot to quickly respond to scene changes (such as moving objects and temporary obstacles), improving the success rate of mission execution. Furthermore, feature compression and region of interest (ROI) extraction reduce the amount of invalid data processing, keeping the environmental perception module's CPU usage low on embedded devices and ensuring real-time performance.

[0072] S120, inputting the operation instructions and environmental information into the diffusion model to generate multiple groups of candidate actions;

[0073] Specifically, the operation instructions and environmental information are encoded into vector representations suitable for processing by the diffusion model. This includes converting language instructions into word embedding vectors, extracting feature vectors from image information through a convolutional neural network (CNN), and encoding proprioceptive information (such as joint angles, speeds, etc.). The diffusion model is trained using a large amount of demonstration data to enable it to learn the ability to recover correct action sequences from random noise. In the generation phase, the encoded operation instructions and environmental information are input as conditions into the diffusion model, and multiple sets of candidate action sequences are generated through a step-by-step denoising process.

[0074] In other words, the diffusion model can capture the complex distribution of actions in dual-arm manipulation and generate a diverse set of candidate action sequences, improving the accuracy and flexibility of action prediction. By combining operational instructions with environmental information, the diffusion model can generate candidate actions that better match actual scenarios and task requirements, enhancing the robot's environmental adaptability and task execution capabilities.

[0075] In one embodiment, inputting the operation instructions and environmental information into the diffusion model to generate multiple sets of candidate actions includes:

[0076] The received operation instructions are converted into vector representations that can be understood by the robot through natural language processing technology to obtain instruction information;

[0077] Specifically, the received natural language operation instructions are cleaned, irrelevant characters and punctuation marks are removed, and word segmentation is performed to decompose the instructions into independent vocabulary units. A pre-trained word vector model (such as Word2Vec, GloVe, or BERT, etc.) is used to convert the segmented words into fixed-dimensional vector representations. These vectors capture the semantic relationship between words, so that words with similar meanings are closer in the vector space. The word vector sequence is input into a recurrent neural network (RNN) or its variants (such as LSTM, GRU) to capture the temporal dependencies in the instructions. Through the hidden state of the RNN, the entire instruction sequence is encoded into a fixed-length vector representation, namely the instruction information vector.

[0078] Using convolutional neural network computer vision technology, the captured environmental information is encoded and task-related visual features are extracted to obtain feature information;

[0079] Specifically, the captured environmental image is normalized, cropped, and scaled to meet the input requirements of the convolutional neural network. The preprocessed image is input into a pre-trained convolutional neural network (such as ResNet and VGG). Through multiple layers of convolution and pooling operations, low-level to high-level features of the image are gradually extracted. These features, such as edges, textures, and shapes, are crucial for understanding environmental information. Depending on the specific task requirements, task-related feature vectors are extracted from the intermediate layers or fully connected layers of the convolutional neural network. These feature vectors can reflect key information in the environment, such as the position, shape, and color of objects, i.e., feature information.

[0080] Obtain and encode the state information perceived by the robot body, and then convert it into a vector representation related to the task to obtain perception information;

[0081] Specifically, onboard sensors (such as encoders, gyroscopes, and accelerometers) collect real-time state information about the robot, including joint angles, velocities, accelerations, and positions. This collected state information undergoes normalization and filtering to eliminate noise and outliers. This processed state information is then fed into a fully connected neural network or recurrent neural network for encoding and feature extraction, yielding a task-specific state vector representation, or perception information.

[0082] The instruction information, feature information and perception information are input into the diffusion model to generate multiple sets of candidate actions.

[0083] Specifically, the instruction information vector, feature information vector, and perception information vector are concatenated or fused through an attention mechanism to produce a comprehensive multimodal information vector. This vector contains comprehensive information about the operation instruction, environmental information, and robot state. The diffusion model is trained using a large dataset containing multimodal information and corresponding action sequences. During training, the diffusion model learns to gradually recover action sequences that match the multimodal information from random noise. During the inference phase, the new multimodal information vector is input into the trained diffusion model. Through a gradual denoising process, the diffusion model generates multiple sets of candidate action sequences. These candidate action sequences both meet the requirements of the operation instruction and are adapted to the current environmental information and robot state.

[0084] Specifically, by combining multimodal information from operational instructions, environmental information, and robot state, the diffusion model is able to generate more accurate and diverse candidate action sequences that not only meet task requirements but also adapt to varying environmental changes and robot states. Furthermore, the diffusion model can learn the influence of environmental information and robot state on action generation, enabling it to generate effective candidate actions even in complex and dynamic environments. This enhances the robot's environmental adaptability and task execution capabilities. Furthermore, through the diffusion model's gradual denoising process, it can generate smooth and continuous action sequences, avoiding the abrupt or discontinuous motions that can occur in traditional methods. This optimizes the action generation process and improves the robot's operational stability and efficiency. Furthermore, the diffusion model boasts high computational efficiency when generating candidate actions, enabling real-time control and decision-making. This allows the robot to respond to operational instructions and environmental changes in a short period of time, improving its real-time performance and flexibility.

[0085] In one embodiment, the diffusion model is obtained by big data training, and the specific training steps are as follows:

[0086] Collect demonstration data, which includes the robot's proprioception information, action sequence, and corresponding control frequency when completing the task;

[0087] Specifically, high-precision sensors (such as joint encoders, force sensors, IMUs, etc.) are installed on the robot body to collect real-time proprioceptive information of the robot when completing tasks, including joint angles, speeds, torques, postures, etc. Develop or use existing robot control software to record the robot's action sequence during task execution, including joint target positions and speed instructions at each time step. While recording the action sequence, mark the control frequency corresponding to each action instruction, that is, the cycle or time interval at which the instruction is issued. Store the collected demonstration data in a database or file system, establish data indexes and tags, and facilitate subsequent data retrieval and processing.

[0088] Preprocess the demonstration data to obtain multimodal input encoding information;

[0089] Specifically, the collected demonstration data is cleaned to remove outliers, noise, and duplicate data to ensure data accuracy and consistency. Key features are extracted from proprioceptive information, such as the rate of change of joint angles, acceleration, peak torque, etc., as a description of the robot state. At the same time, the action sequence is segmented or keyframe extracted to reduce the amount of data and retain key information. Data of different modalities such as proprioceptive information, action sequence, and control frequency are encoded and converted into vector representations suitable for diffusion model processing. For example, the action sequence is encoded using a time series encoding method, and the robot state and control parameters are represented using feature vectors.

[0090] At the beginning of training, random noise is injected into the correct action sequence to generate a set of random action sequences;

[0091] Specifically, a random number generator is used to generate random noise that conforms to a specific distribution (such as a Gaussian distribution). The amplitude and dimension of the noise are adjusted according to the characteristics of the action sequence. The generated random noise is injected into the correct action sequence according to a certain ratio or strategy, resulting in a set of random action sequences. Noise injection can be performed point by point, piecewise replacement, or global perturbation.

[0092] De-noising the random action sequence to obtain denoised actions;

[0093] Specifically, choose a denoising algorithm suitable for the diffusion model, such as a denoising autoencoder (DAE) based on deep learning, a denoising version of a variational autoencoder (VAE), or a dedicated diffusion model denoising network. If using a deep learning-based denoising algorithm, first train the denoising network with a large number of random action sequences and corresponding correct action sequences to learn the mapping from random actions to correct actions. The random action sequence, after injecting noise, is then fed into the trained denoising network, which then outputs the denoised action sequence.

[0094] Based on the current action state, multimodal input encoding information and denoised action, predict the next state that is closer to the real action to obtain the predicted action sequence;

[0095] Specifically, a state prediction model is constructed. This model receives the current action state, multimodal input encoding information, and denoised action as input, and outputs the predicted action state for the next time step. The state prediction model is trained using historical data to learn the transition patterns between action states and the influence of multimodal information on action states. In each iteration, the current action state, multimodal input encoding information, and denoised action are input into the state prediction model to obtain the next state prediction value that is closer to the actual action.

[0096] In each iteration, the loss between the predicted action sequence and the true action sequence is calculated to obtain the loss value;

[0097] Specifically, a loss function suitable for action sequence prediction is selected, such as mean squared error (MSE), mean absolute error (MAE), or a custom loss function, to measure the difference between the predicted action sequence and the true action sequence. In each iteration, the predicted action sequence is compared with the true action sequence, and the loss value is calculated using the selected loss function.

[0098] According to the loss value, the parameters of the core network are updated through the back propagation algorithm to minimize the loss;

[0099] Specifically, based on the calculated loss value, the backpropagation algorithm is used to calculate the gradient of the loss function with respect to the core network parameters. An optimization algorithm (such as stochastic gradient descent SGD, Adam, etc.) is used to update the core network parameters based on the gradient information to minimize the loss value.

[0100] Repeat the above steps until the model reaches the preset performance index to obtain the diffusion model.

[0101] Specifically, the diffusion model is iteratively trained by repeating steps such as data preprocessing, noise injection, denoising, state prediction, loss calculation, and parameter updating. During training, the model's performance is regularly evaluated using a validation set, such as by calculating metrics like the similarity and accuracy between the predicted and true action sequences. Training is terminated when the model's performance on the validation set reaches a preset performance metric (e.g., a loss value below a certain threshold, an accuracy rate above a certain percentage, etc.), resulting in the final diffusion model.

[0102] In other words, the diffusion model trained on big data can learn the complex patterns of robot motion generation, generating motion sequences that are closer to real-world motions, improving the accuracy of motion generation. Furthermore, the process of injecting random noise and denoising makes the model more robust to noise and interference, enabling it to generate reasonable motion sequences even with inaccurate or noisy input. Furthermore, training with a large amount of demonstration data enables the model to learn motion generation patterns across different tasks and environments, enhancing its generalization capabilities. The introduction of multimodal input encoding allows the model to comprehensively consider multiple factors, including the robot's proprioceptive information, motion sequence, and control frequency, further improving its generalization performance. The diffusion model boasts high computational efficiency when generating motion sequences and supports real-time control and decision-making. This enables the robot to respond to operational commands and environmental changes in a short period of time, improving its real-time performance and flexibility.

[0103] S130, sorting the multiple candidate action groups from high to low according to the comprehensive scores, and the candidate action group with the highest comprehensive score is the target action;

[0104] Specifically, the generated candidate action sequences are evaluated based on preset evaluation metrics (such as smoothness, efficiency, and safety). This evaluation process may involve simulation verification of the action sequences and analysis of their compatibility with actual scenarios. By comparing the evaluation scores of different candidate action sequences, the target action is selected.

[0105] In other words, by screening target actions, we ensure that the robot's actions meet the task requirements while also being smooth, efficient, and safe. This evaluation and screening process provides the robot with a scientific basis for decision-making, helping to optimize the decision-making process and increase the success rate of task execution.

[0106] In one embodiment, the plurality of candidate action groups are sorted from high to low according to the comprehensive scores, and the group of candidate actions with the highest comprehensive score is the target action, including:

[0107] Based on the task requirements, set the criteria for evaluating candidate actions;

[0108] Specifically, first, we need to deeply analyze the specific requirements of the current task, including the task objectives, operating environment, robot capability limitations, etc. For example, in a dual-arm robot grasping task, the task objective may be to grasp the target object accurately, quickly and stably; the operating environment may include the position, shape, material, etc. of the object; and the robot capability limitations involve the robot's range of motion, load capacity, accuracy, etc. Based on the task requirement analysis, a set of criteria for evaluating candidate actions is formulated. These criteria may include the accuracy of the action (such as the accuracy of the grasping position), efficiency (such as the time required to complete the action), stability (such as vibration or deviation during the action), safety (such as avoiding collision or damage to objects), etc. Each criterion can be weighted according to the importance and priority of the task.

[0109] For each set of candidate actions, calculate the corresponding evaluation index using the set evaluation criteria;

[0110] Specifically, for each evaluation criterion, a specific calculation method or metric is determined. For example, action accuracy can be measured by calculating the Euclidean distance between the grasping position and the target position; efficiency can be measured by recording the time required to complete the action; stability can be measured by monitoring acceleration or displacement changes during the action; and safety can be measured by checking whether a collision occurs or the action exceeds the safe range. For each set of candidate actions, a specific calculation method or metric is used to calculate the specific value for each action under each evaluation criterion.

[0111] Based on the calculated evaluation indicators, a comprehensive score is calculated for each group of candidate actions, and all candidate actions are sorted from high to low according to the comprehensive score;

[0112] Specifically, a comprehensive score is calculated for each set of candidate actions based on the weights of the various evaluation criteria and the calculated evaluation indicators. This can be achieved through methods such as weighted summation and fuzzy comprehensive evaluation. For example, if the weights for accuracy, efficiency, stability, and safety are 0.4, 0.3, 0.2, and 0.1, respectively, then the comprehensive score is the sum of the products of these four indicators and their corresponding weights. All candidate actions are sorted from high to low according to their comprehensive scores to obtain a candidate action sequence.

[0113] According to the ranking results, a group of candidate actions with the highest comprehensive scores are selected as target actions.

[0114] Specifically, a set of candidate actions with the highest comprehensive scores is selected from the sorted candidate action sequence as the target action. This set of actions performs best in meeting the task requirements and can maximize the accuracy, efficiency, stability, and safety of task completion.

[0115] In other words, by setting clear evaluation criteria and calculating specific evaluation metrics, the performance of each set of candidate actions can be more accurately measured. A comprehensive scoring and ranking mechanism makes the selection of target actions more objective and efficient. Furthermore, the evaluation criteria can be adjusted and optimized based on different task requirements, enabling the robot to adapt to a variety of complex and changing environments. At the same time, by screening target actions, the robot can more flexibly respond to operational requirements in different situations. Furthermore, by considering factors such as stability and safety in the evaluation criteria, the selected target actions can be executed more stably and reliably, avoiding safety issues such as collisions or damage to objects. Furthermore, this technical feature provides strong support for the intelligent and autonomous development of robotics technology. By automatically screening target actions, robots can complete complex tasks more autonomously, reducing the need for human intervention and decision-making.

[0116] S140: Execute the target action.

[0117] Specifically, the selected target action sequences are parsed into specific control instructions, such as joint angle changes and speed control. These instructions are expressed in a form that the robot can understand and execute. These control instructions are then sent to the robot's actuators (such as motors and drivers), driving the robot to perform the target action. During execution, the robot's state and environmental changes are monitored in real time, and the action is adjusted or replanned as necessary.

[0118] In other words, by executing the target action, the robot can complete the task efficiently and accurately, improving work efficiency and task execution quality. Real-time monitoring and adjustment of the action during execution enables the robot to adapt to environmental changes and enhances its real-time control capabilities.

[0119] In one embodiment, performing the target action includes:

[0120] Parse the target action from a vector or encoding form into a specific action sequence;

[0121] Specifically, an action decoder is designed that can receive the target action vector or encoding output by the diffusion model or transformer architecture. The decoder contains mapping rules or neural network models to convert the action in vector or encoding form into a specific action sequence. For example, for a 128-dimensional unified action space vector, the decoder can convert the value of each dimension into specific action parameters such as the angle, speed, and gripper opening and closing degree of the robot joint based on predefined dimensional division and physical meaning mapping. When the robot performs the task, the target action vector or encoding is input into the action decoder in real time for immediate parsing. During the parsing process, factors such as the current state of the robot and the environmental context may also need to be considered to ensure that the generated action sequence meets the requirements of the target action and adapts to the actual situation.

[0122] Send action instructions to the robot one by one according to the action sequence;

[0123] Specifically, based on the action sequence obtained by parsing, the corresponding robot action instructions are generated. These instructions usually include parameters such as joint angle, speed, acceleration, and duration. The instructions are encapsulated into a format suitable for the robot control system to receive, such as a specific communication protocol or data packet. A stable instruction sending mechanism is established to ensure that the action instructions can be sent to the robot control system accurately and in a timely manner. Serial communication, network communication (such as TCP / IP, UDP), etc. can be used for instruction transmission. The specific choice depends on the interface and communication protocol of the robot control system. Send action instructions to the robot one by one in the order of the action sequence. Before sending the next instruction, you can wait for the robot to complete the execution of the current instruction, or set a certain time interval to ensure the continuity and stability of the action.

[0124] The robot performs corresponding actions according to the action instructions.

[0125] Specifically, the robot control system receives and parses motion instructions from the host computer. During the parsing process, the integrity and correctness of the instructions must be verified to ensure that the robot can correctly understand and execute the instructions. Based on the instruction parameters obtained through parsing, the robot control system controls the robot's various joints and actuators to perform the corresponding actions. During the execution process, the robot control system also needs to monitor the robot's status and environmental information in real time, such as joint angles, torque, and position, to ensure the safety and accuracy of the action. After the robot performs an action, it can send feedback information to the control system, such as the completion status of the action and the current position. Based on this feedback information, the control system can adjust or optimize subsequent motion instructions to improve the robot's performance and adaptability.

[0126] Specifically, by parsing the target action from a vector or encoding into a specific action sequence and sending the action instructions to the robot one by one, the robot can ensure that the action is executed according to the predetermined optimal path and parameters, thereby improving the accuracy and efficiency of action execution. Furthermore, the action decoder can convert the target action vector or encoding into an action sequence that is appropriate for the current situation, depending on the task requirements and robot state. This flexibility enables the robot to adapt to various complex and changing environments and perform a variety of tasks. Furthermore, by sending action instructions to the robot one by one and monitoring the robot's state and environmental information in real time, the robot's stability and reliability during action execution are ensured. Furthermore, the feedback and adjustment mechanism enables the robot to respond promptly to abnormal situations, avoiding task failure or damage. Furthermore, this technical feature provides strong support for the development of intelligent and autonomous robotics technology. By automatically parsing the target action and generating action instructions, the robot can complete complex tasks more autonomously, reducing the need for human intervention and decision-making. It also facilitates the integration and collaboration of robots with other intelligent systems.

[0127] In one embodiment, after executing the target action, the method further includes:

[0128] Determine whether the operation instruction is completed; if the operation instruction is not completed, return to execute the received operation instruction and obtain environmental information; if the operation instruction is completed, end the operation.

[0129] Specifically, before the robot starts to execute the operation instruction, the conditions for completing the operation instruction are pre-set. These conditions may include, but are not limited to: the robot reaches the specified position, completes a specific action sequence, and achieves the expected task effect (such as successfully grasping an object, completing an assembly task, etc.). During the process of the robot executing the operation instruction, the robot's status and environmental information are monitored in real time through sensors, cameras and other equipment. This information may include the robot's position, posture, action execution status, the status of the target object, etc. Based on the preset instruction completion conditions and the real-time monitored status information, it is determined whether the operation instruction has been completed. For example, the judgment can be made by comparing the difference between the current position of the robot and the target position, checking whether the action sequence has been fully executed, and evaluating whether the task effect meets expectations.

[0130] A loop mechanism is established whereby the robot control system automatically returns to the step where it received the instruction if it determines that the instruction has not been completed. Before or after each return to receive the instruction, the robot reacquires environmental information. This can be achieved by refreshing sensors or recapturing camera images to ensure that the robot has the latest environmental status. Based on the updated environmental information and the original instruction (which may have been adjusted for the new environment), the robot replans and executes the action sequence until the instruction is completed.

[0131] When the robot control system determines that the operation instruction has been completed, it confirms that the task has been successfully completed. It releases resources related to the current task (such as memory and sensor usage) and resets the robot's state to its initial state or to a state ready for the next task. It records relevant information about the task completion (such as completion time and task performance evaluation) and may send a task completion feedback signal to the user or higher-level system.

[0132] In other words, a looping mechanism continuously checks the completion status of commands and, if incomplete, reacquires environmental information and replans actions. This ensures the robot can continuously and stably execute tasks until they are complete, significantly improving the integrity and reliability of task execution. Furthermore, real-time acquisition of environmental information and the adjustment of action sequences based on this information enable the robot to adapt to environmental changes, such as subtle shifts in the target object's position or the presence of obstacles. This adaptability enhances the robot's robustness, enabling it to operate stably in complex and changing environments. Furthermore, by promptly determining command completion status and terminating operations, unnecessary resource waste and ineffective execution are avoided. Furthermore, the environmental information updates and action replanning within the looping mechanism help optimize task execution paths and improve efficiency. Furthermore, users can obtain timely information about task execution status through feedback signals, eliminating the need to continuously monitor the robot's status. This enhanced automation and intelligence of the system reduces the need for human intervention, improving user experience and system usability.

[0133] In one embodiment, the operation and control of the dual-arm robot can also be achieved in the following manner:

[0134] The diffusion model first injects noise into the latent action space to simulate the uncertainty of actions. It then generates a set of candidate action sequences through a stepwise denoising process. This process is similar to sketching, with the action plan gradually refined through multiple iterations. During each denoising step, the model generates multiple possible action candidates based on the current action state and the latent space distribution, forming a candidate action set. The model encodes image information (such as object position and shape) captured by the camera and language instructions (such as the task description) into a vector representation suitable for processing by the transformer. The transformer, acting as the "brain," analyzes multimodal input information in real time and dynamically adjusts the candidate actions generated by the diffusion model based on the current task requirements and robot state. Using an attention mechanism, the transformer focuses on and optimizes the most relevant action candidates for the current task. A 128-dimensional vector space is then defined as the unified action space, with each dimension corresponding to a specific physical quantity. For example, the first 10 dimensions could represent left arm joint angles, the middle 10 dimensions could represent right arm joint velocities, and subsequent dimensions could represent other relevant physical quantities such as gripper opening and closing, tool posture, etc. The actions of different robots (such as joint angles, velocities, and gripper states) are uniformly mapped into this 128-dimensional vector space. Predefined mapping rules are used to convert the robots' raw motion data into motion vectors in a unified format. All robot motion data is standardized before being input into the model to ensure that it conforms to the format requirements of the unified motion space. Since the motion data of different robots is unified into the same format, the model can extract common motion patterns and regularities by learning from this data. This mutual learning mechanism enables the model to leverage data from different robots to improve its generalization and prediction accuracy.

[0135] During the training or inference process of the diffusion model, image information captured by the camera is injected alternately. This information can include visual features such as the position, shape, and color of objects, providing the model with direct perception of the operating environment. At the same time, language instructions (such as operation task descriptions, user intentions, etc.) are also injected into the model. Language information provides the model with contextual information about task requirements and operation goals. By alternately injecting image and language information, the model can perceive changes in the environment and task requirements in real time and dynamically adjust the generated action sequence. This mechanism enables the model to maintain stable performance even when faced with unseen objects, scenes, or instructions, and adapt to the needs of dynamic environments.

[0136] In other words, by combining the diffusion model with the transformer architecture, the model can generate more diverse and refined candidate actions, and select the most appropriate action sequence through real-time analysis and adjustment by the transformer. This avoids the "average action" problem that can occur in traditional methods, resulting in more accurate action prediction and more coordinated dual-arm coordination. Furthermore, by designing a physically interpretable unified action space, the model can process action data from different robots and unify them into a common format for learning. This design greatly improves data utilization efficiency, allowing the model to leverage more data to improve its generalization and prediction accuracy. Furthermore, by alternately injecting image and language information, the model can perceive environmental changes and changes in task requirements in real time and dynamically adjust the generated action sequences. This mechanism enables the model to maintain stable performance even when faced with unprecedented objects, scenes, or instructions, and adapt to the demands of dynamic environments. Furthermore, the fusion of multimodal information enhances the model's ability to understand and process complex environments.

[0137] In addition, the application of the dual-arm robot operation control is like teaching the robot to learn the correct operation steps from "messy" actions. First, collect a lot of demonstration data (such as "proprioception", "action sequence", "control frequency / cloud time step"). This data is like a video of human demonstration, telling the robot how to complete a task, such as "pour water from a cup into a bowl". Then, use the diffusion model for training: first put the correct action a t (For example, "right arm picks up the cup") Add noise and turn it into a bunch of random actions. The calculation formula is:

[0138] Among them, α is a number that controls the amount of noise, ∈ is random noise, and the formula means "mess up the action"). Then, the core network learns to Recover the correct real action a t , calculated by using the loss function in, The action predicted by the model is fed back through the square error. The larger the error, the less accurate the prediction. The model automatically adjusts the internal parameters through back propagation (similar to reversing the cause of the error) and repeats this process until the error is reduced to below a set threshold. That is, by "calculating the predicted action" The model is optimized by “looking at the gap between the real action and the real action”. After training with a large amount of data, it can guess the correct action from the “mess”.

[0139] In a specific application, the robot is asked to complete a new task, such as "pour water from a cup into a bowl." The robot is fed with information about the current environment ("language instructions," "image input," and "proprioception"), such as photos of the positions of the bowl and cup, and a command. It starts with a completely random action (like a bunch of random movements), and then "denoises" it step by step through the core network using the following formula:

[0140]

[0141] in, represents the action generated after denoising in the k-1th step, and β is a proportional adjustment number, which is a weight coefficient (usually close to 1) and is used to control how many predicted actions are retained; The target action predicted by the model; The k-th step is a noisy action. The small noise is random Gaussian noise, which is used to avoid overly rigid results and increase diversity. Through step-by-step iteration, the final calculated result can be used to infer the target action predicted by the model. In other words, how to "step by step transform random actions back to correct actions" and ultimately output a predicted action sequence ("predicted action"), such as "right arm picks up the cup, left arm stabilizes the bowl, and pour water." These actions are sent to the robot for execution. If the water is successfully poured into the bowl, the task is completed.

[0142] For example, consider a task where a robot is asked to pick up a kettle with its left hand, hold a cup with its right hand, and pour water into the cup without spilling it. First, the diffusion model receives the verbal instruction "Pour water into the cup." Simultaneously, a camera captures the positions of the kettle and cup. The diffusion model then generates multiple possible pouring actions to prevent a single action from causing spillage. The transformer combines visual and verbal information to select the target action sequence. If the cup is knocked askew, for example, the diffusion model uses real-time camera feedback to readjust the action.

[0143] To give another example, dual-arm robotic operation and control are crucial in healthcare. During surgery, dual-arm robots can assist doctors with delicate procedures such as suturing, cutting, and tissue manipulation. By incorporating advanced control methods, dual-arm robots can perform surgical tasks with greater precision and efficiency, improving surgical success rates and patient recovery.

[0144] The doctor issues specific surgical instructions to the robot through voice recognition or text input, such as "suture the wound" or "cut tissue." The robot's onboard high-definition camera or an external medical imaging system (such as an endoscope, CT / MRI scanner, etc.) captures images of the surgical area to obtain environmental information. This information includes tissue morphology, location, and vascular distribution. The received surgical instructions are converted into a vector representation understandable by the robot using natural language processing techniques to obtain instruction information. Computer vision technologies such as convolutional neural networks (CNNs) are used to encode the captured surgical area images and extract visual features relevant to the surgical task, such as tissue edges and vascular orientation, to obtain feature information. State information sensed by the robot (such as joint angles, torques, and position) is encoded and converted into a vector representation relevant to the surgical task to obtain perception information. A diffusion model combines instruction information, feature information, and perception information to generate multiple candidate action sequences. These action sequences include the robot's arm trajectory, force control, and operation speed. Based on the requirements of the surgical task, criteria are set for evaluating candidate actions, such as suturing accuracy, cutting smoothness, and the degree of damage to surrounding tissue. For each candidate action group, corresponding evaluation metrics, such as suturing error and cut surface roughness, are calculated using the set evaluation criteria. Based on the calculated evaluation metrics, a comprehensive score is calculated for each candidate action group, and all candidate actions are ranked from high to low based on their comprehensive scores. Based on the ranking results, the candidate action group with the highest comprehensive score is selected as the target action. The target action is parsed from a vector or encoded form into a specific action sequence, including the robot's arm motion trajectory and force control parameters. Following the action sequence, action commands are sent to the robot one by one, controlling the robot's arms to perform the corresponding surgical operation. The robot executes the corresponding surgical action, such as suturing a wound or cutting tissue, according to the action commands. During the surgical operation, the robot monitors the surgical area's image information and robot status information in real time to determine whether the surgical operation has been completed. If the surgical operation is not completed, the process returns to the steps of receiving the operation command and acquiring environmental information, regenerating candidate actions and executing them until the surgical operation is complete.

[0145] In other words, by incorporating advanced operational control methods, dual-arm robots can perform surgical operations with greater precision, reduce human error, and improve surgical success rates. Furthermore, the robots can perform surgical tasks continuously and stably, unaffected by factors such as fatigue and emotion, thereby improving surgical efficiency. Furthermore, doctors can control the robots through voice or text commands, alleviating the physical strain of prolonged surgeries. Robots can also assist doctors with repetitive and tedious tasks such as suturing and cutting, allowing them to focus more on the critical aspects of the surgery. Furthermore, the application of dual-arm robot operational control methods is a key manifestation of the intelligent development of healthcare, providing strong support for future remote and intelligent surgery. By continuously optimizing and improving operational control methods, the performance and adaptability of robots can be further enhanced, driving the continuous advancement of healthcare technology.

[0146] Another example: In the logistics sector, dual-arm robots can be used to automate cargo sorting and handling tasks in large warehouses. Incorporating advanced operational control methods, these robots can efficiently and accurately perform tasks such as grabbing, handling, sorting, and placing goods, significantly improving the efficiency and accuracy of logistics operations.

[0147] The logistics center's management system sends operational instructions to a dual-arm robot through voice recognition or text input interfaces, such as "Move Class A cargo from Area 1 to Area 2." The robot uses an onboard high-definition camera or external vision system (such as a 3D vision sensor) to capture images within the warehouse, including cargo location, shape, size, and labels, to obtain environmental information. Natural language processing techniques are used to convert these instructions into vector representations understandable to the robot, specifying key information such as the task type (e.g., transport) and target area. Computer vision techniques, such as convolutional neural networks (CNNs), are used to encode the captured warehouse images and extract visual features such as cargo location, posture, and labels to obtain feature information. State information sensed by the robot (e.g., joint angles, position, and load) is encoded and converted into a task-specific vector representation to obtain perception information. A diffusion model combines these command, feature, and perception information to generate multiple candidate action sequences. These action sequences include the robot's dual-arm motion trajectory, gripping force, and transport path. Evaluation criteria for candidate actions, such as transport efficiency, accuracy, and energy consumption, are set based on the needs of the logistics operation. For each set of candidate actions, the corresponding evaluation metrics, such as handling time, error rate, and energy cost, are calculated using the pre-defined evaluation criteria. Based on the calculated evaluation metrics, a comprehensive score is calculated for each candidate action, and all candidate actions are ranked from high to low based on their comprehensive scores. Based on the ranking results, the set of candidate actions with the highest comprehensive scores is selected as the target action. The target action is parsed from a vector or encoded form into a specific action sequence, including the robot's arm motion trajectory, the precise grasping and placement locations, and so on. Following the action sequence, action commands are sent to the robot one by one, controlling the robot's arms to perform the corresponding handling operation. Following the action commands, the robot accurately grasps the goods and moves them to the target area along the planned path. During the handling operation, the robot monitors the goods' location and status in real time to determine whether the operation command has been completed. If the operation command is not completed (for example, if there are still goods to be handled), the robot returns to the steps of receiving the operation command and obtaining environmental information, regenerating candidate actions and executing them until all goods have been handled.

[0148] In other words, by incorporating advanced operational control methods, dual-arm robots can efficiently complete cargo grasping, handling, and sorting tasks, significantly reducing the time and cost of manual operations. Using their visual systems and sensory information, the robots can accurately identify cargo locations and labels, preventing handling errors and damage. Automated handling operations can reduce physical strain on workers, lower labor intensity, and improve job satisfaction. Furthermore, the application of dual-arm robot operational control methods is a key manifestation of the development of intelligent logistics, providing strong support for future unmanned warehouses and smart logistics. Diffusion models can generate a variety of candidate actions based on different operational instructions and environmental information, enabling robots to adapt to a variety of complex logistics scenarios.

[0149] The above-mentioned dual-arm robot operation control method can more accurately capture the complex motion patterns in dual-arm operation through the application of the diffusion model, effectively solving the problem of inaccurate prediction caused by the difficulty of traditional methods in capturing the diversity of motions. By generating multiple sets of candidate actions, the system can more comprehensively consider various possibilities, thereby improving the accuracy of motion prediction.

[0150] Furthermore, a unified action space was designed to map the actions of different robots into a vector space with clear physical meaning, enabling closer and more coordinated movements between the two arms. This design not only preserves the practical meaning of the movements but also promotes effective learning between different robot data, improving the coordination of dual-arm robot operations. By sorting multiple candidate actions from high to low according to their overall scores, with the highest-scoring candidate action becoming the target action, the system ensures optimal coordination between the two arms when performing tasks, further enhancing operational coordination and efficiency.

[0151] In addition, operation instructions and environmental information are input into the diffusion model to generate multiple sets of candidate actions, enabling the system to show stronger adaptability and stability in dynamic environments such as lighting changes and random object positions. By alternately injecting image and language information, the model can maintain stable performance under unseen objects, scenes or instructions, demonstrating strong zero-sample generalization capabilities.

[0152] In addition, the diffusion model generates candidate actions through a fast denoising process. Combined with the optimized algorithm and computing architecture, it significantly reduces the computational complexity of the model and improves the real-time control capability. It can achieve efficient and real-time operation control on resource-limited robot platforms and meet the real-time requirements in practical applications.

[0153] Furthermore, the unified action space design allows for full utilization of data resources from diverse sources, avoiding data waste caused by significant differences in robot structure and action space. Data from different robots can learn from each other within the unified space, improving data utilization efficiency and accelerating model training and optimization.

[0154] Figure 3 FIG is a schematic block diagram of a dual-arm robot operation control device 300 provided by an embodiment of the present invention. Figure 3 As shown, corresponding to the above dual-arm robot operation control method, the present invention also provides a dual-arm robot operation control device 300. The dual-arm robot operation control device 300 includes a unit for executing the above dual-arm robot operation control method, and the device can be configured in a server. Specifically, please refer to Figure 3 The dual-arm robot operation control device 300 includes a receiving and acquiring unit 301 , an input generating unit 302 , a sorting unit 303 and an executing unit 304 .

[0155] Receiving and obtaining unit 301, used to receive operation instructions and obtain environmental information;

[0156] An input generation unit 302, configured to input operation instructions and environmental information into the diffusion model to generate multiple sets of candidate actions;

[0157] A sorting unit 303 is used to sort the multiple groups of candidate actions from high to low according to the comprehensive scores, and the group of candidate actions with the highest comprehensive score is the target action;

[0158] The execution unit 304 is configured to execute the target action.

[0159] In one embodiment, the receiving and obtaining unit 301 includes:

[0160] A receiving module is used to receive operation instructions issued by the user through a voice recognition or text input interface;

[0161] The acquisition module is used to obtain image information of the current environment using the camera carried by the robot or the external vision system to obtain environmental information.

[0162] In one embodiment, the input generation unit 302 includes:

[0163] A conversion module is used to convert the received operation instructions into vector representations understandable to the robot through natural language processing technology to obtain instruction information;

[0164] The encoding and extraction module is used to encode the captured environmental information using convolutional neural network computer vision technology and extract visual features related to the task to obtain feature information;

[0165] The encoding conversion module is used to obtain and encode the state information perceived by the robot body, and then convert it into a vector representation related to the task to obtain perception information;

[0166] The input generation module is used to input instruction information, feature information and perception information into the diffusion model to generate multiple sets of candidate actions.

[0167] In one embodiment, the diffusion model is obtained by training big data, including:

[0168] The collection module is used to collect demonstration data, which includes the robot's proprioceptive information, action sequence, and corresponding control frequency when completing the task;

[0169] A preprocessing module, used to preprocess the demonstration data to obtain multimodal input encoding information;

[0170] The injection generation module is used to inject random noise into the correct action sequence at the beginning of training to generate a set of random action sequences;

[0171] The denoising module is used to denoise the chaotic action sequence to obtain denoised actions;

[0172] The prediction module is used to predict the next state that is closer to the real action based on the current action state, multimodal input encoding information and denoised action to obtain a predicted action sequence;

[0173] The iterative calculation module is used to calculate the loss between the predicted action sequence and the actual action sequence in each iteration to obtain the loss value;

[0174] The update module is used to update the parameters of the core network through the back propagation algorithm according to the loss value to minimize the loss;

[0175] The repeat execution module is used to repeatedly execute the above operations until the model reaches a preset performance index to obtain a diffusion model.

[0176] In one embodiment, the sorting unit 303 includes:

[0177] The setting module is used to set the criteria for evaluating candidate actions according to task requirements;

[0178] A calculation module is used to calculate the corresponding evaluation index for each group of candidate actions using the set evaluation criteria;

[0179] The calculation and ranking module is used to calculate a comprehensive score for each group of candidate actions based on the calculated evaluation indicators, and sort all candidate actions from high to low according to the comprehensive score;

[0180] The selection module is used to select a group of candidate actions with the highest comprehensive scores as the target actions based on the sorting results.

[0181] In one embodiment, the execution unit 304 includes:

[0182] The parsing module is used to parse the target action from a vector or encoding form into a specific action sequence;

[0183] The sending module is used to send action instructions to the robot one by one according to the action sequence;

[0184] The execution module is used for the robot to perform corresponding actions according to the action instructions.

[0185] In one embodiment, the apparatus further comprises:

[0186] The judging unit is used to judge whether the operation instruction is completed; if the operation instruction is not completed, returning to execute the received operation instruction and obtaining environmental information.

[0187] It should be noted that technical personnel in the relevant field can clearly understand that the specific implementation process of the above-mentioned dual-arm robot operation control device 300 and each unit can refer to the corresponding description in the aforementioned method embodiment. For the convenience and conciseness of the description, it will not be repeated here.

[0188] The dual-arm robot operation control device 300 can be implemented in the form of a computer program. The computer program can be used in Figure 4 Runs on the computer equipment shown.

[0189] See also Figure 4 , Figure 4 1 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device 500 may be a server, wherein the server may be an independent server or a server cluster composed of multiple servers.

[0190] See Figure 4 The computer device 500 includes a processor 502 , a memory, and a network interface 505 connected via a system bus 501 , wherein the memory may include a non-volatile storage medium 503 and an internal memory 504 .

[0191] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions, which, when executed, can enable the processor 502 to execute a dual-arm robot operation control method.

[0192] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.

[0193] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a dual-arm robot operation control method.

[0194] The network interface 505 is used to communicate with other devices through the network. Figure 4 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device 500 to which the solution of the present application is applied. The specific computer device 500 may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0195] The processor 502 is configured to execute a computer program 5032 stored in the memory to implement the following steps:

[0196] Receive operation instructions and obtain environmental information; input the operation instructions and environmental information into the diffusion model to generate multiple groups of candidate actions; sort the multiple groups of candidate actions from high to low according to the comprehensive score, and the group of candidate actions with the highest comprehensive score is the target action; execute the target action.

[0197] It should be understood that in the embodiment of the present application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0198] Those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program includes program instructions, which can be stored in a storage medium that is computer-readable. The program instructions are executed by at least one processor in the computer system to implement the steps in the method of the above-described embodiment.

[0199] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, wherein when the computer program is executed by a processor, the processor performs the following steps:

[0200] Receive operation instructions and obtain environmental information; input the operation instructions and environmental information into the diffusion model to generate multiple groups of candidate actions; sort the multiple groups of candidate actions from high to low according to the comprehensive score, and the group of candidate actions with the highest comprehensive score is the target action; execute the target action.

[0201] The storage medium may be any computer-readable storage medium that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.

[0202] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0203] In the several embodiments provided herein, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of various units is merely a logical functional division; actual implementations may employ other divisions. For example, multiple units or components may be combined or integrated into another system, or some features may be omitted or not implemented.

[0204] The steps in the methods of the embodiments of the present invention may be adjusted in order, combined, or deleted as needed. The units in the devices of the embodiments of the present invention may be combined, divided, or deleted as needed. Furthermore, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit.

[0205] If this integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, terminal, or network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present invention.

[0206] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A dual-arm robot operation control method, characterized in that: include: Receive operation instructions and obtain environmental information; Inputting operation instructions and environmental information into the diffusion model to generate multiple sets of candidate actions; Sort multiple groups of candidate actions from high to low according to their comprehensive scores, and the group of candidate actions with the highest comprehensive score is the target action; Perform the target action.

2. The dual-arm robot operation control method according to claim 1, characterized in that: The receiving of the operation instruction and obtaining the environment information includes: Receive operation instructions from users through voice recognition or text input interface; The robot uses the camera or external vision system to obtain image information of the current environment to obtain environmental information.

3. The dual-arm robot operation control method according to claim 1, characterized in that: The operation instructions and environmental information are input into the diffusion model to generate multiple sets of candidate actions, including: The received operation instructions are converted into vector representations that can be understood by the robot through natural language processing technology to obtain instruction information; Using convolutional neural network computer vision technology, the captured environmental information is encoded and task-related visual features are extracted to obtain feature information; Obtain and encode the state information perceived by the robot body, and then convert it into a vector representation related to the task to obtain perception information; The instruction information, feature information and perception information are input into the diffusion model to generate multiple sets of candidate actions.

4. The dual-arm robot operation control method according to claim 1, characterized in that: The diffusion model is obtained by big data training. The specific training steps are as follows: Collect demonstration data, which includes the robot's proprioception information, action sequence, and corresponding control frequency when completing the task; Preprocess the demonstration data to obtain multimodal input encoding information; At the beginning of training, random noise is injected into the correct action sequence to generate a set of random action sequences; De-noising the random action sequence to obtain denoised actions; Based on the current action state, multimodal input encoding information and denoised action, predict the next state that is closer to the real action to obtain the predicted action sequence; In each iteration, the loss between the predicted action sequence and the true action sequence is calculated to obtain the loss value; According to the loss value, the parameters of the core network are updated through the back propagation algorithm to minimize the loss; Repeat the above steps until the model reaches the preset performance index to obtain the diffusion model.

5. The dual-arm robot operation control method according to claim 1, characterized in that: The plurality of candidate actions are sorted from high to low according to the comprehensive scores, and the candidate action with the highest comprehensive score is the target action, including: Based on the task requirements, set the criteria for evaluating candidate actions; For each set of candidate actions, calculate the corresponding evaluation index using the set evaluation criteria; Based on the calculated evaluation indicators, a comprehensive score is calculated for each group of candidate actions, and all candidate actions are sorted from high to low according to the comprehensive score; According to the ranking results, a group of candidate actions with the highest comprehensive scores are selected as target actions.

6. The dual-arm robot operation control method according to claim 1, characterized in that: The executing target action includes: Parse the target action from a vector or encoding form into a specific action sequence; Send action instructions to the robot one by one according to the action sequence; The robot performs corresponding actions according to the action instructions.

7. The dual-arm robot operation control method according to claim 1, characterized in that: After executing the target action, the method further includes: Determine whether the operation instruction is completed; if the operation instruction is not completed, return to execute the received operation instruction and obtain environmental information.

8. A dual-arm robot operation control device, characterized in that: include: A receiving and obtaining unit, configured to receive an operation instruction and obtain environmental information; An input generation unit, configured to input operation instructions and environmental information into the diffusion model to generate multiple sets of candidate actions; A sorting unit is used to sort multiple groups of candidate actions from high to low according to their comprehensive scores. The group of candidate actions with the highest comprehensive score is the target action. Execution unit, used to execute the target action.

9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Gripping method and device based on double-arm robot and double-arm robot

    CN113538576A

  • Voice-driven posture action generation method and device based on diffusion model

    CN117292704A

  • Robot motion trajectory planning method and device and robot

    CN118404590A

  • Image element generation method and device, electronic equipment and storage medium

    CN119273787A

  • Simulation-Based Directed Surgical Development System

    US20220370133A1