Humanoid double-arm robot imitation method and device based on diffusion strategy
By employing a diffusion-based imitation method for humanoid dual-armed robots, and utilizing a task requirement encoder and a target imitation learning generative model, this approach addresses the problem that existing imitation learning algorithms cannot quickly, accurately, and stably simulate complex human movements and interactions in humanoid dual-armed robots, achieving efficient and accurate simulation learning results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-03-27
AI Technical Summary
Existing imitation learning algorithms cannot quickly, accurately, and stably simulate complex human movements and interactions in humanoid dual-arm robots. They suffer from high data dependence and difficulty in handling complex tasks and multimodal data.
A humanoid dual-arm robot imitation method based on diffusion strategy is adopted. By identifying imitation learning instructions and tasks, task requirement encoder is used to analyze tasks, determine task condition parameters, and robot imitation learning is performed based on target imitation learning generative model to generate high-quality action sequences.
It improves the efficiency and accuracy of imitation learning, enhances the diversity and accuracy of robot motion generation, improves the overall performance and stability of the robot, and enables it to better simulate complex human actions and interactions.
Smart Images

Figure CN121745152A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of robot simulation, and in particular to a bimanual humanoid robot imitation method and device based on a diffusion strategy. BACKGROUND
[0002] With the continuous development of artificial intelligence and robot technology, bimanual humanoid robots have attracted widespread attention in various fields such as industry, service, and medical treatment. Among them, bimanual humanoid robots have received particular attention due to their ability to simulate the flexibility and diversity of human hand operations. In order to enable these robots to better complete various tasks, imitation learning algorithms have become one of the research hotspots.
[0003] Imitation learning algorithms are a method for robots to learn skills by observing and learning the behavior of humans or other robots. Traditional imitation learning algorithms mainly include methods such as behavioral cloning and inverse reinforcement learning. The behavioral cloning algorithm directly imitates the input-output pairs demonstrated by experts, but is easily affected by incorrect demonstrations and is difficult to generalize to new task scenarios. The inverse reinforcement learning algorithm learns the reward function to guide the behavior of the robot, but usually requires a large amount of data and a complex optimization process, and performs poorly in multi-modal tasks; although existing imitation learning algorithms have made some progress in the field of robots, there are still some defects and limitations. For traditional behavioral cloning and inverse reinforcement learning algorithms, the main problem is the high dependence on data, and it is difficult to handle complex tasks and multi-modal data. The behavioral cloning algorithm is prone to produce incorrect behavior patterns when facing incorrect demonstrations, and is difficult to adapt to new task environments. The inverse reinforcement learning algorithm requires a large amount of data to learn the reward function, and needs to perform reinforcement learning to optimize, which further reduces its stability. SUMMARY
[0004] The present application provides a bimanual humanoid robot imitation method and device based on a diffusion strategy to solve the technical problem of existing bimanual humanoid robot imitation that cannot quickly and accurately simulate complex human actions and interactions.
[0005] According to an aspect of the present application, a bimanual humanoid robot imitation method based on a diffusion strategy is provided, comprising:
[0006] In response to the imitation learning instruction for the bimanual humanoid robot, identifying the imitation learning demonstration data and the imitation learning task corresponding to the imitation learning instruction;
[0007] Based on the task requirement encoder, performing task analysis on the imitation learning task to determine the task condition parameter;
[0008] determine a target imitation learning generation model based on the task condition parameter, and perform robot imitation learning based on the task condition parameter and the imitation learning demonstration data through the target imitation learning generation model to determine a robot action sequence.
[0009] According to another aspect of the present application, there is provided a diffusion strategy based humanoid dual-arm robot imitation device, comprising:
[0010] An interaction module is configured to, in response to an imitation learning instruction for the humanoid dual-arm robot, identify imitation learning demonstration data and an imitation learning task corresponding to the imitation learning instruction.
[0011] A task analysis module is configured to perform task analysis on the imitation learning task based on a task requirement encoder to determine a task condition parameter.
[0012] An imitation learning module is configured to determine a target imitation learning generation model based on the task condition parameter, and perform robot imitation learning based on the task condition parameter and the imitation learning demonstration data through the target imitation learning generation model to determine a robot action sequence.
[0013] According to another aspect of the present application, there is provided an electronic device, comprising:
[0014] at least one processor; and
[0015] a memory connected to the at least one processor in communication; wherein
[0016] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the diffusion strategy based humanoid dual-arm robot imitation method according to any one of the embodiments of the present application.
[0017] According to another aspect of the present application, there is provided a computer readable storage medium storing computer instructions for enabling a processor to perform the diffusion strategy based humanoid dual-arm robot imitation method according to any one of the embodiments of the present application.
[0018] The technical scheme of the embodiment of the present application can effectively improve the imitation learning efficiency and accuracy by responding to the imitation learning instruction of the humanoid two-arm robot, identifying the imitation learning demonstration data and the imitation learning task corresponding to the imitation learning instruction, and specifying the imitation task of the robot. The task demand encoder analyzes the imitation learning task based on the task demand, determines the task condition parameter, and quickly learns the high-quality two-arm action strategy through the task condition parameter, thereby improving the imitation learning efficiency. The target imitation learning generation model is determined based on the task condition parameter, and the robot imitation learning is performed based on the task condition parameter and the imitation learning demonstration data through the target imitation learning generation model, so as to determine the robot action sequence and make the robot better simulate the complex action and interaction of human beings, thereby improving the overall performance and stability of the robot, solving the technical problem that the humanoid two-arm robot imitation cannot quickly and accurately simulate the complex action and interaction of human beings in the prior art, effectively improving the efficiency and accuracy of robot simulation learning, enhancing the diversity and accuracy of robot action generation, making the robot better simulate the complex action and interaction of human beings, guaranteeing the overall performance and stability of the robot, and improving the working ability of the mass-produced service robot.
[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0021] Figure 1 A flowchart of a humanoid two-arm robot imitation method based on a diffusion strategy is provided for the embodiments of the present application.
[0022] Figure 2 A flowchart of a humanoid two-arm robot imitation method based on a diffusion strategy is provided for the embodiments of the present application.
[0023] Figure 3 A flowchart of a humanoid two-arm robot imitation method based on a diffusion strategy is provided for the embodiments of the present application.
[0024] Figure 4 A structural schematic diagram of a humanoid two-arm robot imitation device based on a diffusion strategy is provided for the embodiments of the present application.
[0025] Figure 5 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. Detailed Implementation
[0026] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0027] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0028] Figure 1 This invention provides a flowchart of a humanoid dual-arm robot imitation method based on a diffusion strategy. This embodiment is applicable to situations where a humanoid dual-arm robot imitates human movements. The method can be executed by a humanoid dual-arm robot imitation device based on a diffusion strategy. This device can be implemented in hardware and / or software and can be configured in a humanoid dual-arm robot and / or electronic devices. Figure 1 As shown, the method includes:
[0029] S110. In response to the imitation learning instruction of the humanoid dual-arm robot, identify the imitation learning demonstration data and imitation learning task corresponding to the imitation learning instruction.
[0030] Optionally, the humanoid dual-arm robot can be a robot composed of two arms and a humanoid configuration. The humanoid dual-arm robot can perform dexterous operations with its two arms and, by combining bipedal, wheeled, or composite mobile platforms, can perform general robotic tasks in complex industrial, daily life, and outdoor scenarios.
[0031] Optionally, the imitation learning instruction can be an instruction that controls the humanoid dual-arm robot to imitate and learn from the demonstrated behavior. It should be noted that the imitation learning instruction typically includes imitation learning demonstration data and imitation learning task. The imitation learning demonstration data can be a set of data, including state parameters, action instructions, and constraints required for the robot to complete the imitation learning task, recorded through worker instruction or environmental perception. The imitation learning task can be a textual description of the task to be imitated by the humanoid dual-arm robot.
[0032] Optionally, when recording imitation learning demonstration data, state data, sensor data, and environmental images are collected synchronously with a unified timestamp, and the multi-source data are aligned to obtain imitation learning demonstration data.
[0033] Optionally, the imitation learning task describes the actions and action requirements of the humanoid dual-arm robot; for example, the humanoid dual-arm robot is required to pick up an apple within 3 seconds with a position error ≤ ±1 mm and a joint angle error ≤ ±1°; or the humanoid dual-arm robot is required to pick up a glass within 10 seconds with a position error ≤ ±0.01 mm and a joint angle error ≤ ±0.001°.
[0034] Optionally, after recording the imitation learning demonstration data and setting the imitation learning task, the imitation learning demonstration data and the imitation learning task are combined into imitation learning instructions for the humanoid dual-arm robot, and the imitation learning instructions are sent to the humanoid dual-arm robot or the control device of the humanoid dual-arm robot.
[0035] Specifically, in response to the imitation learning instructions given to the humanoid dual-arm robot, the imitation learning demonstration data and imitation learning task corresponding to the imitation learning instructions are identified.
[0036] S120. Based on the task requirement encoder, perform task analysis on the imitation learning task to determine the task condition parameters.
[0037] Optionally, the task requirement encoder can be used as a functional unit for semantic analysis and task reasoning in imitation learning tasks.
[0038] Optionally, task condition parameters can be understood as the accuracy or speed parameters of the humanoid dual-arm robot during the imitation learning process. It should be noted that task condition parameters are the top-level decision-making basis in imitation learning, affecting the entire imitation process of the humanoid dual-arm robot. For example, if the task condition parameter is set to r, r→0 represents high-precision assembly, and r→1 represents high-speed handling, r affects the accuracy and speed of the humanoid dual-arm robot when performing the task. For instance, in an imitation learning task, if the humanoid dual-arm robot picks up an apple within 3 seconds with a positional error ≤ ±1mm and a joint angle error ≤ ±1°, then r = 0.97. During task execution, the efficiency priority of the humanoid dual-arm robot is >95%, and the accuracy priority is <5%.
[0039] In another scenario, if a humanoid dual-arm robot is required to pick up a glass within 10 seconds with a positional error ≤ ±0.01mm and a joint angle error ≤ ±0.001°, then r = 0.05 can be understood as the humanoid dual-arm robot having a precision priority > 95% and an efficiency priority < 5% when performing the task.
[0040] Specifically, the task-requirement encoder performs task analysis on the imitation learning task to determine the task condition parameters.
[0041] Optionally, in another optional embodiment of the present invention, the step of performing task analysis on the imitation learning task based on the task requirement encoder to determine task condition parameters includes:
[0042] The imitation learning task is converted into a task semantic vector; the task semantic vector is then used to perform semantic reasoning through a pre-trained semantic reasoning model to determine the task condition parameters.
[0043] Optionally, since the imitation learning task is a linguistic description, the task requirement encoder supports converting the textual and / or data descriptions in the imitation learning task into task semantic vectors. These task semantic vectors are vectorized representations of the imitation learning task. For example, the task requirement encoder integrates Sentence Transformers, using Sentence Transformers to vectorize the imitation learning task and determine the task semantic vectors.
[0044] Optionally, the semantic reasoning model can be a small inference model pre-trained for the reasoning imitation learning task. It should be noted that when designing the semantic reasoning model, a small Transformer model is selected, with the task semantic vector as input and the task condition parameters as output. An initial inference model is constructed, and the initial inference model is trained by collecting a model training dataset to obtain the semantic reasoning model. This semantic reasoning model is then embedded into the task requirement encoder.
[0045] Specifically, by calling the task requirement encoder, the task requirement encoder's vectorization function is used to convert the imitation learning task into a task semantic vector. Then, a pre-trained semantic reasoning model is used to perform semantic reasoning on the task semantic vector to determine the task condition parameters.
[0046] S130. Determine a target imitation learning generation model based on the task condition parameters, and use the target imitation learning generation model to perform robot imitation learning based on the task condition parameters and the imitation learning demonstration data to determine the robot action sequence.
[0047] Optionally, the target imitation learning generative model can be the imitation learning generative model selected by the humanoid dual-arm robot when handling the current imitation learning task. It should be noted that the imitation learning generative model includes a first imitation learning generative model and a second imitation learning generative model, both of which are used to handle the imitation learning task of the humanoid dual-arm robot.
[0048] Optionally, when faced with different task conditions, the humanoid dual-arm robot may select one or both of the first and second imitation learning generation models as the target imitation learning generation model. For example, when the task conditions require high motion accuracy, the first imitation learning generation model can be selected; when the task conditions require high motion smoothness, the second imitation learning generation model can be selected; and when high motion accuracy, smoothness, and real-time performance are also required, the first and second imitation learning generation models can be combined as the target imitation learning generation model.
[0049] Optionally, the robot motion sequence can be a series of continuous actions that a humanoid dual-arm robot can perform. It should be noted that the robot motion sequence describes the joint angles, angular velocities, end-effector positions, and postures of the humanoid dual-arm robot in each consecutive frame along the time dimension. By parsing the robot motion sequence, control commands are obtained to control the humanoid dual-arm robot to perform corresponding actions.
[0050] Specifically, a target imitation learning generative model is determined based on task condition parameters. The robot then uses this model to perform imitation learning based on the task condition parameters and imitation learning demonstration data to determine the robot's action sequence.
[0051] The technical solution of this invention, in response to a humanoid dual-arm robot's imitation learning instruction, identifies the imitation learning demonstration data and imitation learning task corresponding to the instruction. By specifying the robot's imitation task, the efficiency and accuracy of imitation learning can be effectively improved. Based on a task requirement encoder, the imitation learning task is analyzed to determine task condition parameters. These parameters enable the rapid learning of high-quality dual-arm movement strategies, further improving imitation learning efficiency. A target imitation learning generation model is determined based on these parameters. This model, along with the task condition parameters and the imitation learning demonstration data, enables the robot to perform imitation learning and determine its movement sequence. This allows the robot to better simulate complex human movements and interactions, improving its overall performance and stability. This solution addresses the technical problem in existing technologies where humanoid dual-arm robots cannot quickly, accurately, and stably simulate complex human movements and interactions. It effectively improves the efficiency and accuracy of robot simulation learning, enhances the diversity and accuracy of robot movement generation, and ensures the overall performance and stability of the robot, thereby improving the working capabilities of mass-produced service robots.
[0052] Optionally, in this invention, after obtaining the robot action sequence, when directly executing the instructions corresponding to the robot action sequence, there may be jitter or instability in the movement amplitude of the humanoid dual-arm robot. Therefore, further temporal fusion optimization is performed on the robot action sequence to eliminate the jitter or instability, ensuring the smoothness, diversity, and accuracy of the movements, while also solving the drift and cumulative error problems in long sequence movements. The specific process is as follows: calculate the temporal smoothing strategy of the action sequence based on the task condition parameters; perform temporal smoothing on the robot action sequence based on the temporal smoothing strategy to determine the target action sequence.
[0053] Optionally, the robot motion sequence generated based on task condition parameters may exhibit unnatural robot motion transitions and unstable swing amplitudes when the task condition parameters are of high precision, and motion jitter when the task condition parameters are of high efficiency. Furthermore, a corresponding temporal smoothing strategy is matched based on the temporal smoothing strategy.
[0054] Optionally, the temporal smoothing strategy can be a constraint strategy for the humanoid dual-arm robot in the time dimension. It should be noted that the temporal smoothing strategy can constrain the rate of change of motion of the humanoid dual-arm robot, loading the physical limits of each joint of the humanoid dual-arm robot. For example, the temporal smoothing strategy can choose a sliding window averaging...
[0055] Interpolation with Jerk-limited S-curves.
[0056] Optionally, the target action sequence can be an action sequence obtained by temporal smoothing of the robot action sequence.
[0057] Specifically, a temporal smoothing strategy for the action sequence is calculated based on task condition parameters; the robot's action sequence is then smoothed in the temporal domain based on the temporal smoothing strategy to determine the target action sequence.
[0058] Figure 2 This is a flowchart illustrating a humanoid dual-arm robot imitation method based on a diffusion strategy, provided as an embodiment of the present invention. The relationship between this embodiment and the previous embodiments is to specifically explain the detailed process of the humanoid dual-arm robot performing robot imitation learning based on a first imitation learning generative model. Figure 2 As shown, the method includes:
[0059] S210, In response to the imitation learning instruction of the humanoid dual-arm robot, identify the imitation learning demonstration data and imitation learning task corresponding to the imitation learning instruction.
[0060] S220. Based on the task requirement encoder, perform task analysis on the imitation learning task to determine the task condition parameters.
[0061] S230. If the task condition parameters are within the first parameter range, then the first imitation learning generation model is determined as the target imitation learning generation model.
[0062] Optionally, the first parameter range can be a pre-set parameter range for the first imitation learning generative model to meet the task condition parameter requirements. It should be noted that, based on the model performance of the first imitation learning generative model, a parameter range in which the first imitation learning generative model can perform high-precision simulation learning is pre-calculated, and this range is used as the first parameter range. For example, the first imitation learning generative model can be set as a denoising diffusion probabilistic model (DDPM). The second imitation learning generative model can be set as a flow matching model. Specifically, for the imitation learning demonstration data x0, DDPM schedules α with a fixed noise level during the forward phase. T β T σ T By gradually adding noise, we can obtain x. T In the reverse phase, the noise prediction network ε_θ(x) obtained through pre-training is used. T The algorithm performs a denoising update, iterates for T steps, and then outputs the robot action sequence x1 generated by the first imitation learning model. The specific calculation process is as follows:
[0063] μ_θ = (x T -β T / √(1-αT ) · ε_θ(x T ,t)) / √α t
[0064] x T-1 = μ_θ + σ T ·ε,ε~N(0,I)
[0065] The Flow Matching model establishes an ordinary differential equation with the imitation learning demonstration data x0 and the continuous time variable t∈[0,1], namely dx / dt = v_θ(x,t), where v_θ(x,t) is the velocity field network. It is trained with endpoint conditions x(0)~N(0,I) and x(1)=x0. During inference, the robot action sequence x2 output by the second imitation learning generation model is directly obtained through single-step or multi-step integration.
[0066] Specifically, if the task condition parameters are within the range of the first parameter, then the first imitation learning generative model is determined as the target imitation learning generative model.
[0067] S240. Perform data preprocessing on the imitation learning demonstration data to determine the characteristics of the multimodal robot.
[0068] The multimodal robot features can be multimodal robot feature data under a unified time axis. It should be noted that the imitation learning demonstration data consists of state data, sensor data, and environmental images. Data filtering is used to clean the various types of state and sensor data from multiple sources, removing outliers and filling in missing values. The state data, sensor data, and environmental images are then aligned according to timestamps to remove redundant dimensions. Finally, the environmental images are processed to obtain the multimodal robot features.
[0069] Optionally, in this invention, an environmental image is processed by a pre-trained convolutional neural network to extract the robot features corresponding to the humanoid dual-arm robot, and the environmental image is converted into a low-dimensional robot feature representation.
[0070] Specifically, data preprocessing is performed on the imitation learning demonstration data to determine the characteristics of the multimodal robot.
[0071] S250. The robot action sequence is determined by generating an action sequence based on the task condition parameters and the multimodal robot features using the first imitation learning generation model.
[0072] Optionally, the first imitation learning generative model performs forward Gaussian noise based on task condition parameters and multimodal robot features to obtain the noise of the discrete time step under the forward path. The noise of the discrete time step under the forward path is then denoised through the reverse path to obtain the robot action sequence.
[0073] Specifically, the robot action sequence is determined by generating action sequences based on task condition parameters and multimodal robot features through the first imitation learning generative model.
[0074] Optionally, in another optional embodiment of the present invention, the step of training the imitation learning sub-model based on the task condition parameters and the multimodal robot features to determine the robot action features includes:
[0075] The noise coefficient value and the noise reduction intensity value are determined based on the task condition parameters; the robot action sequence is determined by generating an action sequence through the first imitation learning generation model based on the noise coefficient value, the noise reduction intensity value and the multimodal robot features.
[0076] The noise addition coefficient value can be calculated based on the task condition parameters to add noise along the forward path of the first imitation learning generative model; the denoising intensity value can be calculated based on the task condition parameters to denoise along the backward path of the first imitation learning generative model. For example, if the task condition parameters are set to 0.05, the noise addition coefficient value will be 0.024, and the denoising intensity value will be 0.1475.
[0077] Optionally, the forward path of the first imitation learning generative model progressively adds noise to the multimodal robot features by adding noise coefficient values, and the reverse path progressively denoises and updates the robot action sequence based on the denoising intensity values.
[0078] Specifically, the noise coefficient value and the noise reduction intensity value are determined based on the task condition parameters; the robot action sequence is determined by generating the action sequence through the first imitation learning generative model based on the noise coefficient value, the noise reduction intensity value and the multimodal robot features.
[0079] The technical solution of this invention, in response to a humanoid dual-arm robot's imitation learning instruction, identifies the imitation learning demonstration data and imitation learning task corresponding to the instruction. By specifying the robot's imitation task, the efficiency and accuracy of imitation learning can be effectively improved. Based on a task requirement encoder, the imitation learning task is analyzed to determine task condition parameters. These parameters enable the rapid learning of high-quality dual-arm movement strategies, further improving imitation learning efficiency. A target imitation learning generation model is determined based on these parameters. This model, along with the task condition parameters and the imitation learning demonstration data, enables the robot to perform imitation learning and determine its movement sequence. This allows the robot to better simulate complex human movements and interactions, improving its overall performance and stability. This solution addresses the technical problem in existing technologies where humanoid dual-arm robots cannot quickly, accurately, and stably simulate complex human movements and interactions. It effectively improves the efficiency and accuracy of robot simulation learning, enhances the diversity and accuracy of robot movement generation, and ensures the overall performance and stability of the robot, thereby improving the working capabilities of mass-produced service robots.
[0080] Figure 3 This is a flowchart illustrating a humanoid dual-arm robot imitation method based on a diffusion strategy, provided as an embodiment of the present invention. The relationship between this embodiment and the previous embodiments is that this is a detailed process of the humanoid dual-arm robot performing robot imitation learning based on a second imitation learning generative model. Figure 3 As shown, the method includes:
[0081] S310, In response to the imitation learning instruction for the humanoid dual-arm robot, identify the imitation learning demonstration data and imitation learning task corresponding to the imitation learning instruction.
[0082] S320. Based on the task requirement encoder, perform task analysis on the imitation learning task to determine the task condition parameters.
[0083] S330. If the task condition parameters are within the range of the second parameter, then the second imitation learning generation model is determined as the target imitation learning generation model.
[0084] Optionally, the second parameter range can be a pre-set parameter range for the second imitation learning generative model to meet the task condition parameter requirements. It should be noted that, based on the model performance of the second imitation learning generative model, the parameter range in which the second imitation learning generative model can perform high-precision simulation learning is pre-calculated, and this range is used as the second parameter range.
[0085] Specifically, if the task condition parameters are within the range of the second parameter, then the second imitation learning generative model is determined as the target imitation learning generative model.
[0086] S340. Perform data preprocessing on the imitation learning demonstration data to determine the characteristics of the multimodal robot.
[0087] Specifically, data preprocessing is performed on the imitation learning demonstration data to determine the characteristics of the multimodal robot.
[0088] S350. Determine the complex motion sampling density and flow field intensity scaling factor based on the task condition parameters.
[0089] The complex motion sampling density can be the data value obtained from the flow field learning stage based on the task condition parameters; the flow field intensity scaling factor can be the coefficient value obtained from the flow field learning stage based on the task condition parameters. For example, if the task condition parameter is set to 0.97, then the complex motion sampling density = 0.01, and the flow field intensity scaling factor = 0.8.
[0090] Specifically, the sampling density of complex actions and the scaling factor of the flow field intensity are determined based on the task condition parameters.
[0091] S350. The robot action sequence is determined by constructing an action sequence based on the complex action sampling density, the flow field intensity scaling factor, and the multimodal robot features through the second imitation learning generation model.
[0092] Optionally, the second imitation learning generative model serves as an adaptive velocity field network. During the process learning phase, the second imitation learning generative model samples based on the complex action sampling density and the flow field intensity scaling factor. Under the condition of meeting the error requirements, it generates robot action sequences during the generation phase.
[0093] Specifically, the robot action sequence is determined by constructing action sequences based on complex action sampling density, flow field intensity scaling factor and multimodal robot features through the second imitation learning generative model.
[0094] The technical solution of this invention, in response to a humanoid dual-arm robot's imitation learning instruction, identifies the imitation learning demonstration data and imitation learning task corresponding to the instruction. By specifying the robot's imitation task, the efficiency and accuracy of imitation learning can be effectively improved. Based on a task requirement encoder, the imitation learning task is analyzed to determine task condition parameters. These parameters enable the rapid learning of high-quality dual-arm movement strategies, further improving imitation learning efficiency. A target imitation learning generation model is determined based on these parameters. This model, along with the task condition parameters and the imitation learning demonstration data, enables the robot to perform imitation learning and determine its movement sequence. This allows the robot to better simulate complex human movements and interactions, improving its overall performance and stability. This solution addresses the technical problem in existing technologies where humanoid dual-arm robots cannot quickly, accurately, and stably simulate complex human movements and interactions. It effectively improves the efficiency and accuracy of robot simulation learning, enhances the diversity and accuracy of robot movement generation, and ensures the overall performance and stability of the robot, thereby improving the working capabilities of mass-produced service robots.
[0095] Optionally, in another optional embodiment of the present invention, when the task condition parameters satisfy the joint generation of the first imitation learning generation model and the second imitation learning generation model, the forward and backward paths of the first imitation learning generation model can be connected to the generation stage of the second imitation learning generation model. The first imitation learning generation model progressively adds noise to the multimodal robot features through the forward path using noise coefficient values, and progressively updates the noise based on the denoising intensity values in the backward path, gradually learning the motion feature rules corresponding to the task condition parameters. These motion feature rules are then converted into the direction and intensity of the continuous flow field of the second imitation learning generation model. A robot motion sequence is generated based on the second imitation learning generation model, and the robot motion sequence is temporally smoothed by combining the temporal smoothing strategy of the task condition parameters to determine the target motion sequence. By utilizing the high-precision motion adjustment rules of the first imitation learning generation model to guide the second imitation learning generation model in generating the robot motion sequence of the humanoid dual-arm robot in real time and smoothly, the learning of the imitation task can be achieved quickly and accurately.
[0096] Figure 4 This is a schematic diagram of a humanoid dual-arm robot mimicking device based on a diffusion strategy, provided as an embodiment of the present invention. Figure 4 As shown, the device includes: an interaction module 410, a task analysis module 420, and an imitation learning module 430; wherein,
[0097] Interaction module 410 is used to respond to the imitation learning command of the humanoid dual-arm robot and identify the imitation learning demonstration data and imitation learning task corresponding to the imitation learning command;
[0098] Task analysis module 420 is used to perform task analysis on the imitation learning task based on the task requirement encoder and determine task condition parameters;
[0099] The imitation learning module 430 is used to determine a target imitation learning generation model based on the task condition parameters, and to perform robot imitation learning based on the task condition parameters and the imitation learning demonstration data through the target imitation learning generation model to determine the robot action sequence.
[0100] The technical solution of this invention, in response to a humanoid dual-arm robot's imitation learning instruction, identifies the imitation learning demonstration data and imitation learning task corresponding to the instruction. By specifying the robot's imitation task, the efficiency and accuracy of imitation learning can be effectively improved. Based on a task requirement encoder, the imitation learning task is analyzed to determine task condition parameters. These parameters enable the rapid learning of high-quality dual-arm movement strategies, further improving imitation learning efficiency. A target imitation learning generation model is determined based on these parameters. This model, along with the task condition parameters and the imitation learning demonstration data, enables the robot to perform imitation learning and determine its movement sequence. This allows the robot to better simulate complex human movements and interactions, improving its overall performance and stability. This solution addresses the technical problem in existing technologies where humanoid dual-arm robots cannot quickly, accurately, and stably simulate complex human movements and interactions. It effectively improves the efficiency and accuracy of robot simulation learning, enhances the diversity and accuracy of robot movement generation, and ensures the overall performance and stability of the robot, thereby improving the working capabilities of mass-produced service robots.
[0101] Optionally, the imitation learning module 430 is specifically used for:
[0102] If the task condition parameters are within the range of the first parameter, then the first imitation learning generation model is determined as the target imitation learning generation model;
[0103] If the task condition parameters are within the range of the second parameter, then the second imitation learning generation model is determined as the target imitation learning generation model.
[0104] Optionally, the imitation learning module 430 is further used for:
[0105] The target imitation learning generation model is the first imitation learning generation model;
[0106] The imitation learning demonstration data is preprocessed to determine the characteristics of the multimodal robot;
[0107] The robot action sequence is determined by generating an action sequence based on the task condition parameters and the multimodal robot features using the first imitation learning generative model.
[0108] Optionally, the imitation learning module 430 is further used for:
[0109] The noise addition coefficient value and the noise reduction intensity value are determined based on the aforementioned task condition parameters;
[0110] The robot action sequence is determined by generating an action sequence using the first imitation learning generative model based on the noise coefficient value, the noise reduction intensity value, and the multimodal robot features.
[0111] Optionally, the imitation learning module 430 is further used for:
[0112] The target imitation learning generation model is the second imitation learning generation model;
[0113] The imitation learning demonstration data is preprocessed to determine the characteristics of the multimodal robot;
[0114] The sampling density of complex actions and the scaling factor of the flow field intensity are determined based on the aforementioned task condition parameters;
[0115] The robot action sequence is determined by constructing an action sequence based on the complex action sampling density, the flow field intensity scaling factor, and the multimodal robot features using the second imitation learning generation model.
[0116] Optionally, the device further includes an optimization module, which is specifically used for:
[0117] Calculate the temporal smoothing strategy for the action sequence based on the task condition parameters;
[0118] The robot's action sequence is smoothed in the time domain based on the aforementioned time domain smoothing strategy to determine the target action sequence.
[0119] Optionally, the task analysis module 420 is specifically used for:
[0120] The imitation learning task is converted into a task semantic vector;
[0121] The task condition parameters are determined by performing semantic reasoning on the task semantic vector using a pre-trained semantic reasoning model.
[0122] The humanoid dual-arm robot imitation device based on diffusion strategy provided in the embodiments of the present invention can execute the humanoid dual-arm robot imitation method based on diffusion strategy provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0123] Figure 5 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their patterns are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0124] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0125] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of monitors, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer grids such as the Internet and / or various telecommunications grids.
[0126] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the humanoid dual-arm robot mimicry method based on a diffusion strategy.
[0127] In some embodiments, the diffusion-based humanoid dual-arm robot mimicry method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the diffusion-based humanoid dual-arm robot mimicry method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the diffusion-based humanoid dual-arm robot mimicry method by any other suitable means (e.g., by means of firmware).
[0128] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0129] Computer programs used to implement the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the patterns / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs can be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0130] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0131] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0132] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or grid browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication grid). Examples of communication grids include local area networks (LANs), wide area networks (WANs), blockchain grids, and the Internet.
[0133] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS servers, such as high management difficulty and weak business scalability.
[0134] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0135] This embodiment provides a computer-readable storage medium storing a computer program thereon. When executed by a processor, the program implements the steps of the humanoid dual-armed robot imitation method based on a diffusion strategy provided in any embodiment of the present invention. The method includes:
[0136] In response to the imitation learning command of the humanoid dual-arm robot, the imitation learning demonstration data and imitation learning task corresponding to the imitation learning command are identified;
[0137] Based on the task requirement encoder, the imitation learning task is analyzed to determine the task condition parameters;
[0138] Based on the task condition parameters, a target imitation learning generation model is determined. The robot then performs imitation learning based on the task condition parameters and the imitation learning demonstration data using the target imitation learning generation model to determine the robot's action sequence.
[0139] The computer storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0140] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0141] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0142] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of mesh, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0143] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a grid of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computing device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0144] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0145] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for mimicking a humanoid dual-armed robot based on a diffusion strategy, characterized in that, include: In response to the imitation learning command of the humanoid dual-arm robot, the imitation learning demonstration data and imitation learning task corresponding to the imitation learning command are identified; Based on the task requirement encoder, the imitation learning task is analyzed to determine the task condition parameters; Based on the task condition parameters, a target imitation learning generation model is determined. The robot then performs imitation learning based on the task condition parameters and the imitation learning demonstration data using the target imitation learning generation model to determine the robot's action sequence.
2. The method according to claim 1, characterized in that, The step of determining the target imitation learning generation model based on the task condition parameters includes: If the task condition parameters are within the range of the first parameter, then the first imitation learning generation model is determined as the target imitation learning generation model; If the task condition parameters are within the range of the second parameter, then the second imitation learning generation model is determined as the target imitation learning generation model.
3. The method according to claim 2, characterized in that, The target imitation learning generation model is the first imitation learning generation model; The step of using the target imitation learning generation model to perform robot imitation learning based on the task condition parameters and the imitation learning demonstration data to determine the robot action sequence includes: The imitation learning demonstration data is preprocessed to determine the characteristics of the multimodal robot; The robot action sequence is determined by generating an action sequence based on the task condition parameters and the multimodal robot features using the first imitation learning generative model.
4. The method according to claim 3, characterized in that, The step of generating a sequence of actions based on the task condition parameters and the multimodal robot features using the first imitation learning generative model, and determining the robot action sequence, includes: The noise addition coefficient value and the noise reduction intensity value are determined based on the aforementioned task condition parameters; The robot action sequence is determined by generating an action sequence using the first imitation learning generative model based on the noise coefficient value, the noise reduction intensity value, and the multimodal robot features.
5. The method according to claim 2, characterized in that, The target imitation learning generation model is the second imitation learning generation model; The step of using the target imitation learning generation model to perform robot imitation learning based on the task condition parameters and the imitation learning demonstration data to determine the robot action sequence includes: The imitation learning demonstration data is preprocessed to determine the characteristics of the multimodal robot; The sampling density of complex actions and the scaling factor of the flow field intensity are determined based on the aforementioned task condition parameters; The robot action sequence is determined by constructing an action sequence based on the complex action sampling density, the flow field intensity scaling factor, and the multimodal robot features using the second imitation learning generation model.
6. The method according to claim 1, characterized in that, Also includes: Calculate the temporal smoothing strategy for the action sequence based on the task condition parameters; The robot's action sequence is smoothed in the time domain based on the aforementioned time domain smoothing strategy to determine the target action sequence.
7. The method according to claim 1, characterized in that, The task analysis of the imitation learning task based on the task requirement encoder, and the determination of task condition parameters, include: The imitation learning task is converted into a task semantic vector; The task condition parameters are determined by performing semantic reasoning on the task semantic vector using a pre-trained semantic reasoning model.
8. A humanoid dual-arm robot mimicry device based on a diffusion strategy, characterized in that, include: The interaction module is used to respond to the imitation learning instructions of the humanoid dual-arm robot and identify the imitation learning demonstration data and imitation learning task corresponding to the imitation learning instructions. The task analysis module is used to perform task analysis on the imitation learning task based on the task requirement encoder and determine the task condition parameters. The imitation learning module is used to determine a target imitation learning generation model based on the task condition parameters, and to perform robot imitation learning based on the task condition parameters and the imitation learning demonstration data using the target imitation learning generation model to determine the robot action sequence.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the humanoid dual-arm robot imitation method based on the diffusion strategy according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the humanoid dual-arm robot imitation method based on a diffusion strategy as described in any one of claims 1-7.