Humanoid robot imitation learning method and device, computer equipment and storage medium

By using training data including actual joint position and end force in humanoid robot imitation learning, the imitation learning model is solved, and the problem of insufficient whole-body motion control and force control in the prior art is achieved, and more precise and detailed humanoid robot control is achieved.

CN120116218AActive Publication Date: 2025-06-10KEPLER ROBOT CO LTD

Patent Information

Application Number
CN202510320982.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-10
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

The existing imitation learning model has significant shortcomings in whole-body motion control and velocity control, which is difficult to meet the needs of humanoid robots in diverse application scenarios.

Method used

By obtaining the data collected when the imitation robot imitates the original object, including actual joint position and actual end force, the initial model is trained to obtain the imitation learning model. The model is able to make the target humanoid robot perform actions based on the training data, and optimize action prediction and reconstruction losses by updating the encoder and decoder weight parameters.

Benefits of technology

The imitation learning model is achieved to be closer to human movement habits, making the control of humanoid robots more accurate and detailed, especially in the agile operation of hands and fingers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120116218A_ABST
    Figure CN120116218A_ABST
Patent Text Reader

Abstract

The invention relates to a humanoid robot imitation learning method and device, computer equipment and a storage medium. The method comprises the steps of obtaining training data; according to the training data, training an initial model to obtain an imitation learning model; enabling a target humanoid robot to execute an action according to the imitation learning model; wherein the training data is data acquired by taking the simulation robot as a center when the simulation robot simulates an original object, and the training data comprises an actual joint position and an actual tail end force. According to the invention, the simulation learning model can be closer to the action habits of human beings, and the control of the humanoid robot can be more accurate and meticulous.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robots, and in particular to a method, device, computer device and storage medium for a humanoid robot to imitate and learn. Background Art

[0002] With the development of computer technology and artificial intelligence, humanoid robots have become an important research direction in the field of humanoid robots.

[0003] Imitation learning, imitation learning data collection and establishing an imitation model are the basic methods for humanoid robots to imitate human behaviors. Although the existing imitation models have made certain progress in achieving autonomous and dexterous operations, there are still significant deficiencies in whole-body motion control, especially in force control and the like.

[0004] Therefore, there is an urgent need for a new imitation learning method that can better meet the requirements of humanoid robots in diverse application scenarios. Summary of the Invention

[0005] To solve the above technical problems or at least partially solve the above technical problems, the present invention provides a method, device, computer device and storage medium for a humanoid robot to imitate and learn.

[0006] In a first aspect, the present invention provides a method for a humanoid robot to imitate and learn, the method comprising:

[0007] Obtaining training data;

[0008] Training an initial model according to the training data to obtain an imitation learning model;

[0009] Making a target humanoid robot execute an action according to the imitation learning model;

[0010] Wherein, the training data is data collected with the imitation robot as the center when the imitation robot imitates an original object, and the training data includes actual joint positions and actual end forces.

[0011] Optionally, the training the initial model according to the training data to obtain an imitation learning model includes:

[0012] Inputting the current training data into the initial model to obtain the current total loss;

[0013] Judging whether the current total loss meets the convergence condition;

[0014] If the convergence condition is met, end the training, and use the initial model at the end of the training as the imitation learning model;

[0015] If the convergence condition is not met, continue to obtain the next training data.

[0016] Optionally, the current training data includes: the current actual joint positions, the current actual end - effector forces, the current RGB images, the current depth images, and action sequence data;

[0017] The step of inputting the current training data into the initial model to obtain the current total loss includes:

[0018] Updating the previous encoder weight parameters according to the previous total loss to obtain the current encoder weight parameters;

[0019] Obtaining the current style variable according to the current encoder model, the current encoder weight parameters, the current actual joint positions, the current actual end - effector forces, and the action sequence data;

[0020] Updating the previous decoder weight parameters according to the previous total loss to obtain the current decoder weight parameters;

[0021] Obtaining action prediction data according to the current decoder model, the current style variable, the current decoder weight parameters, the current actual joint positions, the current actual end - effector forces, the current RGB images, and the current depth images;

[0022] Obtaining the mean square error between the action prediction data and the action sequence data;

[0023] Obtaining the current reconstruction loss according to the mean square error;

[0024] Obtaining the current regularization term according to the current style variable;

[0025] Obtaining the current total loss according to the current regularization term and the current reconstruction loss.

[0026] Optionally, the step of obtaining the current style variable according to the current encoder model, the current encoder weight parameters, the current actual joint positions, the current actual end - effector forces, and the action sequence data is as follows:

[0027]

[0028] where \(q\) represents the current encoder model, is the current encoder weight parameter, \(a\) t:t+k is the action sequence data, including the current actual joint positions and the current actual end - effector forces, is the current style variable.

[0029] Optionally, the step of obtaining the current style variable according to the current encoder model, the current encoder weight parameters, the current actual joint positions, the current actual end - effector forces, and the action sequence data includes:

[0030] Map the action sequence data to features of a preset dimension through a first linear layer, and perform sine position encoding to obtain a first feature;

[0031] Map the current encoder weights to features of a preset dimension through a second linear layer to obtain a second feature;

[0032] Map the current actual joint position to features of a preset dimension through a third linear layer to obtain a third feature;

[0033] Map the current actual end - effector force to features of a preset dimension through a fourth linear layer to obtain a fourth feature;

[0034] Through the current encoder model, obtain a fifth feature according to the first feature, the second feature, the third feature, and the fourth feature;

[0035] Map the fifth feature through a fifth linear layer to obtain the current style variable.

[0036] Optionally, obtaining action prediction data according to the current decoder model, the current style variable, the current decoder weight parameters, the current actual joint position, the current actual end - effector force, the current RGB image, and the current depth image is performed in the following manner:

[0037]

[0038] Wherein, π represents the current decoder model, θ is the current decoder weight parameter, is the action prediction data, is the current style variable, o t includes the current actual joint position, the current actual end - effector force, the current RGB image, and the current depth image.

[0039] Optionally, obtaining action prediction data according to the current decoder model, the current style variable, the current decoder weight parameters, the current actual joint position, the current actual end - effector force, the current RGB image, and the current depth image includes:

[0040] Update the original depth image according to the number of channels of the current RGB image to obtain the current depth image;

[0041] Map the current depth image and the current RGB image to features of a preset dimension through a sixth linear layer to obtain a sixth feature;

[0042] Map the current actual joint position to features of a preset dimension through a seventh linear layer to obtain a seventh feature;

[0043] Map the current actual end - effector force to a feature of a preset dimension through an eighth linear layer as the eighth feature;

[0044] Map the current style variable to a feature of a preset dimension through a ninth linear layer as the ninth feature;

[0045] Through the current encoder model, obtain the action prediction data according to the sixth feature, the seventh line, the eighth feature, and the ninth feature.

[0046] Optionally, obtain the current reconstruction loss according to the mean squared error in the following manner:

[0047]

[0048] Obtain the current regularization term according to the current style variable in the following manner:

[0049]

[0050] Obtain the current total loss according to the current regularization term and the current reconstruction loss in the following manner:

[0051]

[0052] Wherein, is the current reconstruction loss, is the action prediction data, a t:t+k is the action sequence data, is the current regularization term, represents the standard normal distribution, D KL represents the KL divergence, q represents the current encoder model, is the current encoder weight parameter, is the current style variable, includes the current actual joint positions and the current actual end - effector force, is the current total loss.

[0053] Optionally, updating the original depth image according to the number of channels of the current RGB image to obtain the current depth image includes:

[0054] Copy the original depth image M times;

[0055] Stitch the M copies of the original depth image according to the channels to obtain the current depth image;

[0056] where M is the number of channels of the current RGB image.

[0057] Second aspect, a humanoid robot imitation learning device is provided, and the device includes:

[0058] An acquisition module, configured to acquire training data;

[0059] A training module, configured to train an initial model according to the training data to obtain an imitation learning model;

[0060] An execution module, configured to cause a target humanoid robot to execute an action according to the imitation learning model;

[0061] Wherein, the training data is data collected with the imitation robot as the center when the imitation robot imitates the original object, and the training data includes actual joint positions and actual end forces.

[0062] Third aspect, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method described in any one of the above is implemented.

[0063] Fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method described in any one of the above is implemented.

[0064] The present invention provides a humanoid robot imitation learning method, device, computer device, and storage medium. The method includes: acquiring training data; training an initial model according to the training data to obtain an imitation learning model; causing a target humanoid robot to execute an action according to the imitation learning model; wherein, the training data is data collected with the imitation robot as the center when the imitation robot imitates the original object, and the training data includes actual joint positions and actual end forces. In the embodiments of the present invention, the training data is data collected with the imitation robot as the center when the imitation robot imitates the original object, that is, the training data is from the perspective of the humanoid robot. The most important point of a humanoid robot is to imitate humans. From the main vision of the humanoid robot, the imitation learning model can be made closer to human movement habits. In the embodiments of the present invention, the training data includes end forces, so that the trained model can also output end forces, and the control information of the end forces is also added to the control of the humanoid robot. In various fine operations, especially in the dexterous operations of the hands and fingers of the humanoid robot, the control information of the end forces can make the control of the humanoid robot more precise and detailed, and can also protect the task object. The method of the embodiments of the present invention can make the imitation learning model closer to human movement habits and can also make the control of the humanoid robot more precise and detailed. Description of the Drawings

[0065] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments in accordance with the present invention, and are used together with the specification to explain the principles of the present invention.

[0066] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0067] Figure 1 The following shows an application environment diagram of the humanoid robot imitation learning method according to an embodiment of the present invention;

[0068] Figure 2 The following shows a schematic flowchart of the humanoid robot imitation learning method according to an embodiment of the present invention;

[0069] Figure 3 The following shows a flowchart of obtaining the current style variable according to an embodiment of the present invention;

[0070] Figure 4 The following shows a flowchart of obtaining action prediction data according to an embodiment of the present invention;

[0071] Figure 5 The following shows a structural block diagram of the humanoid robot imitation learning device according to an embodiment of the present invention;

[0072] Figure 6 The following is an internal structure diagram of a computer device in an embodiment of the present invention. Detailed Embodiments

[0073] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0074] Figure 1 The following shows an application environment diagram of the humanoid robot imitation learning method according to an embodiment of the present invention. Refer to Figure 1, the humanoid robot imitation learning method is applied to a humanoid robot imitation learning system. The humanoid robot imitation learning method includes a terminal 110 and / or a server 120. The terminal 110 and the server 120 are connected through a network. The terminal 110 may specifically be a desktop terminal or a mobile terminal, and the mobile terminal may specifically be at least one of a mobile phone, a tablet computer, a laptop computer, etc. The server 120 may be implemented by an independent server or a server cluster composed of multiple servers.

[0075] The humanoid robot imitation learning method of the present invention is applied to the terminal 110 and / or the server 120.

[0076] Figure 2 The flowchart of a humanoid robot imitation learning method according to an embodiment of the present invention is shown as Figure 2 As shown, the method includes:

[0077] Step 210, obtaining training data;

[0078] Step 220, training an initial model according to the training data to obtain an imitation learning model;

[0079] Step 230, making a target humanoid robot execute an action according to the imitation learning model;

[0080] Wherein, the training data is data collected with the imitation robot as the center when the imitation robot imitates the original object, and the training data includes actual joint positions and actual end forces.

[0081] In an embodiment of the present invention, the training data is data collected with the imitation robot as the center when the imitation robot imitates the original object, that is, the training data is from the perspective of the humanoid robot. The most important point of a humanoid robot is to imitate humans. From the main vision of the humanoid robot, the imitation learning model can be made closer to human movement habits.

[0082] In an embodiment of the present invention, the training data includes end forces, so that the trained model can also output end forces, and the control information of the end forces is also added to the control of the humanoid robot. In various fine operations, especially when the humanoid robot performs dexterous operations on its hands and fingers, the control information of the end forces can make the control of the humanoid robot more accurate and detailed, and can also protect the task object.

[0083] The method of the embodiment of the present invention can make the imitation learning model closer to human movement habits, and can also make the control of the humanoid robot more accurate and detailed.

[0084] In an embodiment of the present invention, in step 220, the training the initial model according to the training data to obtain an imitation learning model includes:

[0085] Input the current training data into the initial model to obtain the current total loss;

[0086] Determine whether the current total loss meets the convergence condition;

[0087] If the convergence condition is met, end the training and use the initial model at the end of training as the imitation learning model;

[0088] If the convergence condition is not met, continue to obtain the next training data.

[0089] In the embodiments of the present invention, the current training data includes: the current actual joint position, the current actual end force, the current RGB image, the current depth image, and the action sequence data;

[0090] The step of inputting the current training data into the initial model to obtain the current total loss includes:

[0091] Update the previous encoder weight parameters according to the previous total loss to obtain the current encoder weight parameters;

[0092] Obtain the current style variable according to the current encoder model, the current encoder weight parameters, the current actual joint position, the current actual end force, and the action sequence data;

[0093] Update the previous decoder weight parameters according to the previous total loss to obtain the current decoder weight parameters;

[0094] Obtain the action prediction data according to the current decoder model, the current style variable, the current decoder weight parameters, the current actual joint position, the current actual end force, the current RGB image, and the current depth image;

[0095] Obtain the mean square error between the action prediction data and the action sequence data;

[0096] Obtain the current reconstruction loss according to the mean square error;

[0097] Obtain the current regularization term according to the current style variable;

[0098] Obtain the current total loss according to the current regularization term and the current reconstruction loss.

[0099] In the embodiments of the present invention, the step of obtaining the current style variable according to the current encoder model, the current encoder weight parameters, the current actual joint position, the current actual end force, and the action sequence data is as follows:

[0100]

[0101] Among them, q represents the current encoder model, is the current encoder weight parameter, a t:t+k is the action sequence data, including the current actual joint position and the current actual end force, is the current style variable.

[0102] The above description of the present invention is an implementation manner for obtaining the current style variable. Another way to obtain the current style variable is also provided in the embodiments of the present invention.

[0103] Figure 3 The flowchart of obtaining the current style variable in the embodiment of the present invention is shown. Refer to Figure 3 As shown, in the embodiment of the present invention, obtaining the current style variable according to the current encoder model, the current encoder weight parameter, the current actual joint position, the current actual end force, and the action sequence data includes:

[0104] Mapping the action sequence data to features of a preset dimension through a first linear layer, and performing sinusoidal position encoding, and then using it as the first feature;

[0105] Mapping the current encoder weight to features of a preset dimension through a second linear layer, and using it as the second feature;

[0106] Mapping the current actual joint position to features of a preset dimension through a third linear layer, and using it as the third feature;

[0107] Mapping the current actual end force to features of a preset dimension through a fourth linear layer, and using it as the fourth feature;

[0108] Obtaining a fifth feature according to the first feature, the second feature, the third feature, and the fourth feature through the current encoder model;

[0109] Mapping the fifth feature through a fifth linear layer to obtain the current style variable.

[0110] In the embodiment of the present invention, the preset dimension can be adjusted according to requirements. For example, it can be 512 dimensions, or it can also be 126 dimensions, etc.

[0111] In the embodiment of the present invention, the second feature can be [CLS].

[0112] Such as Figure 3 As shown, 310 is the first linear layer, 320 is the second linear layer, 330 is the third linear layer, 340 is the fourth linear layer, and 350 is the fifth linear layer.

[0113] 3100 is the action sequence data, and 3101 is the first feature.

[0114] 3200 is the current encoder weight, and 3201 is the second feature.

[0115] 3300 is the current actual joint position, and 3301 is the third feature.

[0116] 3400 is the current actual end force, and 3401 is the fourth feature.

[0117] 300 is the current encoder model, and 3501 is the fifth feature.

[0118] is the current style variable.

[0119] In the embodiment of the present invention, to obtain action prediction data according to the current decoder model, the current style variable, the current decoder weight parameter, the current actual joint position, the current actual end force, the current RGB image, and the current depth image, in the following manner:

[0120]

[0121] wherein, π represents the current decoder model, and θ is the current decoder weight parameter, is the action prediction data, is the current style variable, o t includes the current actual joint position, the current actual end force, the current RGB image, and the current depth image.

[0122] The above description of the present invention is an implementation manner for obtaining action prediction data. In the embodiment of the present invention, another way to obtain action prediction data is also provided.

[0123] Figure 4 The flowchart of obtaining action prediction data according to the embodiment of the present invention is shown. Refer to Figure 4 As shown, to obtain action prediction data according to the current decoder model, the current style variable, the current decoder weight parameter, the current actual joint position, the current actual end force, the current RGB image, and the current depth image, includes:

[0124] According to the number of channels of the current RGB image, update the original depth image to obtain the current depth image;

[0125] Map the current depth image and the current RGB image through a sixth linear layer into features of a preset dimension as the sixth feature;

[0126] Map the current actual joint position through a seventh linear layer into features of a preset dimension as the seventh feature;

[0127] Map the current actual end - effector force to a feature of a preset dimension through an eighth linear layer as the eighth feature;

[0128] Map the current style variable to a feature of a preset dimension through a ninth linear layer as the ninth feature;

[0129] Through the current encoder model, obtain the action prediction data according to the sixth feature, the seventh line, the eighth feature, and the ninth feature.

[0130] In an embodiment of the present invention, the sixth feature may include the sixth feature corresponding to the current RGB image and the sixth feature corresponding to the current depth image.

[0131] Figure 4 Among them, 460 is the sixth linear layer, 470 is the seventh linear layer, 480 is the eighth linear layer, and 490 is the ninth linear layer.

[0132] 4101 is the original depth image, 410 is the current depth image, and 420 is the current RGB image.

[0133] 4600 - RGB is the sixth feature corresponding to the current RGB image, and 4600 - Depth is the sixth feature corresponding to the current depth image.

[0134] 4700 is the current actual joint position, 4800 is the current actual end - effector force, 4900 is the current style variable, 400 is the decoder, and 4000 is the action prediction data.

[0135] In an embodiment of the present invention, obtaining the current reconstruction loss according to the mean square error is performed in the following manner:

[0136]

[0137] Obtaining the current regularization term according to the current style variable is performed in the following manner:

[0138]

[0139] Obtaining the current total loss according to the current regularization term and the current reconstruction loss is performed in the following manner:

[0140]

[0141] Among them, is the current reconstruction loss, is the action prediction data, a t:t+k is the action sequence data, is the current regularization term, represents the standard normal distribution, D KLrepresents the KL divergence, q represents the current encoder model, is the current encoder weight parameter, is the current style variable, including the current actual joint position and the current actual end force, is the current total loss.

[0142] In an embodiment of the present invention, the updating the original depth image according to the number of channels of the current RGB image to obtain the current depth image includes:

[0143] Copy the original depth image M times;

[0144] Stitch the M copied original depth images according to the channels to obtain the current depth image;

[0145] where M is the number of channels of the current RGB image.

[0146] In an embodiment of the present invention, the depth image is single-channel and the RGB image is three-channel, so M can be 3.

[0147] For example, the size of the original depth map is (640, 480, 1), which is copied three times and stitched according to the channels, so as to obtain the same size as the RGB image (640, 480, 3), where 1 and 3 represent the number of channels.

[0148] The method of the embodiment of the present invention can make the imitation learning model closer to human movement habits, and can also make the control of the humanoid robot more precise and detailed.

[0149] As Figure 5 shown, the present invention also provides a humanoid robot imitation learning device, and the device includes:

[0150] An acquisition module 510, configured to acquire training data;

[0151] A training module 520, configured to train an initial model according to the training data to obtain an imitation learning model;

[0152] An execution module 530, configured to make a target humanoid robot execute an action according to the imitation learning model;

[0153] where the training data is data collected with the imitation robot as the center when the imitation robot imitates the original object, and the training data includes the actual joint position and the actual end force.

[0154] In an embodiment of the present invention, the training module 520 is further configured to:

[0155] Input the current training data into the initial model to obtain the current total loss;

[0156] Determine whether the current total loss satisfies the convergence condition;

[0157] If the convergence condition is satisfied, end the training, and use the initial model at the end of the training as the imitation learning model;

[0158] If the convergence condition is not satisfied, continue to obtain the next training data.

[0159] In an embodiment of the present invention, the current training data includes: the current actual joint position, the current actual end force, the current RGB image, the current depth image, and the action sequence data;

[0160] The training module 520 is further configured to:

[0161] Update the previous encoder weight parameters according to the previous total loss to obtain the current encoder weight parameters;

[0162] Obtain the current style variable according to the current encoder model, the current encoder weight parameters, the current actual joint position, the current actual end force, and the action sequence data;

[0163] Update the previous decoder weight parameters according to the previous total loss to obtain the current decoder weight parameters;

[0164] Obtain the action prediction data according to the current decoder model, the current style variable, the current decoder weight parameters, the current actual joint position, the current actual end force, the current RGB image, and the current depth image;

[0165] Obtain the mean square error between the action prediction data and the action sequence data;

[0166] Obtain the current reconstruction loss according to the mean square error;

[0167] Obtain the current regularization term according to the current style variable;

[0168] Obtain the current total loss according to the current regularization term and the current reconstruction loss.

[0169] In an embodiment of the present invention, the training module 520 is further configured to obtain the current style variable according to the following method, according to the current encoder model, the current encoder weight parameters, the current actual joint position, the current actual end force, and the action sequence data:

[0170]

[0171] where q represents the current encoder model, is the current encoder weight parameter, at:t+k is action sequence data, including the current actual joint position and the current actual end - effector force, is the current style variable.

[0172] In the embodiment of the present invention, the training module 520 is further configured to:

[0173] Map the action sequence data to features of a preset dimension through a first linear layer, and after performing sine - position encoding, use it as the first feature;

[0174] Map the current encoder weight through a second linear layer to features of a preset dimension, and use it as the second feature;

[0175] Map the current actual joint position through a third linear layer to features of a preset dimension, and use it as the third feature;

[0176] Map the current actual end - effector force through a fourth linear layer to features of a preset dimension, and use it as the fourth feature;

[0177] Through the current encoder model, obtain a fifth feature according to the first feature, the second feature, the third feature, and the fourth feature;

[0178] Map the fifth feature through a fifth linear layer to obtain the current style variable.

[0179] In the embodiment of the present invention, the training module 520 is further configured to obtain action prediction data according to the following method, based on the current decoder model, the current style variable, the current decoder weight parameter, the current actual joint position, the current actual end - effector force, the current RGB image, and the current depth image:

[0180]

[0181] where, π represents the current decoder model, θ is the current decoder weight parameter, is the action prediction data, is the current style variable, o t including the current actual joint position, the current actual end - effector force, the current RGB image, and the current depth image.

[0182] In the embodiment of the present invention, the training module 520 is further configured to:

[0183] Obtain the current RGB image according to the current depth image;

[0184] Map the current depth image and the current RGB image through a sixth linear layer to features of a preset dimension, and use it as the sixth feature;

[0185] Map the current actual joint position into features of a preset dimension through a seventh linear layer as the seventh feature;

[0186] Map the current actual end - effector force into features of a preset dimension through an eighth linear layer as the eighth feature;

[0187] Map the current style variable into features of a preset dimension through a ninth linear layer as the ninth feature;

[0188] Through the current encoder model, obtain the action prediction data according to the sixth feature, the seventh feature, the eighth feature, and the ninth feature.

[0189] In an embodiment of the present invention, the training module 520 is further configured to obtain the current reconstruction loss according to the following method based on the mean square error:

[0190]

[0191] The training module 520 is further configured to obtain the current regularization term according to the following method based on the current style variable:

[0192]

[0193] The training module 520 is further configured to obtain the current total loss according to the following method based on the current regularization term and the current reconstruction loss:

[0194]

[0195] Where, is the current reconstruction loss, is the action prediction data, a t:t+k is the action sequence data, is the current regularization term, represents the standard normal distribution, D KL represents the KL divergence, q represents the current encoder model, is the current encoder weight parameter, is the current style variable, includes the current actual joint position and the current actual end - effector force, is the current total loss.

[0196] In an embodiment of the present invention, the training module 520 is further configured to:

[0197] Copy the original depth image M times;

[0198] Stitch the M copied original depth images according to channels to obtain the current depth image;

[0199] Wherein, M is the number of channels of the current RGB image.

[0200] The method according to the embodiments of the present invention can make the imitation learning model closer to human movement habits, and can also make the control of the humanoid robot more accurate and detailed.

[0201] The embodiments of the present invention also provide a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following method is implemented: obtaining training data; training an initial model according to the training data to obtain an imitation learning model; making a target humanoid robot execute an action according to the imitation learning model; wherein, the training data is data collected centered on the imitation robot when the imitation robot imitates the original object, and the training data includes actual joint positions and actual end forces.

[0202] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following method is implemented: obtaining training data; training an initial model according to the training data to obtain an imitation learning model; making a target humanoid robot execute an action according to the imitation learning model; wherein, the training data is data collected centered on the imitation robot when the imitation robot imitates the original object, and the training data includes actual joint positions and actual end forces.

[0203] The above humanoid robot imitation learning method achieves the beneficial effect of being able to solve the technical problems proposed in the background art.

[0204] Figure 2 It is a schematic flowchart of a humanoid robot imitation learning method in an embodiment. It should be understood that although Figure 2 the steps in the flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figure 2 at least a part of the steps in

[0205] Figure 6 can include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps. Figure 1Server 120 therein. As Figure 6 shown, the computer device includes a processor, a memory, a network interface, an input device, and a display screen connected via a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor can implement the imitation learning method of the humanoid robot. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor can implement the imitation learning method of the humanoid robot. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the computer device, or an external keyboard, a touchpad, or a mouse, etc.

[0206] Those skilled in the art can understand that Figure 6 the structure shown in is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0207] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or an external cache. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0208] It should be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.

[0209] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.

Claims

1. A humanoid robot imitation learning method, characterized in that: The method comprises: Get training data; Training an initial model based on the training data to obtain an imitation learning model; According to the imitation learning model, causing the target humanoid robot to perform actions; The training data is data collected with the imitation robot as the center when the imitation robot imitates the original object, and the training data includes actual joint positions and actual end forces.

2. The method according to claim 1, characterized in that The step of training the initial model according to the training data to obtain the imitation learning model comprises: Input the current training data into the initial model to obtain the current total loss; Determine whether the current total loss meets the convergence condition; If the convergence condition is met, the training is terminated, and the initial model after the training is completed is used as the imitation learning model; If the convergence condition is not met, continue to obtain the next training data.

3. The method according to claim 2, characterized in that The current training data includes: current actual joint position, current actual end force, current RGB image, current depth image and action sequence data; The current training data is input into the initial model to obtain the current total loss, including: According to the previous total loss, update the weight parameters of the previous encoder to obtain the weight parameters of the current encoder; Acquire a current style variable according to a current encoder model, the current encoder weight parameter, the current actual joint position, the current actual end force and the action sequence data; According to the previous total loss, update the weight parameters of the previous decoder to obtain the weight parameters of the current decoder; Acquire action prediction data according to the current decoder model, the current style variable, the current decoder weight parameter, the current actual joint position, the current actual end force, the current RGB image and the current depth image; Obtain the mean square error between the action prediction data and the action sequence data; Get the current reconstruction loss based on the mean square error; According to the current style variable, obtain a current regularization item; The current total loss is obtained according to the current regularization term and the current reconstruction loss.

4. The method according to claim 3, characterized in that The current style variable is obtained according to the current encoder model, the current encoder weight parameter, the current actual joint position, the current actual end force and the action sequence data, in the following manner: Among them, q represents the current encoder model, is the current encoder weight parameter, a t:t+k is the action sequence data, Including the current actual joint position and the current actual end force, The current style variable.

5. The method according to claim 3, characterized in that: The obtaining of the current style variable according to the current encoder model, the current encoder weight parameter, the current actual joint position, the current actual end force and the action sequence data comprises: Mapping the action sequence data into features of a preset dimension through a first linear layer, and performing forward rotation position encoding as the first feature; Mapping the current encoder weight into a feature of a preset dimension through a second linear layer as a second feature; Mapping the current actual joint position into a feature of a preset dimension through a third linear layer as a third feature; Mapping the current actual end force into a feature of a preset dimension through a fourth linear layer as a fourth feature; Obtaining a fifth feature according to the first feature, the second feature, the third feature, and the fourth feature through the current encoder model; The fifth feature is mapped through a fifth linear layer to obtain the current style variable.

6. The method according to claim 3, characterized in that: The obtaining of action prediction data according to the current decoder model, the current style variable, the current decoder weight parameter, the current actual joint position, the current actual end force, the current RGB image and the current depth image is performed in the following manner: Among them, π represents the current decoder model, θ is the current decoder weight parameter, is the action prediction data, is the current style variable, o t Includes the current actual joint position, the current actual end force, the current RGB image, and the current depth image.

7. The method according to claim 3, characterized in that The step of acquiring action prediction data according to the current decoder model, the current style variable, the current decoder weight parameter, the current actual joint position, the current actual end force, the current RGB image, and the current depth image includes: Update the original depth image according to the number of channels of the current RGB image to obtain the current depth image; Mapping the current depth image and the current RGB image into features of a preset dimension through a sixth linear layer as a sixth feature; Mapping the current actual joint position into a feature of a preset dimension through a seventh linear layer as a seventh feature; Mapping the current actual end force into a feature of a preset dimension through an eighth linear layer as an eighth feature; Mapping the current style variable into a feature of a preset dimension through a ninth linear layer as a ninth feature; The action prediction data is obtained through the current encoder model according to the sixth feature, the seventh line, the eighth feature and the ninth feature.

8. The method according to claim 3, characterized in that The current reconstruction loss is obtained according to the mean square error in the following manner: The current regularization term is obtained according to the current style variable in the following manner: According to the current regularization term and the current reconstruction loss, the current total loss is obtained in the following manner: in, is the current reconstruction loss, is the action prediction data, a t:t+k is the action sequence data, is the current regularization term, represents the standard normal distribution, D KL represents the KL divergence, q represents the current encoder model, is the current encoder weight parameter, is the current style variable, Including the current actual joint position and the current actual end force, is the current total loss.

9. The method according to claim 6, characterized in that The updating of the original depth image according to the number of channels of the current RGB image to obtain the current depth image includes: Copying the original depth image M times; splicing the M copies of the original depth image according to the channels to obtain the current depth image; Wherein, M is the number of channels of the current RGB image.

10. A humanoid robot imitation learning device, characterized in that: The device comprises: An acquisition module, used to obtain training data; A training module, used to train an initial model according to the training data to obtain an imitation learning model; An execution module, used to enable the target humanoid robot to perform actions according to the imitation learning model; The training data is data collected with the imitation robot as the center when the imitation robot imitates the original object, and the training data includes actual joint positions and actual end forces.

11. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 9 is implemented.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Machine translation model training method and device, electronic equipment and storage medium

    CN115600613A

  • Joint angle control method, device and model, computer equipment and storage medium

    CN117885094A

  • Self-adaptive neural network control method for hydraulic mechanical arm

    CN118848994A

  • Robot imitation learning system and method based on multi-modal somatosensory data

    CN119260729A

  • Mitigating reality gap through simulating compliant control and / or compliant contact in robotic simulator

    US20210107157A1

Cited By

  • Compensation method for robot imitation learning training process

    CN121028677A

  • A compensation method for a robot imitation learning training process

    CN121028677B

  • Humanoid robot motion control model training method and device based on imitation learning

    CN121403415A

  • Humanoid robot motion control model training method and device based on imitation learning

    CN121403415B