Humanoid robot imitation learning data acquisition method and device, computer equipment and storage medium

By splitting the imitation learning task into sub-tasks and matching it with the primitive library, and acquiring and fusing data, the problems of high data acquisition cost and insufficient authenticity in the existing technology are solved, and efficient and low-cost data acquisition is achieved, taking into account the authenticity and accuracy of the data.

CN119974005APending Publication Date: 2025-05-13KEPLER ROBOT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510320979.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, humanoid robot imitation learning data acquisition methods cannot take into account both authenticity and low cost, low data quality or inaccurate demonstration may lead to the model learning incorrect strategies.

Method used

By splitting the imitation learning task into multiple subtasks and matching it with the primitive tasks in the primitive library, the first and second types of data are obtained and fused as learning data for the imitation learning task.

Benefits of technology

It reduces the tasks that humanoid robots need to be demonstrated, shortens the time for data acquisition, reduces the cost, and takes into account the authenticity and accuracy of the data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119974005A_ABST
    Figure CN119974005A_ABST
Patent Text Reader

Abstract

The invention relates to a humanoid robot imitation learning data acquisition method and device, computer equipment and a storage medium. The method comprises the following steps: generating an imitation learning task; splitting the imitation learning task into a plurality of sub-tasks; judging whether each subtask is matched with a primitive task in a primitive library or not, taking the subtask matched with any primitive task as a first type of subtask, and taking the subtasks which cannot be matched with any primitive task as a second type of subtask; obtaining a first type of data corresponding to the first type of subtasks from the primitive library; acquiring a second type of data corresponding to the second type of subtasks; and according to the sequence of the plurality of sub-tasks, fusing the first type of data and the second type of data to serve as learning data corresponding to the imitation learning task. The data collection time can be shortened, and the cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data collection, and in particular to a method, device, computer equipment and storage medium for collecting data for imitation learning of a humanoid robot. Background Art

[0002] Imitation learning, imitation learning data collection and imitation model building are the basic methods for humanoid robots to imitate human behavior. Among them, imitation learning is that humanoid robots learn and complete tasks by imitating the behavior of experts, imitation learning data collection is to collect various data in the process of humanoid robots imitating expert behavior, and imitation model building is to train the model through the collected imitation learning data.

[0003] The training results of the imitation model depend on the collection of imitation learning data. If the data quality is not high or the demonstration is inaccurate, the model may learn the wrong strategy.

[0004] In the prior art, there are two main ways to collect robot imitation learning data. One is to rely on humanoid robots to imitate expert behavior, that is, the robot needs to operate according to the expert behavior. This method has the disadvantages of high data collection cost, long time consumption, and difficulty in covering sufficient environment and task diversity. The second is virtual synthetic data, which mainly generates virtual synthetic data. This method is fast and low-cost, does not require real robot imitation behavior, and does not require a real environment. Especially in high-risk tasks, it can avoid the risk of hardware damage, but the synthetic data may have errors with the real imitation behavior, and the data distribution in the simulation environment is also different from the actual environment.

[0005] It can be seen that the data collection method for robot imitation learning in the prior art cannot take into account both authenticity and low cost. Summary of the invention

[0006] In order to solve the above technical problems or at least partially solve the above technical problems, the present invention provides a method, device, computer equipment and storage medium for collecting data for imitation learning of a humanoid robot.

[0007] In a first aspect, the present invention provides a method for collecting data from imitation learning of a humanoid robot, the method comprising:

[0008] Generate imitation learning tasks;

[0009] Splitting the imitation learning task into multiple subtasks;

[0010] Determine whether each of the subtasks matches a primitive task in the primitive library, and regard the subtasks that match any primitive task as first-category subtasks, and regard the subtasks that cannot match any primitive task as second-category subtasks;

[0011] Acquire, from the primitive library, first-category data corresponding to the first-category subtask;

[0012] Obtaining the second type of data corresponding to the second type of subtask;

[0013] According to the order of the multiple subtasks, the first category of data and the second category of data are fused as learning data corresponding to the imitation learning task.

[0014] Optionally, if the second-type subtask is a non-contact task, the acquiring the second-type data corresponding to the second-type subtask includes:

[0015] Obtaining the starting position, the ending position and the humanoid robot parameters of the second type of subtask;

[0016] Generate a trajectory from the starting position to the ending position according to the starting position, the ending position and the humanoid robot parameters;

[0017] According to the trajectory, the corresponding second-category data is generated.

[0018] Optionally, the obtaining the second type of data corresponding to the second type of subtask includes:

[0019] enabling the humanoid robot to complete the second type of subtask;

[0020] Collect motion data of the humanoid robot in the process of completing the second type of subtask, and use the motion data as the corresponding second type of data.

[0021] Optionally, the primitive library includes a plurality of primitive tasks, and primitive data corresponding to each of the primitive tasks.

[0022] After collecting the motion data of the humanoid robot in the process of completing the second type of subtask, the method further includes:

[0023] The second type of subtask and the corresponding action data are placed in the primitive library as a primitive task of the primitive library and the corresponding primitive data.

[0024] Optionally, before generating a plurality of imitation learning tasks, the method further includes:

[0025] Acquire a primitive task, wherein the primitive task includes environmental information;

[0026] placing the humanoid robot in a virtual environment matching the environmental information of the primitive task;

[0027] enabling the humanoid robot to complete the primitive task in the virtual environment;

[0028] Obtaining initial data corresponding to the primitive task;

[0029] The initial data is generalized according to the random variables, and the generalized initial data is used as primitive data.

[0030] Optionally, a camera is provided on the head of the humanoid robot, the camera includes an RGB camera and a depth camera, the RGB camera is used to collect an RGB image from a main perspective of the humanoid robot, and the depth camera is used to collect a depth image from the main perspective of the humanoid robot;

[0031] The obtaining of initial data corresponding to the primitive task includes:

[0032] According to the camera, initial data of the main perspective of the robot is collected during the process of the humanoid robot completing the primitive task.

[0033] Optionally, the initial data includes:

[0034] An RGB image of the main perspective of the humanoid robot;

[0035] A depth image of the main perspective of the humanoid robot;

[0036] External control instructions for each joint of the humanoid robot;

[0037] Real-time status feedback data of each joint of the robot;

[0038] Timestamp.

[0039] Optionally, after fusing the first category of data and the second category of data according to the order of the multiple subtasks as learning data corresponding to the imitation learning task, the method further includes:

[0040] Get the transformation matrix between the robot coordinate system and the world coordinate system;

[0041] Convert the robot's main perspective learning data into world coordinate system learning data.

[0042] In a second aspect, a humanoid robot imitation learning data acquisition device is provided, the device comprising:

[0043] A task generation unit, used to generate imitation learning tasks;

[0044] A splitting unit, used for splitting the imitation learning task into multiple subtasks;

[0045] A judging unit, used for judging whether each of the subtasks matches a primitive task in the primitive library, treating a subtask that matches any primitive task as a first-category subtask, and treating a subtask that cannot match any primitive task as a second-category subtask;

[0046] A data acquisition unit, configured to acquire, from the primitive library, first-category data corresponding to the first-category subtask;

[0047] The data acquisition unit is further used to acquire the second type of data corresponding to the second type of subtask;

[0048] A fusion unit is used to fuse the first category of data and the second category of data according to the order of the multiple subtasks as learning data corresponding to the imitation learning task.

[0049] According to a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above methods when executing the computer program.

[0050] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the method as described in any one of the above items is implemented.

[0051] The present invention relates to a humanoid robot imitation learning data collection method, device, computer equipment and storage medium, the method comprising: generating an imitation learning task; splitting the imitation learning task into multiple subtasks; judging whether each of the subtasks matches a primitive task in a primitive library, taking the subtasks that match any primitive task as the first type of subtask, and taking the subtasks that cannot match any primitive task as the second type of subtask; obtaining the first type of data corresponding to the first type of subtask from the primitive library; obtaining the second type of data corresponding to the second type of subtask; according to the order of the multiple subtasks, fusing the first type of data and the second type of data as the learning data corresponding to the imitation learning task. In the method of the embodiment of the present invention, the imitation learning task is split into multiple subtasks, and the first type of subtask and the second type of subtask are distinguished by comparing with the primitive tasks in the primitive library, and the first type of data and the second type of data are fused as the learning data. Since the first type of subtask is a subtask that can be matched with any primitive task, and the primitive library contains the first type of data corresponding to the first type of subtask, the method of the present invention can reduce the tasks that the humanoid robot needs to demonstrate, shorten the time for collecting data, and reduce costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0053] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0054] Figure 1 The figure shows an application environment diagram of the humanoid robot imitation learning data collection method according to an embodiment of the present invention;

[0055] Figure 2 It is a schematic diagram of the flow chart of the humanoid robot imitation learning data collection method according to an embodiment of the present invention;

[0056] Figure 3 Shown is a schematic diagram of the main perspective of a robot according to an embodiment of the present invention;

[0057] Figure 4 Shown is a schematic diagram of an imitation learning task according to an embodiment of the present invention;

[0058] Figure 5 FIG. 1 is a structural block diagram of a humanoid robot imitation learning data acquisition device according to an embodiment of the present invention;

[0059] Figure 6 1 is a diagram showing the internal structure of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION

[0060] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0061] Figure 1 FIG. 1 is a diagram showing an application environment of a humanoid robot imitation learning data collection method in one embodiment. Figure 1 , the humanoid robot imitation learning data collection method is applied to a humanoid robot imitation learning data collection system. The humanoid robot imitation learning data collection method includes a terminal 110 (and / or a server 120). The terminal 110 and the server 120 are connected via a network. The terminal 110 can be a desktop terminal or a mobile terminal, and the mobile terminal can be at least one of a mobile phone, a tablet computer, a laptop computer, etc. The server 120 can be implemented by an independent server or a server cluster composed of multiple servers.

[0062] The terminal 110 and / or the server 120 may be arranged inside the humanoid robot, outside the humanoid robot, or partially inside the humanoid robot and partially outside the humanoid robot.

[0063] The humanoid robot imitation learning data collection method of the present invention is applied to the terminal 110 and / or the server 120 .

[0064] Figure 2 FIG. 1 is a flow chart of a method for collecting data through imitation learning of a humanoid robot according to an embodiment of the present invention. Figure 2 As shown, the method includes:

[0065] Step 210, generating an imitation learning task;

[0066] Step 220, splitting the imitation learning task into multiple subtasks;

[0067] Step 230, determining whether each of the subtasks matches a primitive task in the primitive library, treating the subtasks that match any primitive task as first-category subtasks, and treating the subtasks that cannot match any primitive task as second-category subtasks;

[0068] Step 240, obtaining the first type of data corresponding to the first type of subtask from the primitive library;

[0069] Step 250, obtaining the second type of data corresponding to the second type of subtask;

[0070] Step 260: According to the order of the multiple subtasks, the first category of data and the second category of data are integrated as learning data corresponding to the imitation learning task.

[0071] In an embodiment of the present invention, the imitation learning task may include environmental information, task starting position, task ending position, action, action object, etc.; wherein the environmental information may include friction coefficient, gravity, background, interference, distance between the object and the robot, etc.

[0072] In the method of the embodiment of the present invention, the primitive library includes primitive tasks and primitive data corresponding to the primitive tasks. The primitive tasks may be the behavior of a humanoid robot imitating an expert, and the primitive data are data actually collected during the imitation process.

[0073] The more imitation learning tasks used for subsequent humanoid robot training, the more accurate the subsequent model will be and the more application scenarios it can adapt to. However, there are too many imitation learning tasks generated, and each imitation learning task requires the humanoid robot to demonstrate it from beginning to end and collect data at the same time, which will result in too many tasks for the humanoid robot to demonstrate, and the data collection process will be too long. In addition, the large number of tasks demonstrated by the humanoid robot will also cause wear and tear on the humanoid robot, resulting in huge material and time costs.

[0074] In the method of the embodiment of the present invention, the imitation learning task is divided into multiple subtasks, and the first type of subtasks and the second type of subtasks are distinguished by comparing with the primitive tasks in the primitive library, and the first type of data and the second type of data are fused as learning data. Since the first type of subtask is a subtask that can match any primitive task, and the primitive library has the first type of data corresponding to the first type of subtask, the method of the present invention can reduce the tasks that the humanoid robot needs to demonstrate, shorten the time for collecting data, and reduce costs. In the embodiment of the present invention, the primitive task can be the behavior of the humanoid robot imitating the expert, and the primitive data is the data actually collected during the imitation process. Therefore, the method of the embodiment of the present invention can also take into account both authenticity and accuracy.

[0075] In the embodiment of the present invention, if the second type of subtask is a non-contact task, then obtaining the second type of data corresponding to the second type of subtask includes:

[0076] Obtaining the starting position, the ending position and the humanoid robot parameters of the second type of subtask;

[0077] Generate a trajectory from the starting position to the ending position according to the starting position, the ending position and the humanoid robot parameters;

[0078] According to the trajectory, the corresponding second-category data is generated.

[0079] In the embodiment of the present invention, a non-contact task refers to a task in which the humanoid robot does not need to make any additional contact with other objects from the start time to the end time of the task, for example, the manipulator of the humanoid robot moves from one position to another.

[0080] New contact refers to the situation where the humanoid robot does not put down or pick up objects from the start time to the end time of the task. For example, the humanoid robot's manipulator does not grab any object and moves from one position to another, which can be called a non-contact task; and the humanoid robot's manipulator grabs an object and moves it from one position to another without putting down the object, which can also be called a non-contact task.

[0081] For non-contact tasks, it is equivalent to a trajectory movement task, so the second type of data can be directly generated based on the trajectory and robot parameters.

[0082] In the embodiment of the present invention, the step of obtaining the second type of data corresponding to the second type of subtask includes:

[0083] enabling the humanoid robot to complete the second type of subtask;

[0084] Collect motion data of the humanoid robot in the process of completing the second type of subtask, and use the motion data as the corresponding second type of data.

[0085] In the embodiment of the present invention, the primitive library includes a plurality of primitive tasks and primitive data corresponding to each of the primitive tasks.

[0086] After collecting the motion data of the humanoid robot in the process of completing the second type of subtask, the method further includes:

[0087] The second type of subtask and the corresponding action data are placed in the primitive library as a primitive task of the primitive library and the corresponding primitive data.

[0088] In an embodiment of the present invention, the second type of subtask is a subtask not included in the primitive library, which enables the humanoid robot to complete the second type of subtask and collect data. After collecting the data, the second type of subtask is placed in the primitive library for use in other imitation learning tasks.

[0089] In the embodiment of the present invention, before generating a plurality of imitation learning tasks, the method further includes:

[0090] Acquire a primitive task, wherein the primitive task includes environmental information;

[0091] placing the humanoid robot in a virtual environment matching the environmental information of the primitive task;

[0092] enabling the humanoid robot to complete the primitive task in the virtual environment;

[0093] Obtaining initial data corresponding to the primitive task;

[0094] The initial data is generalized according to the random variables, and the generalized initial data is used as primitive data.

[0095] Random variables can refer to environmental random variables, such as slight changes in friction, wind force, and the flatness of the ground, etc.; they can also refer to some changes in the robot's own drive motor, such as accuracy, driving force, etc. Through random variables, data can be generalized, so that the generalized data has more application scenarios.

[0096] In an embodiment of the present invention, in the initial stage, a humanoid robot can be used to imitate expert behavior for demonstration and establish a primitive library. In the subsequent imitation learning data collection process, the primitive library is expanded so that more subtasks for collecting imitation learning data can find corresponding subtasks in the primitive library, thereby reducing the time consumption and robot consumption of data collection.

[0097] In the embodiment of the present invention, the primitive task may be a humanoid robot imitating the behavior of an expert, and the primitive data is the data actually collected during the imitation process. Therefore, the method of the embodiment of the present invention can also take into account both authenticity and accuracy.

[0098] In an embodiment of the present invention, a camera is provided on the head of the humanoid robot, and the camera includes an RGB camera and a depth camera, the RGB camera is used to collect an RGB image of the main perspective of the humanoid robot, and the depth camera is used to collect a depth image of the main perspective of the humanoid robot;

[0099] The obtaining of initial data corresponding to the primitive task includes:

[0100] According to the camera, initial data of the main perspective of the robot is collected during the process of the humanoid robot completing the primitive task.

[0101] In the method of the embodiment of the present invention, a camera is provided on the head of the humanoid robot, so that the primitive tasks collected are tasks from the main perspective of the humanoid robot. The humanoid robot is an imitation of humans. Collecting data with the main vision of the humanoid robot can make the collected data closer to human action habits. For example, when a person takes an object from a table, he usually needs to visually see the object, that is, he needs to visually see the hand and the object. The posture of the limbs of the human body at this time is not necessarily exactly the same as when the object is not visually seen. Therefore, collecting data with the main vision of the humanoid robot can make the joint angles of the humanoid robot, the trajectory of the end of the robotic arm, the driving method of the robotic arm motor, etc. closer to human habits.

[0102] In the embodiment of the present invention, the initial data includes:

[0103] An RGB image of the main perspective of the humanoid robot;

[0104] A depth image of the main perspective of the humanoid robot;

[0105] External control instructions for each joint of the humanoid robot;

[0106] Real-time status feedback data of each joint of the robot;

[0107] Timestamp.

[0108] In the method of the embodiment of the present invention, after fusing the first type of data and the second type of data according to the order of the multiple subtasks as the learning data corresponding to the imitation learning task, the method further includes:

[0109] Get the transformation matrix between the robot coordinate system and the world coordinate system;

[0110] Convert the robot's main perspective learning data into world coordinate system learning data.

[0111] In the method of the present invention, the humanoid robot's primary vision is used to collect data, but the learning data of the embodiment of the present invention can also be converted into learning data of the world coordinate system, which can expand the application scenarios of the method of the embodiment of the present invention.

[0112] Figure 3 It is a schematic diagram of the main perspective of the robot according to an embodiment of the present invention, and a camera is arranged on the head of the humanoid robot.

[0113] like Figure 3 As shown on the left, the robot is in a virtual environment. Figure 3 The right side shows the image of the humanoid robot from the main perspective in the collected data. If the humanoid robot imitates humans and has two cameras on the left and right, the image collected from the main perspective is as follows: Figure 3 The two images shown on the right can be considered as the left camera image and the right camera image.

[0114] Figure 4 Shown is a schematic diagram of an imitation learning task according to an embodiment of the present invention.

[0115] In the embodiment of the present invention, an imitation learning task may be that a humanoid robot uses a mechanical right hand to pick up an object from table A and put it on table B. The friction coefficient between the table and the object is a, the mass of the object is M, and the gravity is g.

[0116] The imitation learning task described above can be divided into multiple subtasks, which can be:

[0117] Subtask 1: The robot right hand moves from the initial position of the task to the preset position above the table A;

[0118] Subtask 2: Open the mechanical right hand;

[0119] Subtask 3: The robot's right hand moves to the vicinity of the object on the table A;

[0120] Subtask 4: The mechanical right hand grasps the object;

[0121] Subtask 5: The robot's right hand grabs the object and moves it to the preset position above table B;

[0122] Subtask 6: The robotic right hand places the object on table B;

[0123] Subtask 7: Open the mechanical right hand;

[0124] Subtask 8: The mechanical right hand moves to the task end position.

[0125] In the above subtasks, if the initial position of the task is the preset position above desktop A, then subtask 1 can be omitted.

[0126] Among the above subtasks, subtask 1, subtask 3, subtask 5, subtask 6 and subtask 8 are non-contact tasks, involving only the movement of the trajectory. Subtask 2 can be simply called opening, subtask 4 can be simply called grasping or closing, and subtask 7 can be simply called opening.

[0127] Therefore, the above task can be considered to consist of two primitive actions, namely opening and grasping, and multiple non-contact trajectory movement subtasks.

[0128] Opening and grasping can obtain corresponding primitive data from the primitive library; trajectory movement tasks can have corresponding primitive data in the primitive library. If there is no corresponding primitive data, the second type of data can be generated according to the starting position and the end position. Therefore, the above tasks can be generated by fusing primitive data and / or second type of data.

[0129] The method of the embodiment of the present invention can reduce the time cost and economic cost of data collection.

[0130] Figure 4 The following is a schematic diagram of the subtasks: Figure 4 (a) shows a non-contact trajectory movement task. Figure 4 (b) to open, Figure 4 (c) for crawling, Figure 4 (d) Open.

[0131] Figure 4 The following is a schematic diagram of the subtasks of the dual-arm grasping task. Figure 4 Each subtask consists of two images, the left image is an RGB image and the right image is a depth image.

[0132] Figure 5 FIG. 1 is a structural block diagram of a humanoid robot imitation learning data acquisition device according to an embodiment of the present invention. Figure 5 As shown, the device comprises:

[0133] A task generating unit 510, used to generate an imitation learning task;

[0134] A splitting unit 520, used for splitting the imitation learning task into multiple subtasks;

[0135] A judging unit 530 is used to judge whether each of the subtasks matches a primitive task in the primitive library, and to regard a subtask that matches any primitive task as a first-category subtask, and to regard a subtask that cannot match any primitive task as a second-category subtask;

[0136] A data acquisition unit 540 is used to acquire the first type of data corresponding to the first type of subtask from the primitive library;

[0137] The data acquisition unit 540 is further used to acquire the second type of data corresponding to the second type of subtask;

[0138] The fusion unit 550 is used to fuse the first category of data and the second category of data according to the order of the multiple subtasks as learning data corresponding to the imitation learning task.

[0139] In the embodiment of the present invention, if the second type of subtask is a non-contact task, the data acquisition unit 540 is further configured to:

[0140] Obtaining the starting position, the ending position and the humanoid robot parameters of the second type of subtask;

[0141] Generate a trajectory from the starting position to the ending position according to the starting position, the ending position and the humanoid robot parameters;

[0142] According to the trajectory, the corresponding second-category data is generated.

[0143] In the embodiment of the present invention, the data acquisition unit 540 is further used for:

[0144] enabling the humanoid robot to complete the second type of subtask;

[0145] Collect motion data of the humanoid robot in the process of completing the second type of subtask, and use the motion data as the corresponding second type of data.

[0146] In the embodiment of the present invention, the primitive library includes a plurality of primitive tasks and primitive data corresponding to each of the primitive tasks.

[0147] The data acquisition unit 540 is also used to: after collecting the action data of the humanoid robot in the process of completing the second type of subtask, put the second type of subtask and the corresponding action data into the primitive library as a primitive task of the primitive library and the corresponding primitive data.

[0148] In an embodiment of the present invention, the apparatus further comprises a primitive library, which is used for:

[0149] Acquire a primitive task, wherein the primitive task includes environmental information;

[0150] placing the humanoid robot in a virtual environment matching the environmental information of the primitive task;

[0151] enabling the humanoid robot to complete the primitive task in the virtual environment;

[0152] Obtaining initial data corresponding to the primitive task;

[0153] The initial data is generalized according to the random variables, and the generalized initial data is used as primitive data.

[0154] In an embodiment of the present invention, a camera is provided on the head of the humanoid robot, and the camera includes an RGB camera and a depth camera, the RGB camera is used to collect an RGB image of the main perspective of the humanoid robot, and the depth camera is used to collect a depth image of the main perspective of the humanoid robot;

[0155] The primitive library is also used for:

[0156] According to the camera, initial data of the main perspective of the robot is collected during the process of the humanoid robot completing the primitive task.

[0157] In the embodiment of the present invention, the initial data includes:

[0158] An RGB image of the main perspective of the humanoid robot;

[0159] A depth image of the main perspective of the humanoid robot;

[0160] External control instructions for each joint of the humanoid robot;

[0161] Real-time status feedback data of each joint of the robot;

[0162] Timestamp.

[0163] The apparatus of the embodiment of the present invention further includes a transformation unit, which is used to fuse the first type of data and the second type of data according to the order of the multiple subtasks as the learning data corresponding to the imitation learning task:

[0164] Get the transformation matrix between the robot coordinate system and the world coordinate system;

[0165] Convert the robot's main perspective learning data into world coordinate system learning data.

[0166] The embodiments of the present invention can reduce the time cost and economic cost of data collection.

[0167] An embodiment of the present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following method when executing the computer program: generating an imitation learning task; splitting the imitation learning task into multiple subtasks; determining whether each of the subtasks matches a primitive task in a primitive library, treating the subtasks that match any primitive task as first-category subtasks, and treating the subtasks that cannot match any primitive task as second-category subtasks; obtaining first-category data corresponding to the first-category subtasks from the primitive library; obtaining second-category data corresponding to the second-category subtasks; and, according to the order of the multiple subtasks, fusing the first-category data and the second-category data as learning data corresponding to the imitation learning task.

[0168] An embodiment of the present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the following method is implemented: generating an imitation learning task; splitting the imitation learning task into multiple subtasks; determining whether each of the subtasks matches a primitive task in a primitive library, treating the subtasks that match any primitive task as first-category subtasks, and treating the subtasks that cannot match any primitive task as second-category subtasks; obtaining first-category data corresponding to the first-category subtasks from the primitive library; obtaining second-category data corresponding to the second-category subtasks; and according to the order of the multiple subtasks, fusing the first-category data and the second-category data as learning data corresponding to the imitation learning task.

[0169] The above-mentioned humanoid robot imitation learning data collection method, device, computer equipment and storage medium achieve the beneficial effect of solving the technical problems raised in the background technology.

[0170] Figure 2 FIG. 1 is a flow chart of a method for collecting data for imitating learning of a humanoid robot in one embodiment. It should be understood that although Figure 2 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0171] Figure 6The internal structure diagram of a computer device in one embodiment is shown. The computer device may specifically be Figure 1 The server 120 in FIG. Figure 6 As shown, the computer device includes a processor, a memory, a network interface, an input device and a display screen connected through a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor can implement the humanoid robot imitation learning data collection method. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor can execute the humanoid robot imitation learning data collection method. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a button, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.

[0172] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0173] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0174] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0175] The foregoing is merely a specific embodiment of the present invention, which enables those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A humanoid robot imitation learning data collection method, characterized in that: The method comprises: Generate imitation learning tasks; Splitting the imitation learning task into multiple subtasks; Determine whether each of the subtasks matches a primitive task in the primitive library, and regard the subtasks that match any primitive task as first-category subtasks, and regard the subtasks that cannot match any primitive task as second-category subtasks; Acquire, from the primitive library, first-category data corresponding to the first-category subtask; Obtaining the second type of data corresponding to the second type of subtask; According to the order of the multiple subtasks, the first category of data and the second category of data are fused as learning data corresponding to the imitation learning task.

2. The method according to claim 1, characterized in that If the second-type subtask is a non-contact task, the step of obtaining the second-type data corresponding to the second-type subtask includes: Obtaining the starting position, the ending position and the humanoid robot parameters of the second type of subtask; Generate a trajectory from the starting position to the ending position according to the starting position, the ending position and the humanoid robot parameters; According to the trajectory, the corresponding second-category data is generated.

3. The method according to claim 1, characterized in that: The obtaining the second type of data corresponding to the second type of subtask includes: enabling the humanoid robot to complete the second type of subtask; Collect motion data of the humanoid robot in the process of completing the second type of subtask, and use the motion data as the corresponding second type of data.

4. The method according to claim 3, characterized in that: The primitive library includes a plurality of primitive tasks and primitive data corresponding to each of the primitive tasks. After collecting the motion data of the humanoid robot in the process of completing the second type of subtask, the method further includes: The second type of subtask and the corresponding action data are placed in the primitive library as a primitive task of the primitive library and the corresponding primitive data.

5. The method according to claim 1, characterized in that Before generating a plurality of imitation learning tasks, the method further includes: Acquire a primitive task, wherein the primitive task includes environmental information; placing the humanoid robot in a virtual environment matching the environmental information of the primitive task; enabling the humanoid robot to complete the primitive task in the virtual environment; Obtaining initial data corresponding to the primitive task; The initial data is generalized according to the random variables, and the generalized initial data is used as primitive data.

6. The method according to claim 3, characterized in that The head of the humanoid robot is provided with a camera, the camera includes an RGB camera and a depth camera, the RGB camera is used to collect an RGB image of the main perspective of the humanoid robot, and the depth camera is used to collect a depth image of the main perspective of the humanoid robot; The obtaining of initial data corresponding to the primitive task includes: According to the camera, initial data of the main perspective of the robot is collected during the process of the humanoid robot completing the primitive task.

7. The method according to claim 6, characterized in that The initial data includes: An RGB image of the main perspective of the humanoid robot; A depth image of the main perspective of the humanoid robot; External control instructions for each joint of the humanoid robot; Real-time status feedback data of each joint of the robot; Timestamp.

8. The method according to claim 1, characterized in that: After fusing the first type of data and the second type of data according to the order of the multiple subtasks to serve as learning data corresponding to the imitation learning task, the method further includes: Get the transformation matrix between the robot coordinate system and the world coordinate system; Convert the robot's main perspective learning data into world coordinate system learning data.

9. A humanoid robot imitation learning data acquisition device, characterized in that: The device comprises: A task generation unit, used to generate imitation learning tasks; A splitting unit, used for splitting the imitation learning task into multiple subtasks; A judging unit, used for judging whether each of the subtasks matches a primitive task in the primitive library, treating a subtask that matches any primitive task as a first-category subtask, and treating a subtask that cannot match any primitive task as a second-category subtask; A data acquisition unit, configured to acquire, from the primitive library, first-category data corresponding to the first-category subtask; The data acquisition unit is further used to acquire the second type of data corresponding to the second type of subtask; A fusion unit is used to fuse the first category of data and the second category of data according to the order of the multiple subtasks as learning data corresponding to the imitation learning task.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.