Method for Creating Embodied Skills and Related Devices

User data is sent to cloud agents through terminal devices for expansion, generating diversified training data, solving the problem of inefficient training of embodied robots in complex skills learning, and achieving efficient skill creation.

CN119849543BActive Publication Date: 2025-07-22SHANGHAI FOURIER INTELLIGENCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510329705.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-22
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

In the prior art, when embossed robots learn complex skills, traditional imitation learning methods are difficult to expand to more complex scenarios and more diverse tasks, resulting in inefficient training.

Method used

The multimodal demonstration data uploaded by the user is sent to the cloud agent for expansion through the terminal device, and a diversified training data is generated and sent to the embodied robot for training. The multimodal data is processed using motion capture and audio capture technologies to ensure the similarity between the expanded data and the data uploaded by the user.

Benefits of technology

It reduces the cost of data acquisition, improves the efficiency of embodied robots to learn new skills, and ensures the quality and accuracy of expanded data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119849543B_ABST
    Figure CN119849543B_ABST
Patent Text Reader

Abstract

The present application discloses a method for creating embodied skills and related devices. The method includes: receiving multi-modal demonstration data uploaded by a user for a target skill; creating a first data augmentation request message; sending the first data augmentation request message to the cloud intelligent agent, where the first data augmentation request message is used to instruct the cloud intelligent agent to determine first training data; receiving a first data augmentation response message from the cloud intelligent agent, displaying the first training data on a training data management interface, and recording the total data volume of the current training data; if the total data volume of the current training data is greater than or equal to a preset data volume and a training request from the user for the target skill is detected, creating a first skill training instruction; and sending the first skill training instruction to the embodied robot. The present application can improve the efficiency of the embodied robot's imitation learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robot control, and in particular, to a method for creating embodied skills and related devices. Background Art

[0002] Training a robot's skills through imitation learning has become a popular robot training method. The traditional approach is for a human operator to remotely operate the robot through different control interfaces to complete various operation tasks, and use this demonstration data to train the robot to autonomously execute tasks. Although this method has achieved good results in some simple tasks, when it comes to expanding to more complex scenarios and more diverse tasks, it may not be able to learn sufficiently complex motion patterns from the current demonstration data. Summary of the Invention

[0003] Embodiments of this application provide a method for creating embodied skills and related devices to improve the efficiency of imitation learning of embodied robots.

[0004] In a first aspect, embodiments of this application provide a method for creating embodied skills, which is applied to a terminal device of a skill creation system. The skill creation system further includes an embodied robot and a cloud agent, and the cloud agent is deployed on at least one server. The method includes:

[0005] Receiving multi-modal demonstration data for a target skill uploaded by a user, where the target skill is a skill that the user determines needs to be added; and creating a first data augmentation request message; and sending the first data augmentation request message to the cloud agent. The first data augmentation request message includes the multi-modal demonstration data, and the first data augmentation request message is used to instruct the cloud agent to determine first training data, where the first training data is first action data that the embodied robot needs to simulate and train to achieve the target skill;

[0006] Receiving a first data augmentation response message from the cloud agent, where the first data augmentation response message includes the first training data, displaying the first training data on a training data management interface, and recording the total data volume of the current training data;

[0007] If the total data volume of the current training data is greater than or equal to a preset data volume, and a training request for the target skill from the user is detected, creating a first skill training instruction; and sending the first skill training instruction to the embodied robot, so that the embodied robot trains the target skill according to the first training data to achieve the creation of the target skill.

[0008] Wherein, the method further includes:

[0009] If the total amount of the current training data is less than the preset amount of data, a first prompt window is displayed on the training data management interface. The first prompt window includes a first prompt message, a first confirmation control, and a first denial control. The first prompt message is used to indicate to the user that the amount of the current training data is insufficient and determine whether to continue to expand the training data.

[0010] Upon detecting a trigger operation on the first confirmation control, a second data expansion request message is created; and the second data expansion request message is sent to the cloud intelligent agent. The second data expansion request message includes the total amount of the current training data. The second data expansion request message is used to instruct the cloud intelligent agent to determine second training data, where the second training data is second action data that the embodied robot needs to simulate and train to achieve the target skill.

[0011] Upon receiving a second data expansion response message from the cloud intelligent agent, where the second data expansion response message includes the second training data, the second training data is displayed on the training data management interface, and the total amount of the current training data is updated for the first time according to the second training data.

[0012] If the total amount of the current training data obtained after the first update is greater than or equal to the preset amount of data, and a training request for the target skill is detected by the user, a second skill training instruction is created; and the second skill training instruction is sent to the embodied robot. The second skill training instruction is used to instruct the embodied robot to train the target skill according to the first training data and the second training data.

[0013] A training feedback message is received from the embodied robot. The training feedback message is used to indicate the completion degree of the embodied robot for the target skill.

[0014] If the completion degree is greater than or equal to the preset completion degree, a control for the target skill is added to the embodied robot skill demonstration interface.

[0015] Wherein, the method further includes:

[0016] If the total amount of the current training data obtained after the first update is less than the preset amount of data, the first prompt window is displayed on the training data management interface.

[0017] Detect a triggering operation on the first determination control, and create a third data augmentation request message; and send the third data augmentation request message to the cloud agent, where the third data augmentation request message is used to instruct the cloud agent to determine third training data, and the third training data is third action data that the embodied robot needs to simulate and train to achieve the target skill;

[0018] Receive a third data augmentation response message from the cloud agent, where the third data augmentation response message includes the third training data, display the third training data on the training data management interface, and update the total data volume of the currently updated training data according to the third training data;

[0019] If the total data volume of the currently updated training data after the second update is greater than or equal to the preset data volume, and a training request for the target skill is detected, create a third skill training instruction; and send the third skill training instruction to the embodied robot, where the third skill training instruction is used to instruct the embodied robot to train the target skill according to the first training data, the second training data, and the third training data;

[0020] Receive a training feedback message from the embodied robot, where the training feedback message is used to indicate the completion degree of the embodied robot for the target skill;

[0021] If the completion degree is greater than or equal to the preset completion degree, add a control for the target skill to the embodied robot skill demonstration interface.

[0022] Wherein, the determination process of the first training data includes the following steps:

[0023] Perform action capture and audio capture on the multi-modal demonstration data to obtain fourth action data and first audio data;

[0024] Perform single action decomposition on the fourth action data to obtain multiple first sub-action data; and perform keyword search on each first sub-action data in the multiple first sub-action data to obtain multiple second sub-action data; and perform parsing on each first sub-action data and each second sub-action data in the multiple second sub-action data to obtain sub-action training data;

[0025] Determine the waveform of the first audio data; and, based on the waveform, find M fifth action data corresponding to the first audio data; and, determine the matching degrees between the fourth action data and each of the M fifth action data to obtain M matching degrees; and, screen the M second action data according to the M matching degrees to obtain N second action data, where N is less than or equal to M;

[0026] Determine the N fifth action data as the fourth training data;

[0027] Determine the first training data according to the sub-action training data and the fourth training data.

[0028] Wherein, the first data augmentation response message further includes second prompt information, and the second prompt information is used to indicate that the demonstration angle of the sub-action data is sufficient, or is used to indicate that the demonstration angle of the sub-action data is insufficient. Wherein, if the second prompt information is used to indicate that the demonstration angle of the sub-action data is insufficient, the second prompt information includes information on the sub-action data with insufficient demonstration angle.

[0029] Wherein, the process of determining the second training data includes the following steps:

[0030] Determine the difference between the total data volume of the current training data and the preset data volume to obtain a target difference;

[0031] Determine the target demonstration data according to the target difference, and the target demonstration data includes demonstration data of at least one modality in the multi-modal demonstration data;

[0032] Extract the feature data of the target demonstration data;

[0033] Find the corresponding action data according to the feature data to obtain P sixth action data;

[0034] Perform action capture on the multi-modal demonstration data to obtain seventh action data;

[0035] Determine the matching degrees between the seventh action data and each of the P sixth action data to obtain P matching degrees;

[0036] Screen the P sixth action data according to the P matching degrees to obtain Q sixth action data, where Q is less than or equal to P;

[0037] Determine the Q sixth action data as the second training data.

[0038] Wherein, after updating the total data volume of the current training data obtained after the first update according to the third training data, the method further includes:

[0039] Upon receiving the request from the user to perform an editing operation on the target training data, a second prompt window is displayed on the training data management interface. The second prompt window includes a third prompt message, a second confirmation control, and a second denial control. The third prompt message is used to prompt the user whether to continue performing the editing operation. The target training data includes the first training data, and / or the second training data, and / or the third training data;

[0040] Upon detecting a triggering operation on the second confirmation control, perform the editing operation and update the training data management interface.

[0041] In a second aspect, an embodiment of the present application provides a creation device for embodied skills, which is applied to a terminal device of a skill creation system. The skill creation system further includes an embodied robot and a cloud intelligent agent. The cloud intelligent agent is deployed on at least one server and includes:

[0042] A first processing unit, configured to receive multi-modal demonstration data uploaded by the user for a target skill, where the target skill is a skill that the user determines needs to be added; and create a first data augmentation request message; and send the first data augmentation request message to the cloud intelligent agent. The first data augmentation request message includes the multi-modal demonstration data, and the first data augmentation request message is used to instruct the cloud intelligent agent to determine first training data, where the first training data is first action data that the embodied robot needs to simulate and train to implement the target skill;

[0043] A display unit, configured to receive a first data augmentation response message from the cloud intelligent agent. The first data augmentation response message includes the first training data, display the first training data on the training data management interface, and record the total data volume of the current training data;

[0044] A second processing unit, configured to, if the total data volume of the current training data is greater than or equal to a preset data volume and a training request from the user for the target skill is detected, create a first skill training instruction; and send the first skill training instruction to the embodied robot, so that the embodied robot trains the target skill according to the first training data to implement the creation of the target skill.

[0045] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory storing execution instructions. The memory stores one or more programs; when the processor executes the execution instructions stored in the memory, the processor executes the method described in the first aspect.

[0046] Fourthly, an embodiment of the present application provides a computer-readable storage medium storing an energy data management program, including execution instructions. When a processor of an electronic device executes the execution instructions, the processor executes the method described in the first aspect.

[0047] Fifthly, an embodiment of the present application provides a computer program product. The computer program product includes a non-transitory computer-readable storage medium storing a computer program. The computer program is operable to cause a computer to execute some or all of the steps described in the first aspect of the embodiments of the present application. The computer program product may be a software installation package.

[0048] It can be seen that in the embodiment of the present application, first, multi-modal demonstration data for a target skill uploaded by a user is received. The target skill is a skill that the user determines needs to be newly added. Then, a first data augmentation request message is created. Then, the first data augmentation request message is sent to the cloud agent. The first data augmentation request message includes the multi-modal demonstration data. The first data augmentation request message is used to instruct the cloud agent to determine first training data. The first training data is first action data required for the embodied robot to simulate and train to achieve the target skill. Then, a first data augmentation response message from the cloud agent is received. The first data augmentation response message includes the first training data. The first training data is displayed on a training data management interface, and the total data volume of the current training data is recorded. If the total data volume of the current training data is greater than or equal to a preset data volume, and a training request for the target skill by the user is detected, a first skill training instruction is created. Then, the first skill training instruction is sent to the embodied robot, so that the embodied robot trains the target skill according to the first training data, and the creation of the target skill is achieved.

[0049] After receiving an instruction for the user to create a new skill, compared with the embodiment where the embodied robot trains the skill with a large amount of demonstration data collected manually, in this solution, the terminal device sends the demonstration data uploaded by the user to the cloud agent. The cloud agent augments the demonstration data according to the demonstration data uploaded by the user, that is, a small amount of demonstration data is augmented to generate a large-scale and diverse demonstration data, reducing the cost of collecting demonstration data. At the same time, the similarity between the augmented demonstration data and the demonstration data uploaded by the user is ensured. After the cloud agent obtains diverse demonstration data through augmentation, it sends the data to the terminal device, so that the terminal device sends a skill training instruction to the robot. Then, the robot trains the skill according to the diverse demonstration data, thereby achieving the creation of a new skill, improving the learning efficiency of the new skill, and further improving the imitation learning efficiency of the embodied robot. Description of the Drawings

[0050] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0051] Figure 1 is the system architecture diagram of a skill creation system provided by an embodiment of the present application;

[0052] Figure 2 is the flow schematic diagram of a method for creating an embodied skill provided by an embodiment of the present application;

[0053] Figure 3 is the schematic diagram of a training data management interface provided by an embodiment of the present application;

[0054] Figure 4 is the schematic diagram of another training data management interface provided by an embodiment of the present application;

[0055] Figure 5 is the schematic diagram of a skill demonstration interface provided by an embodiment of the present application;

[0056] Figure 6 is the schematic diagram of a prompt window provided by an embodiment of the present application;

[0057] Figure 7 is the schematic diagram of a training data editing interface provided by an embodiment of the present application;

[0058] Figure 8 is the block diagram of the functional units of a device for creating an embodied skill provided by an embodiment of the present application;

[0059] Figure 9 is the block diagram of the functional units of another device for creating an embodied skill provided by an embodiment of the present application;

[0060] Figure 10 is the schematic diagram of the structure of an electronic device proposed by an embodiment of the present application. Detailed implementation manners

[0061] In order to enable those skilled in the art to better understand the solution of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present application.

[0062] In the description and claims of this application and the above-mentioned drawings, terms such as "first" and "second" are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.

[0063] Referring to "embodiments" herein means that the specific features, structures or characteristics described in connection with the embodiments may be included in at least one embodiment of this application. The appearance of this phrase in various positions in the description does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.

[0064] When an embodied robot learns new skills, the current demonstration data cannot be extended to more complex scenarios and more diverse tasks.

[0065] In view of the above problems, the embodiments of this application provide a method for creating embodied skills and related devices. The embodiments of this application will be introduced in detail below with reference to the drawings.

[0066] Please refer to Figure 1 , Figure 1 which is a system architecture diagram of a skill creation system provided by an embodiment of this application. As Figure 1 shown, the skill creation system 100 includes a terminal device 101, an embodied robot 102, and a cloud agent 103. The terminal device 101 can communicate with the embodied robot 102 and the cloud agent 103 respectively through a network or other connection methods. The terminal device 101 refers to the device used by the user. The user can remotely control the embodied robot through the application program interface installed on the terminal device 101 to learn skills. Specifically, the user can click on the new skill control and upload demonstration data in the display screen interface of the terminal device. After the terminal device 101 receives the demonstration data, it sends it to the cloud agent 103. The terminal device 101 may include a tablet computer, a personal digital assistant, in-vehicle electronic devices, a server, a laptop computer, a mobile Internet device (MID), or wearable electronic devices (such as smart watches, Bluetooth headsets), smart phones (such as Android phones, iOS phones, Windows Phone phones, etc.). The above are only examples, not an exhaustive list, and include but are not limited to the above-mentioned electronic devices.

[0067] Among them, the cloud agent 103 can deploy multiple servers, which are communicatively connected to each other. After receiving the user's demonstration data, the cloud agent 103 expands the demonstration data to obtain diversified demonstration data, and transmits the diversified demonstration data to the terminal device 101, so that the terminal device 101 sends a training instruction to the embodied robot 102 based on the diversified demonstration data, enabling the embodied robot 102 to train new skills according to the diversified demonstration data. After the training is completed, a control for the new skill is added to the skill interface of the embodied robot.

[0068] Among them, the embodied robot 102 embeds artificial intelligence into a tangible entity such as a robot, enabling it to have the ability to perceive, learn, and dynamically participate in the surrounding environment. It can have a human-like appearance, including two legs, two arms, and a head. It is a robot that mimics human functions and intelligence and can perform various tasks in the working and living environments of humans.

[0069] Please refer to Figure 2 , Figure 2 FIG. is a schematic flowchart of a method for creating an embodied skill provided by an embodiment of the present application. This method is applied to the terminal device of a skill creation system, and the skill creation system further includes an embodied robot and a cloud agent. The cloud agent is deployed on at least one server. The method includes the following steps.

[0070] S210, receiving multimodal demonstration data for a target skill uploaded by a user, where the target skill is a skill that the user determines needs to be newly added; and creating a first data augmentation request message; and sending the first data augmentation request message to the cloud agent.

[0071] Among them, the first data augmentation request message includes the multimodal demonstration data, and the first data augmentation request message is used to instruct the cloud agent to determine first training data, where the first training data is first action data that the embodied robot needs to simulate and train to achieve the target skill.

[0072] Among them, upon receiving a request from the user to add a skill to the embodied robot, a training data management interface is displayed on the display screen of the terminal device. Please refer to Figure 3 , Figure 3 FIG. is a schematic diagram of a training data management interface provided by an embodiment of the present application. As shown in Figure 3As shown, the training data management interface includes multiple new data controls and training controls, and counts the amount of data currently being trained. The amount of data currently being trained is 0. The new data control is represented by a '+'. When the user clicks on this control, a data upload window will be entered to facilitate data upload. The training control is used for the user to click on this training control to train the target skill when the current amount of data is sufficient; the ellipsis is used to indicate that more new data controls can be added.

[0073] Among them, the multi-modal demonstration data can be data in various forms such as text, images, audio, and video. In the case where the data uploaded by the user is in the form of text or images, the cloud intelligent agent can convert the text or images into data such as video, and then perform an expansion operation on the data.

[0074] Exemplarily, the skill selected by the user can be dancing along with the target music. The multi-modal demonstration data uploaded can be video data showing the dance from multiple angles according to the target music. Further, click on the new data control to enter the data upload window. The data upload window includes an upload control and a real-time recording control. If the user selects the upload control, the user's dance demonstration data can be obtained. If the user selects the real-time recording control, the user's demonstration video is recorded through the camera configured on the embodied robot. Specifically, it can be that the user plays the music and shows dance movements based on the music, and the embodied robot captures the user's demonstration data based on the camera to obtain the initial demonstration data, that is, the multi-modal demonstration data.

[0075] In a possible embodiment, the determination process of the first training data includes the following steps: performing motion capture and audio capture on the multi-modal demonstration data to obtain fourth motion data and first audio data; decomposing the fourth motion data into a single motion to obtain multiple first sub-motion data; and, performing keyword search on each of the multiple first sub-motion data to obtain multiple second sub-motion data; and, parsing each of the first sub-motion data and each of the multiple second sub-motion data to obtain sub-motion training data; determining the waveform of the first audio data; and, searching for M fifth motion data corresponding to the first audio data according to the waveform; and, determining the matching degree between the fourth motion data and each of the M fifth motion data to obtain M matching degrees; and, screening the M second motion data according to the M matching degrees to obtain N second motion data, where N is less than or equal to M; determining the N fifth motion data as the fourth training data; and determining the first training data according to the sub-motion training data and the fourth training data.

[0076] Among them, through motion capture technology, the motion trajectory and posture changes of the user can be accurately recorded, making the demonstration process more vivid and intuitive. At the same time, audio capture technology can obtain various sound signals generated by the user during the demonstration, such as motion noise, voice output, etc., thus providing richer demonstration information.

[0077] Among them, it is necessary to segment the motion data of the user. Through motion recognition algorithms, the specific motions performed by the user during the demonstration can be recognized and divided into a series of basic motions. Then, relevant video data corresponding to the basic motions is queried in the Internet database. Exemplarily, relevant video data can be retrieved by inputting keywords or descriptive information of the basic motions, and then the queried video data and the corresponding multiple basic motions are parsed. Exemplarily, through video processing technologies, such as inter-frame difference, optical flow estimation, etc., the motions in the video are recognized and tracked to extract the basic motion information therein, obtaining sub-action training data.

[0078] Among them, the waveform of the first audio data can be obtained through audio processing technology. The waveform is the representation of the audio signal in the time domain, which reflects the change of the amplitude of the audio signal over time. By analyzing the waveform, the basic characteristics of the audio signal, such as frequency, amplitude, etc., are obtained. In the database, the corresponding motion data is searched according to the waveform characteristics of the first audio data, and this motion data represents the motion sequence matching the audio data. And the similarity between the calculated multiple motion data and the motion data obtained by motion capture is calculated. When the similarity is greater than the similarity threshold, it is determined as training data. Among them, this training data and the sub-action training data constitute the first training data.

[0079] Exemplarily, after performing dance motion capture and music capture on the dance video uploaded by the user, the user's dance motions are divided into multiple basic motions, and then relevant video data corresponding to the basic motions is queried in the Internet database, and then the queried video data is parsed into training data; and the music played by the user is determined, the dance video corresponding to the same music is selected from the database, the obtained dance video is matched with the dance video demonstrated by the user, and the dance video with a high matching degree is determined as the demonstration video, obtaining training data.

[0080] It can be seen that in the embodiments of the present application, through the parsing and processing of multi-modal data, similar data is searched in the database and screened according to the similarity degree, while expanding the training data, ensuring the similarity between the expanded demonstration data and the demonstration data uploaded by the user, and improving the quality of the expanded data.

[0081] In a possible embodiment, the first data augmentation response message further includes second prompt information, where the second prompt information is used to indicate that the demonstration angle of the sub-action data is sufficient, or the second prompt information is used to indicate that the demonstration angle of the sub-action data is insufficient. Wherein, if the second prompt information is used to indicate that the demonstration angle of the sub-action data is insufficient, the second prompt information includes information about the sub-action data with insufficient demonstration angle.

[0082] Wherein, after obtaining the sub-action training data, it is determined whether each sub-action in the sub-action training data includes data of multiple angles. If there are multiple video angles in the demonstration data corresponding to the current basic action, for example, data of four angles: front, back, left, and right, it is determined that the demonstration data of this basic action is sufficient. Otherwise, it is considered that the demonstration angle of this basic action needs to be supplemented, and a prompt message is output to prompt the user to demonstrate the basic action of the corresponding angle.

[0083] It can be seen that in this embodiment, by determining whether the video angle of each basic action is sufficient, the integrity of the training data is ensured, thereby improving the accuracy of the embodied robot in performing the target skill.

[0084] S220. Receive the first data augmentation response message from the cloud agent, display the first training data on the training data management interface, and record the total data volume of the current training data.

[0085] Wherein, the first data augmentation response message includes the first training data.

[0086] Exemplarily, please refer to Figure 3 and Figure 4 , Figure 4 is a schematic diagram of another training data management interface provided by an embodiment of the present application. As shown in Figure 4 , the first new data control is replaced by the demonstration data. If this control is clicked, the demonstration video can be viewed. The second new data control is replaced by the first augmented data. If this control is clicked, the first augmented data can be viewed. The first training data includes the demonstration data and the first augmented data. The total number of the demonstration data and the first augmented data is 1789, and the current training data volume has changed from 0 before to 1789 now.

[0087] S230. If the total data volume of the current training data is greater than or equal to the preset data volume, and a training request for the target skill is detected by the user, create a first skill training instruction; and send the first skill training instruction to the embodied robot.

[0088] Wherein, the first skill training instruction is used to instruct the embodied robot to train according to the first training data.

[0089] Among them, a training data volume threshold can be set for each newly added skill. Specifically, it can be determined according to the difficulty of each skill. The higher the difficulty of the skill, the larger the corresponding training data volume threshold. If the training data volume threshold corresponding to the skill that the current user needs to add is 1500, it is determined that the total data volume of the current training data is greater than or equal to the preset data volume, and the user can be prompted whether to start training. If it is detected that the user triggers an operation on the training control, a first skill training instruction is created to instruct the embodied robot to train according to the first training data. After receiving the training instruction, the embodied robot trains according to the first training data and gives real-time feedback on the training progress to the user terminal. Specifically, the training progress can be reflected by the accuracy of the current executed action, reaction time, etc.

[0090] Among them, if the terminal device receives a message that the training of the embodied robot for the target skill has been completed, a skill demonstration interface is displayed on the display screen, and a control for the target skill is added to the skill demonstration interface.

[0091] Exemplarily, please refer to Figure 5 , Figure 5 is a schematic diagram of a skill demonstration interface provided by an embodiment of the present application. As Figure 5 shown, the original skill demonstration controls include shaking hands, bending down, pouring water, squatting, and taking out the trash, etc. Now a new skill section is added, and a control for dancing to music is added. Specifically, the user can click on the control for dancing to music to play the target music, and the embodied robot 102 demonstrates this action simultaneously in the demonstration scene 402 and the real scene.

[0092] In a possible embodiment, if the total amount of the current training data is less than the preset amount of data, a first prompt window is displayed on the training data management interface. The first prompt window includes first prompt information, a first confirmation control, and a first denial control. The first prompt information is used to indicate to the user that the amount of the current training data is insufficient and to determine whether to continue to expand the training data. Detecting a trigger operation on the first confirmation control, a second data expansion request message is created; and the second data expansion request message is sent to the cloud intelligent agent. The second data expansion request message includes the total amount of the current training data. The second data expansion request message is used to instruct the cloud intelligent agent to determine second training data, where the second training data is second action data that the embodied robot needs to simulate and train to achieve the target skill. Receiving a second data expansion response message from the cloud intelligent agent, where the second data expansion response message includes the second training data, the second training data is displayed on the training data management interface, and the total amount of the current training data is updated for the first time according to the second training data; if the total amount of the current training data obtained after the first update is greater than or equal to the preset amount of data, and detecting a training request from the user for the target skill, a second skill training instruction is created; and the second skill training instruction is sent to the embodied robot. The second skill training instruction is used to instruct the embodied robot to train the target skill according to the first training data and the second training data; receiving a training feedback message from the embodied robot, where the training feedback message is used to indicate the completion degree of the embodied robot for the target skill; if the completion degree is greater than or equal to the preset completion degree, a control for the target skill is added to the embodied robot skill demonstration interface.

[0093] Exemplarily, if the threshold of the amount of training data corresponding to the skill that the current user needs to add is 2000, it is determined that the total amount of the current training data is less than the preset amount of data. Please refer to Figure 6 , Figure 6 which is a schematic diagram of a prompt window provided by an embodiment of the present application. As Figure 6 shown, a prompt window pops up. The prompt window includes prompt information, a confirmation control, and a denial control. The prompt information is that the current training data is insufficient. Whether to continue to expand is determined by clicking the confirmation control and the denial control.

[0094] In a possible embodiment, detecting that the user clicks the confirmation control, a continue expansion request is sent to the cloud intelligent agent.

[0095] In a possible embodiment, the process of determining the second training data includes the following steps: determining the difference between the total amount of the current training data and the preset amount of data to obtain a target difference; determining the target demonstration data according to the target difference, where the target demonstration data includes the demonstration data of at least one modality in the multi-modal demonstration data; extracting the feature data of the target demonstration data; finding the corresponding action data according to the feature data to obtain P sixth action data; performing action capture on the multi-modal demonstration data to obtain seventh action data; determining the matching degree between the seventh action data and each of the P sixth action data to obtain P matching degrees; screening the P sixth action data according to the P matching degrees to obtain Q sixth action data, where Q is less than or equal to P; and determining the Q sixth action data as the second training data.

[0096] Among them, after receiving the expansion request, the cloud intelligent agent determines according to the difference between the total amount of the current training data carried in the request message and the preset amount of data. In the case of a large difference, key features are extracted from the demonstration data of each modality; in the case of a small difference, the demonstration data of a certain modality is randomly selected for key feature extraction, such as extracting key features for audio data or for action videos.

[0097] Specifically, for action videos, the key features can be middle-level feature extraction. Through a random forest framework, multiple low-level features are fused to establish a middle-level feature representation with strong discriminative ability and descriptive ability, so as to effectively represent actions and accurately identify the combination and segmentation of multiple actions. It can also be spatio-temporal context features, that is, taking the time context features as important information for describing actions, reflecting the context information between local spatio-temporal blocks, which helps to more accurately represent and identify actions.

[0098] Specifically, for audio data, the key features can be time-domain features, frequency-domain features, time-frequency domain features, etc.

[0099] Among them, the time-domain features can include short-time energy, which reflects the energy change of the audio signal in a short time interval, such as the zero-crossing rate, the number of times the signal crosses from positive to negative or from negative to positive per unit time, effectively reflecting the frequency characteristics of the audio signal; for example, the peak value, the maximum amplitude of the audio signal in a certain time period, which can be used to describe the intensity and dynamic range of the audio signal.

[0100] Among them, the frequency-domain features may include spectral flatness, which describes the smoothness of the audio signal spectrum; for example, spectral centroid, the central frequency of the spectral energy distribution, which can reflect the main frequency components of the audio signal; for example, Mel-frequency cepstral coefficients, which transform the spectrum into a set of coefficients by simulating the auditory perception mechanism of the human ear and are applied to speech recognition and music classification.

[0101] Among them, the time-frequency domain features may include short-time Fourier transform, which can obtain the time-frequency representation of the signal by windowing the audio signal and calculating the Fourier transform of each window; for example, waveform analysis, which can effectively capture the local characteristics of the signal by analyzing the short-time waveform features of the audio signal, such as short-time energy, short-time zero-crossing, etc.; for example, wavelet transform, which can obtain information in both time and frequency by performing wavelet decomposition of the audio signal at different scales.

[0102] Among them, according to the determined key features, search for the corresponding action data in the database under the same key features, and calculate the similarity between the found action data and the action data captured by the action capture. When the similarity is greater than the similarity threshold, it is determined as the second training data.

[0103] Exemplarily, according to the training data volume threshold of 2000 and the current training data of 1789, the difference of 211 is obtained, and the difference is small, so the features of a certain modality can be searched. Exemplarily, the beats per minute of the music played by the user can be determined, and the dance video corresponding to the music with the same beats per minute can be selected, and then the second training data is determined based on the similarity between the currently obtained dance video and the dance video demonstrated by the user.

[0104] Among them, after the cloud agent expands to obtain the second training data, it transmits it to the terminal device, so that the terminal device receives the training data and counts the quantity of the training data.

[0105] Exemplarily, the third new data control in the training data management interface is replaced by the second expanded data, and the second expanded data is the second training data. If this control is clicked, the second expanded data can be viewed, and the current training data volume changes from 1789 before to 3612 now. It is determined that the total data volume of the current training data is greater than or equal to the preset data volume, and at the same time, it is prompted whether the user needs to start training. If it is detected that the user triggers the training control, a second skill training instruction is created, instructing the embodied robot to perform skill training according to the first training data and the second training data; after receiving the training instruction, the embodied robot performs training according to the first training data and the second training data and gives real-time feedback on the training stage to the user terminal. If the terminal device receives the message that the training of the target skill by the embodied robot has been completed, a skill demonstration interface is displayed on the display screen, and a control for the target skill is added to the skill demonstration interface.

[0106] It can be seen that in the embodiment of the present application, key features of multiple demonstration data are extracted, and training data with high similarity is found according to the key features, which ensures the similarity between the expanded demonstration data and the demonstration data uploaded by the user, reduces the cost of collecting demonstration data, and improves the quality of the expanded data.

[0107] In a possible embodiment, if the total data volume of the current training data obtained after the first update is less than the preset data volume, the first prompt window is displayed on the training data management interface; upon detecting a trigger operation on the first confirmation control, a third data expansion request message is created; and the third data expansion request message is sent to the cloud intelligent agent, where the third data expansion request message is used to instruct the cloud intelligent agent to determine third training data, and the third training data is the third action data that the embodied robot needs to simulate and train to achieve the target skill; upon receiving a third data expansion response message from the cloud intelligent agent, where the third data expansion response message includes the third training data, the third training data is displayed on the training data management interface, and the total data volume of the current training data obtained after the first update is updated according to the third training data; if the total data volume of the current training data obtained after the second update is greater than or equal to the preset data volume, and a training request for the target skill is detected from the user, a third skill training instruction is created; and the third skill training instruction is sent to the embodied robot, where the third skill training instruction is used to instruct the embodied robot to train the target skill according to the first training data, the second training data, and the third training data; upon receiving a training feedback message from the embodied robot, where the training feedback message is used to indicate the completion degree of the embodied robot for the target skill; if the completion degree is greater than or equal to the preset completion degree, a control for the target skill is added to the embodied robot skill demonstration interface.

[0108] Among them, after obtaining the second training data, if the training data volume threshold corresponding to the skill that the current user needs to add is 4000, it is determined that the total data volume of the current training data is less than the preset data volume, a prompt window is popped up to determine whether to continue expanding the training data, and upon detecting that the user clicks the confirmation control in the prompt window, a continue expansion request is sent to the cloud intelligent agent. After receiving the expansion request, the cloud intelligent agent searches for action data with high similarity in the database according to the captured action data of the user as the to-be-processed action data, and clips and synthesizes the to-be-processed action data, the first training data, and the second training data to obtain the third training data.

[0109] Specifically, the clip synthesis can be performed by synthesizing based on the matching of body postures, the similarity of movement amplitudes, the unity of picture composition, the coordination of color matching, the coordination of movements and audio, etc.

[0110] Exemplarily, a dance video with a relatively high similarity to the dance video of the user's initial demonstration can be selected as the dance video to be processed, and then the first training data, the second training data, and the dance video to be processed are used for clip synthesis to obtain the third training data.

[0111] Furthermore, the video frames in the dance video to be processed that are the same as the dance actions demonstrated by the user can be left, and the video frames of other dance actions can be replaced with the video frames of the dance actions demonstrated by the user or the corresponding dance actions in the first training data and the second training data.

[0112] Among them, after the cloud intelligent agent expands to obtain the third training data, it transmits it to the terminal device, so that the terminal device receives the training data and counts the quantity of the training data.

[0113] Exemplarily, the fourth new data control in the training data management interface is replaced by the third expanded data, and the third expanded data is the third training data. If this control is clicked, the third expanded data is viewed, and the current training data volume changes from 3612 before to 5781 now. It is determined that the total data volume of the current training data is greater than or equal to the preset data volume, and at the same time, the user is prompted whether to start training. If it is detected that the user triggers the training control, a third skill training instruction is created, instructing the embodied robot to perform skill training based on the first training data, the second training data, and the third training data; after receiving the training instruction, the embodied robot performs training based on the first training data, the second training data, and the third training data, and gives real-time feedback on the training stage to the user terminal. If the terminal device receives the message that the training of the target skill by the embodied robot has been completed, a skill demonstration interface is displayed on the display screen, and a control for the target skill is added to the skill demonstration interface.

[0114] It can be seen that in the embodiment of the present application, through the integration of multiple training data for expansion, while enriching the training data, the similarity between the demonstration data and the demonstration data uploaded by the user is ensured, the cost of collecting demonstration data is reduced, and the quality of the expanded data is improved.

[0115] It can be seen that in the embodiment of the present application, after receiving the instruction that the user needs to create a new skill, compared with the embodied robot training the skill with a large amount of demonstration data collected manually, in this solution, the terminal device sends the demonstration data uploaded by the user to the cloud intelligent agent, and the cloud intelligent agent expands the demonstration data according to the demonstration data uploaded by the user, that is, expands a small amount of demonstration data to generate a large-scale and diverse demonstration data, reducing the cost of collecting demonstration data, and at the same time ensuring the similarity between the expanded demonstration data and the demonstration data uploaded by the user. After the cloud intelligent agent expands to obtain diverse demonstration data, it sends it to the terminal device, so that the terminal device sends a skill training instruction to the robot, and then the robot trains the skill according to the diverse demonstration data, thereby realizing the creation of a new skill, improving the learning efficiency of the new skill, and further improving the imitation learning efficiency of the embodied robot.

[0116] In a possible embodiment, after updating the total data volume of the current training data obtained after the first update according to the third training data, the method further includes: receiving a request from the user to perform an editing operation on the target training data, and displaying a second prompt window on the training data management interface, where the second prompt window includes third prompt information, a second confirmation control, and a second denial control, and the third prompt information is used to prompt the user whether to continue to perform the editing operation, and the target training data includes the first training data, and / or the second training data, and / or the third training data; detecting a trigger operation on the second confirmation control, performing the editing operation, and updating the training data management interface.

[0117] Among them, after the user obtains multiple training data, the user can perform editing operations such as viewing or deleting the training data. Exemplarily, it can be that the user clicks on any one of the training data to view it and enters the training data editing interface. Please refer to Figure 7 , Figure 7 is a schematic diagram of a training data editing interface provided by an embodiment of the present application. As Figure 7 shown, the training data editing interface includes multiple training data, such as training data 1, training data 2, training data 3, etc. There is a viewing control and a deletion control behind each training data. If the user clicks on the viewing control to view training data 2 and feels that the matching degree of training data 2 is poor or the visual experience is not good, the user can click on the deletion control in the interface to delete it.

[0118] It can be seen that in the embodiment of the present application, it can enhance the user experience, optimize the storage space of the terminal device, and improve the quality of the training data.

[0119] In a possible embodiment, the obtained training data can be classified into categories and weights can be set, such as the user's initial demonstration data, basic action training data, and similar training data. When the embodied robot is undergoing training, it can train skills based on the weights of the training data, paying more attention to important data types, thereby improving the overall efficiency of training.

[0120] Consistent with the above embodiment, please refer to Figure 8 , Figure 8 which is a functional unit composition block diagram of an embodied skill creation device provided by an embodiment of the present application. As Figure 8 shown, the embodied skill creation device 80 includes: a first processing unit 81, configured to receive multimodal demonstration data for a target skill uploaded by a user, where the target skill is a skill that the user determines needs to be newly added; and create a first data augmentation request message; and send the first data augmentation request message to the cloud intelligent agent, where the first data augmentation request message includes the multimodal demonstration data, and the first data augmentation request message is used to instruct the cloud intelligent agent to determine first training data, where the first training data is first action data that the embodied robot needs to simulate and train to achieve the target skill; a display unit 82, configured to receive a first data augmentation response message from the cloud intelligent agent, where the first data augmentation response message includes the first training data, display the first training data on a training data management interface, and record the total data volume of the current training data; a second processing unit 83, configured to create a first skill training instruction if the total data volume of the current training data is greater than or equal to a preset data volume and a training request for the target skill from the user is detected; and send the first skill training instruction to the embodied robot, so that the embodied robot trains the target skill according to the first training data to achieve the creation of the target skill.

[0121] In a possible embodiment, in terms of creating skills, the creation device 80 of the embodied skills is further specifically configured to: if the total data volume of the current training data is less than the preset data volume, display a first prompt window on the training data management interface, where the first prompt window includes a first prompt message, a first confirmation control, and a first denial control, and the first prompt message is used to indicate that the data volume of the current training data of the user is insufficient and determine whether to continue to expand the training data; detect a trigger operation on the first confirmation control, and create a second data expansion request message; and send the second data expansion request message to the cloud intelligent agent, where the second data expansion request message includes the total data volume of the current training data, and the second data expansion request message is used to instruct the cloud intelligent agent to determine second training data, and the second training data is second action data that the embodied robot needs to simulate and train to achieve the target skill; receive a second data expansion response message from the cloud intelligent agent, where the second data expansion response message includes the second training data, display the second training data on the training data management interface, and update the total data volume of the current training data for the first time according to the second training data; if the total data volume of the current training data obtained after the first update is greater than or equal to the preset data volume, and a training request for the target skill is detected by the user, create a second skill training instruction; and send the second skill training instruction to the embodied robot, where the second skill training instruction is used to instruct the embodied robot to train the target skill according to the first training data and the second training data; receive a training feedback message from the embodied robot, where the training feedback message is used to indicate the completion degree of the embodied robot for the target skill; if the completion degree is greater than or equal to the preset completion degree, add a control for the target skill to the embodied robot skill demonstration interface.

[0122] Among them, the process of determining the first training data includes the following steps: performing action capture and audio capture on the multi-modal demonstration data to obtain fourth action data and first audio data; decomposing the fourth action data into single actions to obtain a plurality of first sub-action data; and, performing keyword search on each of the plurality of first sub-action data to obtain a plurality of second sub-action data; and, parsing each of the first sub-action data and each of the second sub-action data among the plurality of second sub-action data to obtain sub-action training data; determining the waveform of the first audio data; and, searching for M fifth action data corresponding to the first audio data according to the waveform; and, determining the matching degree between the fourth action data and each of the M fifth action data to obtain M matching degrees; and, screening the M fifth action data according to the M matching degrees to obtain N fifth action data, where N is less than or equal to M; determining the N fifth action data as the fourth training data; and determining the first training data according to the sub-action training data and the fourth training data.

[0123] Among them, the first data augmentation response message further includes second prompt information, where the second prompt information is used to indicate that the demonstration angle of the sub-action data is sufficient, or, the second prompt information is used to indicate that the demonstration angle of the sub-action data is insufficient, and where, if the second prompt information is used to indicate that the demonstration angle of the sub-action data is insufficient, the second prompt information includes information on the sub-action data with insufficient demonstration angle.

[0124] In a possible embodiment, in terms of creating a skill, the creation device 80 of the embodied skill is further specifically configured to: if the total amount of the current training data obtained after the first update is less than the preset data amount, display the first prompt window on the training data management interface; detect a trigger operation on the first determination control, create a third data augmentation request message; and send the third data augmentation request message to the cloud intelligent agent, where the third data augmentation request message is used to instruct the cloud intelligent agent to determine third training data, and the third training data is third action data that the embodied robot needs to simulate and train to achieve the target skill; receive a third data augmentation response message from the cloud intelligent agent, where the third data augmentation response message includes the third training data, display the third training data on the training data management interface, and update the total amount of the current training data obtained after the first update according to the third training data; if the total amount of the current training data obtained after the second update is greater than or equal to the preset data amount, and detect a training request of the user for the target skill, create a third skill training instruction; and send the third skill training instruction to the embodied robot, where the third skill training instruction is used to instruct the embodied robot to train the target skill according to the first training data, the second training data, and the third training data; receive a training feedback message from the embodied robot, where the training feedback message is used to indicate the completion degree of the embodied robot for the target skill; if the completion degree is greater than or equal to the preset completion degree, add a control for the target skill to the embodied robot skill demonstration interface.

[0125] Among them, the determination process of the second training data includes the following steps: determine the difference between the total amount of the current training data and the preset data amount to obtain a target difference; determine the target demonstration data according to the target difference, where the target demonstration data includes demonstration data of at least one modality in the multi-modal demonstration data; extract the feature data of the target demonstration data; find the corresponding action data according to the feature data to obtain P sixth action data; perform action capture on the multi-modal demonstration data to obtain seventh action data; determine the matching degree between the seventh action data and each of the P sixth action data to obtain P matching degrees; screen the P sixth action data according to the P matching degrees to obtain Q sixth action data, where Q is less than or equal to P; and determine the Q sixth action data as the second training data.

[0126] In a possible embodiment, after updating the total amount of the current training data obtained after the first update according to the third training data, the embodied skill creation device 80 is further specifically configured to: receive a request from the user to perform an editing operation on the target training data, and display a second prompt window on the training data management interface. The second prompt window includes third prompt information, a second confirmation control, and a second denial control. The third prompt information is used to prompt the user whether to continue to perform the editing operation. The target training data includes the first training data, and / or the second training data, and / or the third training data; detect a trigger operation on the second confirmation control, perform the editing operation, and update the training data management interface.

[0127] It can be understood that since the method embodiment and the device embodiment are different presentation forms of the same technical concept, therefore, the content of the method embodiment part in this application should be synchronously adapted to the device embodiment part, and will not be elaborated here.

[0128] In the case of adopting an integrated unit, please refer to Figure 9 , Figure 9 is a functional unit composition block diagram of another embodied skill creation device provided by the embodiments of the present application. As Figure 9 shown, the embodied skill creation device 80 includes: a processing module 802 and a communication module 801. The processing module 802 is used to control and manage the actions of the embodied skill creation device 80. For example, it executes the steps of the first processing unit 81, the display unit 82, and the second processing unit 83, and / or is used to execute other processes of the technologies described herein. The communication module 801 is used for the interaction between the embodied skill creation device 80 and other devices. As Figure 9 shown, the embodied skill creation device 80 may further include a storage module 803. The storage module 803 is used to store the program code and data of the embodied skill creation device 80.

[0129] Among them, the processing module 802 can be a processor or a controller. For example, it can be a Central Processing Unit (CPU), a general-purpose processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in connection with the disclosure of this application. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and so on. The communication module 801 can be a transceiver, an RF circuit, a communication interface, or the like. The storage module 803 can be a memory.

[0130] Among them, all relevant contents of each scenario involved in the above method embodiment can be cited in the function description of the corresponding functional module, and will not be elaborated here. The above-mentioned embodied skill creation device 80 can execute the above Figure 2 shown embodied skill creation method.

[0131] Please refer to Figure 10 , Figure 10 which is a schematic structural diagram of an electronic device proposed in an embodiment of this application. As Figure 10 shown, the electronic device 1000 includes a processor 1010, a memory 1020, a communication interface 1030, and one or more programs 1021. The above one or more programs 1021 are stored in the above memory and are configured to be executed by the above processor. When the program is executed, it includes some or all of the steps of any of the embodied robot control methods described in the above method embodiment. The processor, the memory, and the communication interface are interconnected and complete communication work with each other.

[0132] Among them, the memory can be a volatile memory such as a Dynamic Random Access Memory (DRAM), or a non-volatile memory such as a mechanical hard disk. The above memory is used to store a set of executable program codes. The above processor is used to call the executable program codes stored in the memory and can execute some or all of the steps of any embodied skill creation method described in the above embodiment of the embodied skill creation method.

[0133] It can be seen that for the electronic device 1000 described in the embodiments of the present application. It can be seen that in the embodiments of the present application, first, multimodal demonstration data for a target skill uploaded by a user is received. The target skill is a skill that the user determines needs to be newly added. And a first data augmentation request message is created. And the first data augmentation request message is sent to the cloud intelligent agent. The first data augmentation request message includes the multimodal demonstration data. The first data augmentation request message is used to instruct the cloud intelligent agent to determine first training data. The first training data is the first action data that the embodied robot needs to simulate and train to implement the target skill. Then, a first data augmentation response message from the cloud intelligent agent is received. The first data augmentation response message includes the first training data. The first training data is displayed on the training data management interface, and the total data volume of the current training data is recorded. If the total data volume of the current training data is greater than or equal to a preset data volume, and a training request from the user for the target skill is detected, a first skill training instruction is created. And the first skill training instruction is sent to the embodied robot, so that the embodied robot trains the target skill according to the first training data, and the creation of the target skill is realized.

[0134] After the present application receives an instruction from the user to create a new skill, compared with the embodied robot training the skill with a large amount of demonstration data collected manually, in this solution, the terminal device sends the demonstration data uploaded by the user to the cloud intelligent agent. The cloud intelligent agent augments the demonstration data according to the demonstration data uploaded by the user, that is, a small amount of demonstration data is augmented to generate a large-scale and diverse demonstration data, reducing the cost of collecting demonstration data. At the same time, the similarity between the augmented demonstration data and the demonstration data uploaded by the user is ensured. After the cloud intelligent agent obtains diverse demonstration data through augmentation, it sends it to the terminal device, so that the terminal device sends a skill training instruction to the robot, and then the robot trains the skill according to the diverse demonstration data, thereby realizing the creation of a new skill, improving the learning efficiency of the new skill, and further improving the imitation learning efficiency of the embodied robot.

[0135] The embodiments of the present application further provide a computer storage medium. Wherein, the computer storage medium stores a computer program for electronic data exchange. The computer program enables the computer to execute some or all of the steps of any method recorded in the above method embodiments. The above computer includes an electronic device.

[0136] The embodiments of the present application also provide a computer program product. The computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute some or all of the steps of any of the methods described in the foregoing method embodiments. The computer program product may be a software installation package, and the computer includes an electronic device.

[0137] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, some steps may be in other sequences or performed simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0138] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0139] In several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.

[0140] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0141] In addition, the functional units in the various embodiments of the present application may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software program module.

[0142] If the integrated unit is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned memory includes various media that can store program codes, such as USB flash drives, read-only memories (ROM), random access memories (RAM), mobile hard disks, magnetic disks, or optical discs.

[0143] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable memory, and the memory can include: flash drives, read-only memories, random access memories, magnetic disks, or optical discs, etc.

[0144] The above has introduced the embodiments of this application in detail. Specific examples are used in this article to elaborate on the principles and embodiments of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific embodiments and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A method for creating embodied skills, characterized in that, A terminal device applied to a skill creation system, the skill creation system further including an embodied robot and a cloud agent, the cloud agent being deployed on at least one server, the method comprising: Receiving multimodal demonstration data uploaded by a user for a target skill, the target skill being a skill that the user determines needs to be newly added; and creating a first data augmentation request message; and sending the first data augmentation request message to the cloud agent, the first data augmentation request message including the multimodal demonstration data, the first data augmentation request message being used to instruct the cloud agent to determine first training data, the first training data being first action data that the embodied robot needs to simulate and train to implement the target skill; wherein, the determination process of the first training data includes the following steps: performing action capture and audio capture on the multimodal demonstration data to obtain fourth action data and first audio data; decomposing the fourth action data into single actions to obtain a plurality of first sub-action data; and performing keyword search on each of the plurality of first sub-action data to obtain a plurality of second sub-action data; and parsing each of the first sub-action data and each of the plurality of second sub-action data to obtain sub-action training data; determining the waveform of the first audio data; and searching for M fifth action data corresponding to the first audio data according to the waveform; and determining the matching degree between the fourth action data and each of the M fifth action data to obtain M matching degrees; and screening the M fifth action data according to the M matching degrees to obtain N fifth action data, where N is less than or equal to M; determining the N fifth action data as fourth training data; determining the first training data according to the sub-action training data and the fourth training data; Receiving a first data augmentation response message from the cloud agent, the first data augmentation response message including the first training data, displaying the first training data on a training data management interface, and recording the total data volume of the current training data; If the total data volume of the current training data is greater than or equal to a preset data volume, and a training request from the user for the target skill is detected, creating a first skill training instruction; and sending the first skill training instruction to the embodied robot, so that the embodied robot trains the target skill according to the first training data to achieve the creation of the target skill.

2. The method according to claim 1, characterized in that, The method further includes: If the total data volume of the current training data is less than the preset data volume, then a first prompt window is displayed on the training data management interface, the first prompt window including a first prompt message, a first confirmation control, and a first denial control, the first prompt message being used to indicate to the user that the data volume of the current training data is insufficient and to determine whether to continue to augment the training data; Upon detecting a trigger operation on the first determination control, create a second data augmentation request message; and send the second data augmentation request message to the cloud agent, where the second data augmentation request message includes the total data volume of the current training data, and the second data augmentation request message is used to instruct the cloud agent to determine second training data, and the second training data is second action data that the embodied robot needs to simulate and train to achieve the target skill; Upon receiving a second data augmentation response message from the cloud agent, where the second data augmentation response message includes the second training data, display the second training data on the training data management interface, and update the total data volume of the current training data for the first time according to the second training data; If the total data volume of the current training data obtained after the first update is greater than or equal to a preset data volume, and a training request for the target skill is detected from the user, create a second skill training instruction; and send the second skill training instruction to the embodied robot, where the second skill training instruction is used to instruct the embodied robot to train the target skill according to the first training data and the second training data; Receive a training feedback message from the embodied robot, where the training feedback message is used to indicate the completion degree of the embodied robot for the target skill; If the completion degree is greater than or equal to a preset completion degree, add a control for the target skill to the embodied robot skill demonstration interface.

3. The method according to claim 2, wherein The method further includes: If the total data volume of the current training data obtained after the first update is less than the preset data volume, display the first prompt window on the training data management interface; Upon detecting a trigger operation on the first determination control, create a third data augmentation request message; and send the third data augmentation request message to the cloud agent, where the third data augmentation request message is used to instruct the cloud agent to determine third training data, and the third training data is third action data that the embodied robot needs to simulate and train to achieve the target skill; Upon receiving a third data augmentation response message from the cloud agent, where the third data augmentation response message includes the third training data, display the third training data on the training data management interface, and update the total data volume of the current training data obtained after the first update according to the third training data; If the total data volume of the current training data obtained after the second update is greater than or equal to the preset data volume, and a training request for the target skill is detected from the user, create a third skill training instruction; and send the third skill training instruction to the embodied robot, where the third skill training instruction is used to instruct the embodied robot to train the target skill according to the first training data, the second training data, and the third training data; Receive a training feedback message from the embodied robot, where the training feedback message is used to indicate the completion degree of the embodied robot for the target skill; If the completion degree is greater than or equal to the preset completion degree, a control for the target skill is added to the embodied robot skill demonstration interface.

4. The method according to claim 1, wherein The first data augmentation response message further includes second prompt information, where the second prompt information is used to indicate that the demonstration angle of the sub-action data is sufficient, or the second prompt information is used to indicate that the demonstration angle of the sub-action data is insufficient. If the second prompt information is used to indicate that the demonstration angle of the sub-action data is insufficient, the information of the sub-action data with insufficient demonstration angle is included in the second prompt information.

5. The method according to claim 2, wherein The process of determining the second training data includes the following steps: Determine the difference between the total data volume of the current training data and the preset data volume to obtain a target difference; Determine target demonstration data according to the target difference, where the target demonstration data includes demonstration data of at least one modality in the multi-modal demonstration data; Extract the feature data of the target demonstration data; Search for corresponding action data according to the feature data to obtain P sixth action data; Perform action capture on the multi-modal demonstration data to obtain seventh action data; Determine the matching degree between the seventh action data and each of the P sixth action data to obtain P matching degrees; Screen the P sixth action data according to the P matching degrees to obtain Q sixth action data, where Q is less than or equal to P; Determine the Q sixth action data as the second training data.

6. The method according to claim 3, wherein After updating the total data volume of the current training data obtained after the first update according to the third training data, the method further includes: Receiving a request from the user to perform an editing operation on the target training data, and displaying a second prompt window on the training data management interface. The second prompt window includes third prompt information, a second confirmation control, and a second denial control. The third prompt information is used to prompt the user whether to continue performing the editing operation. The target training data includes the first training data, and / or the second training data, and / or the third training data; Detecting a trigger operation on the second confirmation control, performing the editing operation, and updating the training data management interface.

7. An embodied skill creation device, characterized in that A terminal device applied to a skill creation system, where the skill creation system further includes an embodied robot and a cloud intelligent agent, and the cloud intelligent agent is deployed on at least one server, including: The first processing unit is configured to receive the multimodal demonstration data uploaded by the user for the target skill, where the target skill is the skill that the user determines needs to be added; and, create a first data augmentation request message; and, send the first data augmentation request message to the cloud intelligent agent, where the first data augmentation request message includes the multimodal demonstration data, and the first data augmentation request message is used to instruct the cloud intelligent agent to determine the first training data, where the first training data is the first action data that the embodied robot needs to simulate and train to achieve the target skill; wherein, the process of determining the first training data includes the following steps: perform action capture and audio capture on the multimodal demonstration data to obtain the fourth action data and the first audio data; decompose the fourth action data into a single action to obtain multiple first sub-action data; and, perform keyword search on each first sub-action data in the multiple first sub-action data to obtain multiple second sub-action data; and, parse each first sub-action data and each second sub-action data in the multiple second sub-action data to obtain sub-action training data; determine the waveform of the first audio data; and, search for M fifth action data corresponding to the first audio data according to the waveform; and, determine the matching degree between the fourth action data and each of the M fifth action data to obtain M matching degrees; and, screen the M fifth action data according to the M matching degrees to obtain N fifth action data, where N is less than or equal to M; determine the N fifth action data as the fourth training data; determine the first training data according to the sub-action training data and the fourth training data; The display unit is configured to receive the first data augmentation response message from the cloud intelligent agent, where the first data augmentation response message includes the first training data, display the first training data on the training data management interface, and record the total data volume of the current training data; The second processing unit is configured to, if the total data volume of the current training data is greater than or equal to the preset data volume and a training request for the target skill is detected by the user, create a first skill training instruction; and, send the first skill training instruction to the embodied robot, so that the embodied robot trains the target skill according to the first training data to achieve the creation of the target skill.

8. An electronic device, characterized in that, It includes a processor and a memory storing execution instructions, and the memory stores one or more programs; when the processor executes the execution instructions stored in the memory, the processor executes the method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, It stores an energy data management program, including execution instructions, and when the processor of the electronic device executes the execution instructions, the processor executes the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Natural Transfer of Knowledge Between Human and Artificial Intelligence

    US20180173999A1

  • Methods and systems for augmentation and feature cache

    US20240028963A1