Task Configuration Method and Related Devices for Humanoid Robots

By introducing terminal devices, VR wearable devices and servers into the humanoid robot control system, collecting user's task configuration information and operation data, generating sub-task sequences and atomic skill model parameters, the problem of insufficient flexibility in humanoid robot task execution is solved, and more abundant and flexible task execution capabilities are achieved.

CN119910668BActive Publication Date: 2025-06-20SHANGHAI FOURIER INTELLIGENCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510412752.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-06-20
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

When existing humanoid robots imitate human behavior to perform tasks, they usually can only perform pre-set fixed types of tasks, and cannot flexibly configure tasks according to the actual usage needs of users, resulting in insufficient flexibility in task execution.

Method used

By introducing terminal devices, virtual reality VR wearable devices and servers into the humanoid robot control system, the task configuration method is adopted to obtain the user's task configuration information, simulate task operations through the VR environment, collect user's posture data and video streams, generate sub-task sequences and atomic skill model parameters, and update the robot's task execution capabilities.

Benefits of technology

The diversity and flexibility of humanoid robot task execution is realized, which can better match the actual needs of users and improve the flexibility and diversity of task execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119910668B_ABST
    Figure CN119910668B_ABST
Patent Text Reader

Abstract

The present application discloses a task configuration method and related device for a humanoid robot, which are applied to a terminal device of a humanoid robot control system. The method includes: responding to an input operation based on a robot function configuration interface to obtain task configuration information; outputting a first prompt message to prompt the user to wear a VR wearable device for task operation demonstration; sending a first message to the VR wearable device to instruct the VR wearable device to display a VR environment corresponding to the task type; responding to a selection operation on a target control of the robot function configuration interface and sending a second message to the VR wearable device; receiving first pose data and a first video stream of the VR wearable device; sending a third message carrying the task type, the first video stream, and the first pose data to the server to instruct the server to send the association relationship between the task type and the subtask sequence and at least one atomic skill model parameter to the robot. The present application is beneficial to improving the diversity and flexibility of task execution of the humanoid robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of humanoid robot control, and particularly relates to a task configuration method and related device for a humanoid robot. Background Art

[0002] With the continuous development of robot technology and intelligent technology, robots have been widely used in various fields such as the service industry, education, medical care, entertainment, and research. Among them, humanoid robots, due to their similar appearance characteristics to humans and the ability to imitate human behavior to perform tasks, have become an important branch in the development of robots.

[0003] However, currently, when humanoid robots imitate human behavior to perform tasks, they usually can only perform pre-set fixed types of tasks and cannot flexibly configure tasks according to the actual usage needs of users, resulting in insufficient flexibility in task execution of humanoid robots. Summary of the Invention

[0004] Embodiments of this application provide a task configuration method and related device for a humanoid robot, aiming to improve the diversity and flexibility of task execution of the humanoid robot.

[0005] In a first aspect, embodiments of this application provide a task configuration method for a humanoid robot, which is applied to a terminal device in a humanoid robot control system. The humanoid robot control system includes: the terminal device, a virtual reality (VR) wearable device, a server, and the humanoid robot. The method includes:

[0006] In response to an input operation based on a robot function configuration interface, obtain task configuration information, where the task configuration information includes a task type;

[0007] Output a first prompt message for prompting the user to wear the VR wearable device to perform a task operation demonstration;

[0008] Send a first message to the VR wearable device, where the first message is used to instruct the VR wearable device to display a VR environment corresponding to the task type, and the VR environment is used to simulate the operation scenario of the task operation demonstration;

[0009] In response to a selection operation on a target control in the robot function configuration interface, send a second message to the VR wearable device, where the second message is used to instruct the VR wearable device to send the first pose data of the user collected during the task operation demonstration and a first video stream corresponding to the actual display content of the VR wearable device;

[0010] Receive the first pose data and the first video stream from the VR wearable device;

[0011] Send third information carrying the task type, the first video stream, and the first pose data to the server. The third information is used to instruct the server to send the association relationship between the task type and the subtask sequence, and at least one atomic skill model parameter to the robot. The subtask sequence is obtained by performing task splitting processing on the first video stream. The subtask sequence includes multiple subtasks and the execution order of the multiple subtasks. The atomic skill model parameter is determined according to the first pose data, and the atomic skill model parameter is used to update at least one atomic skill corresponding to the robot and the multiple subtasks.

[0012] In a second aspect, an embodiment of the present application provides a task configuration device for a humanoid robot, which is applied to a terminal device in a humanoid robot control system. The humanoid robot control system includes: the terminal device, a virtual reality (VR) wearable device, a server, and the humanoid robot. The device includes:

[0013] An acquisition unit, configured to acquire task configuration information in response to an input operation based on a robot function configuration interface. The task configuration information includes a task type.

[0014] An output unit, configured to output a first prompt message, where the first prompt message is used to prompt the user to wear the VR wearable device to perform a task operation demonstration.

[0015] A first sending unit, configured to send first information to the VR wearable device. The first information is used to instruct the VR wearable device to display a VR environment corresponding to the task type, and the VR environment is used to simulate an operation scenario of the task operation demonstration.

[0016] A second sending unit, configured to send second information to the VR wearable device in response to a selection operation on a target control in the robot function configuration interface. The second information is used to instruct the VR wearable device to send the first pose data of the user collected during the task operation demonstration and a first video stream corresponding to the actual display content of the VR wearable device.

[0017] A receiving unit, configured to receive the first pose data and the first video stream from the VR wearable device.

[0018] A third sending unit, configured to send third information carrying the task type, the first video stream, and the first pose data to the server, where the third information is used to instruct the server to send the association relationship between the task type and the subtask sequence, and at least one atomic skill model parameter to the robot. The subtask sequence is obtained by performing task splitting processing on the first video stream, and includes a plurality of subtasks and the execution order of the plurality of subtasks. The atomic skill model parameter is determined according to the first pose data, and is used to update at least one atomic skill corresponding to the robot and the plurality of subtasks.

[0019] In a third aspect, an embodiment of the present application provides a terminal device, where the terminal device includes:

[0020] A processor, a memory, and a communication interface, where the processor, the memory, and the communication interface are connected to each other. The communication interface is configured to receive or send data, the memory is configured to store application program code for the terminal device to execute the method in the first aspect above, and the processor is configured to execute the method in the first aspect above.

[0021] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program for electronic data exchange, and the computer program causes a computer to execute the steps of the task configuration method of the humanoid robot as described in the first aspect.

[0022] In a fifth aspect, an embodiment of the present application provides a computer program product, where the computer program product includes a computer program, and the computer program is operable to cause a computer to execute some or all of the steps described in the task configuration method of the humanoid robot as described in the first aspect of the embodiments of the present application. The computer program product may be a software installation package.

[0023] It can be seen that in the embodiments of the present application, the terminal can obtain task configuration information based on interface input operations, and output prompt information to prompt the user to wear a VR wearable device for task operation demonstration, and send information to the VR wearable device to instruct it to display a VR environment corresponding to the task type of the task configuration information to simulate the operation scenario of the task operation demonstration. Then, the terminal can also respond to the selection operation for the target control, send information to the VR wearable device, instruct it to send the user posture data collected during the task operation demonstration and the video stream corresponding to the actual display content of the VR wearable device, and send the received posture data, video stream and the obtained task type from the VR wearable device to the server. The server sends the association relationship between the task type and the subtask sequence and at least one atomic skill model parameter to the humanoid robot. The subtask sequence is obtained by performing task splitting processing on the video stream, including multiple subtasks and the execution order of multiple subtasks. The atomic skill model parameter is determined based on the posture data and is used to update at least one atomic skill of the robot corresponding to the multiple subtasks. In this way, the terminal can receive the task types of the humanoid robot customized by the user, and can generate the subtask sequence of the humanoid robot in the task type by obtaining the posture data obtained by the user in the actual operation demonstration of the customized task type and the video stream of the actual display screen of the VR wearable device, and update the robot atomic skills corresponding to the subtasks in the subtask sequence based on the actual user posture data, so that the task types that the robot can execute are more abundant, and the task execution can better match the actual needs of the user, which is beneficial to improving the diversity and flexibility of the task execution of the humanoid robot. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0025] Figure 1 is a schematic architecture diagram of a task configuration system for a humanoid robot provided by an embodiment of the present application;

[0026] Figure 2 is a schematic structural diagram of a terminal device provided by an embodiment of the present application;

[0027] Figure 3 is a schematic flowchart of a task configuration method for a humanoid robot provided by an embodiment of the present application;

[0028] Figure 4 is a schematic diagram of a robot function configuration interface provided by an embodiment of the present application;

[0029] Figure 5 It is a schematic diagram of a robot function selection interface provided by an embodiment of the present application;

[0030] Figure 6 It is a schematic diagram of a prompt information display interface provided by an embodiment of the present application;

[0031] Figure 7 It is a schematic diagram of another prompt information display interface provided by an embodiment of the present application;

[0032] Figure 8 It is a schematic diagram of a demonstration video editing interface provided by an embodiment of the present application;

[0033] Figure 9 It is a schematic diagram of another demonstration video editing interface provided by an embodiment of the present application;

[0034] Figure 10 It is a block diagram of the functional units of a task configuration device for a humanoid robot provided by an embodiment of the present application;

[0035] Figure 11 It is a block diagram of the functional units of another task configuration device for a humanoid robot provided by an embodiment of the present application. Detailed implementation manners

[0036] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0037] The terms "first", "second", etc. in the specification and claims of the present application and the above accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0038] Reference to "embodiments" in this document means that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0039] An embodiment of the present application provides a method and related device for task configuration of a humanoid robot. The embodiments of the present application will be introduced in detail below with reference to the accompanying drawings.

[0040] Please refer to Figure 1 , Figure 1 FIG. is a schematic architecture diagram of a task configuration system for a humanoid robot provided by an embodiment of the present application. The method for task configuration of a humanoid robot and the device for task configuration of a humanoid robot in the embodiments of the present application can be applied to the task configuration system for a humanoid robot. Among them, the task configuration system 10 for a humanoid robot may include a terminal device 101, a server 102, a virtual reality (VR) wearable device 103, and a humanoid robot 104. Any two devices in the task configuration system 10 for a humanoid robot can be directly or indirectly connected through a wired or wireless communication method, which is not limited in the embodiments of the present application.

[0041] In the embodiments of the present application, the terminal device 101 is mainly responsible for obtaining interactions with the user and the VR wearable device, and then obtaining the task configuration information, the video stream corresponding to the user posture data and the actual display content of the VR wearable device during the task operation demonstration, and sending it to the server for processing. The user can interact with the system through the terminal device. For example, input operations and selection operations are performed on the display interface of the terminal device, thereby triggering the terminal device to execute steps such as obtaining task configuration information and sending information to the VR wearable device. The VR wearable device is mainly responsible for displaying the VR environment for simulating the task operation demonstration scenario, collecting the posture data of the user during the task operation demonstration based on the VR environment, and collecting the video stream corresponding to the actual display content of the VR wearable device during the task operation demonstration, and sending it to the terminal device. The server is mainly responsible for processing the video stream and posture data from the terminal device. For example, the video stream is split into tasks to generate corresponding subtask sequences, and the atomic skill model parameters are determined according to the posture data to update the atomic skills of the humanoid robot corresponding to the subtasks in the subtask sequence.

[0042] In a specific implementation, after the humanoid robot 104 receives the association relationship between the task type and the subtask sequence from the server 102, as well as at least one atomic skill model parameter, it can update the atomic skill corresponding to the subtask associated with the subtask sequence based on the atomic skill model parameter, and can save the association relationship between the subtask sequence and the task type. Subsequently, the user can select the configured task type through the display interface of the terminal device 101. After the terminal device 101 confirms the task type to be executed based on the user's operation, it can send a task execution request carrying the task type to the humanoid robot 104. After receiving the task execution request, the humanoid robot 104 can execute the user-defined task type.

[0043] Based on this system, the terminal can receive the user-defined task type of the humanoid robot, and can generate the subtask sequence of the humanoid robot in this task type by obtaining the posture data obtained by the user in the actual operation demonstration of the user-defined task type and the video stream of the actual display screen of the VR wearable device, and update the robot atomic skill corresponding to the subtask of the subtask sequence based on the actual posture data of the user, so that the task types that the robot can execute are more abundant, and the task execution can better match the actual needs of the user, which is beneficial to improving the diversity and flexibility of the robot task execution.

[0044] Of course, in other embodiments, the terminal device 101 and the server 102 can also be the same device, that is, a single computer can support receiving the input operation and selection operation of the user to obtain the task configuration information and trigger the interaction with the VR wearable device, and support processing the posture data and video stream data collected by the VR wearable device to generate the subtask sequence and the atomic skill model parameter, and sending them to the robot.

[0045] It can be understood that Figure 1 the forms and quantities of the terminal device 101, the server 102, the VR wearable device 103 and the humanoid robot 104 shown in

[0046] In this application, the terminal can be various types of terminals such as a laptop computer, a tablet computer, a desktop computer, a set-top box, a smart phone, a smart speaker, a smart watch, a smart TV, a vehicle-mounted terminal, etc. The server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The VR wearable device can include, for example, a VR head-mounted display, VR hand controllers, VR gloves, a VR full-body tracking device, earphones and other accessories to support the VR wearable device to display a VR environment and support the VR wearable device to collect the pose data of the user (such as the pose data of the joints or parts of the user's hands, head, wrists, body, etc.). The humanoid robot can be a robot with a body structure similar to that of a human, capable of performing functional operations similar to human behaviors such as walking, grasping, and turning the head, and can also be equipped with various sensor systems such as cameras, microphones, and touch sensors for perceiving the environment.

[0047] Please refer to Figure 2 , the composition structure of the terminal device in this application can be as Figure 2 shown. The terminal device can include a processor 110, a memory 120, a communication interface 130, and one or more programs 121. Among them, the one or more programs 121 are stored in the above-mentioned memory 120 and are configured to be executed by the above-mentioned processor 110. The one or more programs 121 include instructions for executing any step performed by the terminal device in the following method embodiments. Among them, the communication interface 130 is used to support the communication between the terminal device and other devices. In a specific implementation, the processor 110 is used to execute any step performed by the terminal device in the following method embodiments, and when performing data transmission such as sending, the communication interface 130 can be selectively called to complete the corresponding operation. It should be noted that the structural schematic diagram of the above terminal device is only an example, and the specific components included can be more or less, and there is no unique limitation here.

[0048] Please refer to Figure 3 , Figure 3 is a schematic flowchart of a task configuration method for a humanoid robot provided by an embodiment of this application. The task configuration method for the humanoid robot can be applied to a terminal device in a task configuration system of a humanoid robot as Figure 1 shown. The humanoid robot control system includes: a terminal device, a virtual reality VR wearable device, a server, and a humanoid robot. As Figure 3 shown, the task configuration method for this humanoid robot includes the following steps:

[0049] S201. The terminal device obtains task configuration information in response to an input operation based on the robot function configuration interface.

[0050] Wherein, the task configuration information includes a task type.

[0051] In a specific implementation, a humanoid robot control application program can run on the terminal device. Specifically, the user can use this humanoid robot control application program to control the robot and configure the functions of the robot.

[0052] For example, the robot function configuration interface can be as Figure 4 shown. In this robot function configuration interface, an input box corresponding to the task type can be displayed. The user can complete the setting of the task type by entering a task name in the input box.

[0053] In a specific implementation, the display of the robot function configuration interface can also be triggered based on an operation detected on a control of another interface. For example, before step S201, the electronic device can also display a robot function selection interface. This robot function selection interface can display the system preset functions and user-defined functions of the robot. A "New Task" control can also be displayed on this user-defined function interface. The terminal device can also jump to the robot function configuration interface in S201 when detecting a preset operation (such as a click operation) on the "New Task" control.

[0054] In addition, taking Figure 5 the shown robot function selection interface as an example, a "System Preset Function" control and a "Custom Function" control can also be displayed at the top of this function selection interface. When the terminal device detects that the user clicks the "System Preset Function" control, it can display a list of system preset humanoid robot functions. If the terminal device detects that the user clicks the "Custom Function" control, it can switch to display a list of task types created by the user, and this "New Task" control. Specifically, Figure 5 taking the detection that the user clicks this "Custom Function" control as an example for illustration. At this time, the function selection page can display a list of task types created by the user. For example Figure 5 it includes a function list of multiple created task types such as "Go downstairs to pick up the express delivery", "Clean the living room", and "Clean the kitchen". The terminal device can send a task execution request carrying the selected created task type information to the corresponding humanoid robot when detecting a preset operation (such as a click operation) on any of the created task types, and then the selected created task type information can be executed by the humanoid robot.

[0055] It is understandable that the interface schematic diagrams provided in the embodiments of the present application are all for exemplary illustration. In actual applications, the positions, quantities, and presentation forms of the controls in each functional interface (such as the aforementioned robot function configuration interface and robot function selection interface) can be set differently from the interface schematic diagrams provided in the embodiments of the present application. For example, the "New Task" control in the robot function selection interface can also be set in the upper right corner of the interface, or any other position, or the "New Task" control can also be directly displayed as a "+" control, and no specific limitation is made here.

[0056] Furthermore, task configuration prompt information can also be displayed in the robot function configuration interface. The task configuration prompt information can be generated according to the atomic skills of the humanoid robot. Specifically, the terminal device can send identification information such as the model of the humanoid robot to the server. The server queries the task types of the created functions corresponding to the identification information of the humanoid robot uploaded by the user as the reference task types, and sends the obtained reference task types to the terminal device. The terminal device can then generate task configuration prompt information based on the reference task types.

[0057] In particular, the terminal device can obtain a task configuration prompt information template, use the reference task type as a keyword, and fill it into the information template to finally generate the task configuration prompt information. Taking Figure 4 the shown robot function configuration page as an example, in other embodiments, in the input box corresponding to the task type, in addition to displaying the prompt information of "Please enter the task name", task configuration prompt information such as "Try to create a task to let the robot help you collect express delivery" can also be displayed.

[0058] In particular, please continue to refer to Figure 4 , when obtaining the task configuration information, the terminal device can also display a robot list for the user. The robot list includes the robots associated with the user's logged-in account or bound to the terminal device. The terminal device can establish an association relationship between the robot and the task type created by the user according to the user's selection operation on the robots in the robot list. Then, in scenarios such as sending the robot identification information to the server, the terminal device can determine the robot identification information to be sent based on the association relationship from multiple robot identification information.

[0059] Correspondingly, in other embodiments, when the terminal device displays the list of created tasks and the system preset task list, it can also generate a task sub-list corresponding to each robot according to the association relationship between the robot and each task, and display the tasks corresponding to different robots in a partitioned manner, which is convenient for the user to view and operate.

[0060] S202, the terminal device outputs the first prompt information.

[0061] Among them, the first prompt message is used to prompt the user to wear the VR wearable device for task operation demonstration.

[0062] Among them, the first prompt message can be output through a display interface, or the first prompt message can also be other information such as audio information or vibration information.

[0063] For example, taking the first prompt message being output to the user through the display interface as an example, in S202, the terminal device can, for example, output the first prompt message by displaying an interface as shown below. By displaying a text prompt message such as "The VR device is connected. Please wear the VR device to start the task operation demonstration", it prompts the user to wear the VR wearable device for task operation demonstration. Or, in other embodiments, the terminal device can also prompt by outputting a voice such as "Please wear the VR device to start the task operation demonstration". It can be understood that the specific text content and voice content of the prompt message here are only for illustrative purposes. In actual applications, other forms of prompt messages can also be selected, and no matter what form of prompt message is used, the specific content of the prompt message can be set according to needs, and no specific limitation is made here. Figure 6

[0064] In specific implementation, before step S202, the terminal device can also first detect whether a connection has been established with the VR wearable device. If it detects that there is no connection with the VR wearable device, it can also jump to display the VR wearable device connection setting interface to prompt the user to perform VR wearable device connection setting. After the terminal device detects that the user's operation has established a connection with the VR wearable device, it can output the first prompt message.

[0065] Figure 7 In specific implementation, after step S201 and before step S202, the terminal device can also display an interface as shown below to prompt the user whether to start the task operation demonstration. Specifically, the terminal device can display this interface when it detects a confirmation operation for the task configuration information. Still taking the robot function configuration interface as an example, the confirmation operation can be a selection operation for the "Confirm" control. Specifically, as shown below, the terminal device can display information such as "The task has been created. Do you want to start the task operation demonstration immediately". If it detects that the user selects the "Start now" control, it can execute step S202, for example, detect whether the VR wearable device is connected and jump to display the interface as shown below when connected. If the terminal device detects that the user selects the "Wait a while" control, it does not need to continue to execute steps S202 - S206, and the terminal device can jump to display an interface as shown below. Figure 7 Figure 4 Figure 7 Figure 6 Figure 5 ​​​​​​The shown robot function selection interface can also display a list of tasks to be demonstrated. The list of tasks to be demonstrated includes task types that the user has only created tasks in the robot function configuration interface but has not completed the task demonstration operation. For example Figure 5 "Water the flowers on the balcony" and "Take out the trash" in Figure 5 . When the terminal device detects a selection operation for the task type in the list of tasks to be demonstrated again later, steps S201 - S206 can be executed again.

[0066] S203, the terminal device sends a first message to the VR wearable device.

[0067] Among them, the first message is used to instruct the VR wearable device to display a VR environment corresponding to the task type, and the VR environment is used to simulate the operation scenario of the task operation demonstration.

[0068] In a specific implementation, for example, after extracting keywords from the task type in the task configuration information input by the user, the keywords are sent to the server. The server determines a matching VR environment based on the keywords and feeds back the environment data of the VR environment or the environment identifier of the VR environment. Then, the electronic device carries the environment data or the environment identifier in the first message and sends it to the VR wearable device for display. Or, the VR wearable device can save a VR general environment. When it does not receive the environment data or the environment identifier from the electronic device, it can directly use the VR general environment as the VR environment corresponding to the task type for display.

[0069] S204, in response to a selection operation on the target control in the robot function configuration interface, the terminal device sends a second message to the VR wearable device.

[0070] Among them, the second message is used to instruct the VR wearable device to send the first pose data of the user collected during the task operation demonstration and the first video stream corresponding to the actual display content of the VR wearable device.

[0071] In a specific implementation, the pose data of the user can specifically be the pose data of the user's hands, head, wrists, body and other joints or parts collected by the VR wearable device. The video stream corresponding to the actual display content of the VR wearable device is the video stream corresponding to the display screen of the VR wearable device when the user wears the VR wearable device and interacts with the VR environment during the task operation demonstration.

[0072] In a specific implementation, the selection operation can be, for example, a click operation, or it can also be a long - press operation for a preset duration, etc. The target control can be, for example Figure 6For the "Complete Demonstration" control shown in [Figure 0], after the terminal device detects that the "Complete Demonstration" control is selected, it can send a second message to the VR wearable device, and the second message can also instruct the VR wearable device to stop collecting attitude data and video streams.

[0073] In addition, Figure 6 a "Quit Demonstration" control can also be displayed in [Figure 0]. After the terminal device detects the "Quit Demonstration" control, it can also output a save prompt message, prompting the user to select whether to save the previously collected user attitude information and video streams. If a selection operation for the save control is detected, it can also send a second message to the VR wearable device, save the received user attitude information and video streams, and can also jump to display Figure 5 the robot function selection interface of [Figure 0], update the task type information in the to-be-demonstrated task list in the robot function selection interface, and when a selection operation for the task type in the to-be-demonstrated task list is detected subsequently and steps S201 - S206 are restarted, based on the actual display image of the VR wearable device recorded when exiting the demonstration previously, determine the current screen that needs to be displayed by the VR wearable device, and confirm the interaction states of each interaction object in the VR environment, so that the user can continue the operation demonstration based on the display screen when exiting the demonstration last time through the VR wearable device, and the interaction objects in the interaction states when starting the operation demonstration are also the same as those when exiting the demonstration last time, improving the task configuration efficiency and saving system resources.

[0074] In specific implementation, the VR wearable device can specifically stop collecting attitude data and video streams when receiving the second message, and start displaying the VR environment when detecting a connection with the terminal device or receiving a start request from the terminal device. The start request can be issued by the terminal device when detecting a user's selection operation for the Figure 7 "Start Now" control in [Figure 0].

[0075] S205. The terminal device receives the first attitude data and the first video stream from the VR wearable device.

[0076] S206. The terminal device sends a third message carrying the task type, the first video stream, and the first attitude data to the server.

[0077] Among them, the third message is used to instruct the server to send the association relationship between the task type and the sub-task sequence, and at least one atomic skill model parameter to the robot. The sub-task sequence is obtained by performing task splitting processing on the first video stream. The sub-task sequence includes multiple sub-tasks and the execution order of the multiple sub-tasks. The atomic skill model parameter is determined based on the first attitude data, and the atomic skill model parameter is used to update at least one atomic skill corresponding to the robot and the multiple sub-tasks.

[0078] In a specific implementation, the server can split the first video stream into a sequence of subtasks, each sequence of subtasks corresponding to a skill group, and each skill group corresponding to a set of atomic skills that need to be mapped to the humanoid robot. The set of atomic skills includes at least one atomic skill corresponding to the multiple subtasks.

[0079] Taking the task type of "taking out the trash" as an example, the multiple subtasks included in the sequence of subtasks may specifically include: opening the entrance door, going out, grabbing the garbage bag, closing the door, walking from the entrance door to the elevator entrance, pressing the elevator down button, waiting for the elevator door to open and entering the elevator, pressing the button for the 1st floor, waiting for the elevator door to open and the elevator display to show the 1st floor, getting out of the elevator, walking to the building entrance door, pushing the door open after clicking the door opening button, walking to the garbage disposal point, putting the garbage bag into the garbage entrance, walking to the building entrance door, waiting for the door to open, entering the building, walking to the elevator entrance, detecting and pressing the up button, waiting for the elevator door to open and entering, pressing the button for the target floor such as the 6th floor, waiting for the elevator door to open and the display to show the 6th floor, getting out, walking to the entrance door, waiting for the door to open, and entering. Among them, the execution sequence relationship between adjacent subtasks is constructed by the sequence of subtasks and synchronized to the humanoid robot.

[0080] Among them, the atomic skills corresponding to the above subtasks, that is, the atomic skills of the robot required in each subtask, can be pre-trained. The server can create a fine-tuning data set based on the first pose data and train each atomic skill required in the current task type of the humanoid robot according to the fine-tuning data set, that is, at least one atomic skill corresponding to the multiple subtasks, so that each atomic skill can better adapt to the actual needs of the user. After the server trains the humanoid robot with at least one atomic skill corresponding to the multiple subtasks based on the fine-tuning data set, the new model parameters of each trained atomic skill, that is, at least one atomic skill model parameter in the embodiments of the present application, can be obtained and sent to the humanoid robot, so that the humanoid robot can complete the user-defined task type subsequently.

[0081] It can be seen that in the embodiments of the present application, the terminal can obtain task configuration information based on interface input operations, and output a prompt message to prompt the user to wear a VR wearable device for task operation demonstration, and send information to the VR wearable device to instruct it to display a VR environment corresponding to the task type of the task configuration information to simulate the operation scenario of the task operation demonstration. Then, the terminal can also respond to the selection operation for the target control, send information to the VR wearable device, instruct it to send the user posture data collected during the task operation demonstration and the video stream corresponding to the actual display content of the VR wearable device, and send the received posture data, video stream and the obtained task type from the VR wearable device to the server. The server sends the association relationship between the task type and the subtask sequence and at least one atomic skill model parameter to the humanoid robot. The subtask sequence is obtained by performing task splitting processing on the video stream, including multiple subtasks and the execution order of the multiple subtasks. The atomic skill model parameter is determined according to the posture data and is used to update at least one atomic skill of the robot corresponding to the multiple subtasks. In this way, the terminal can receive the task type of the humanoid robot customized by the user, and can generate the subtask sequence of the humanoid robot in the task type by obtaining the posture data obtained by the user in the actual operation demonstration of the customized task type and the video stream of the actual display screen of the VR wearable device, and update the robot atomic skill corresponding to the subtask in the subtask sequence based on the actual user posture data, so that the task types that the robot can execute are more abundant, and the task execution can better match the actual needs of the user, which is beneficial to improving the diversity and flexibility of the task execution of the humanoid robot.

[0082] In a possible example, the task configuration information further includes task environment description information, and the VR environment is determined through the following steps: determining a VR reference environment corresponding to the task type from a preset plurality of VR basic environments; determining target environmental elements according to the task environment description information, where the target environmental elements include at least one of the following: building form, form of interactive object; adjusting the VR reference environment according to the target environmental elements to obtain the VR environment, and the VR environment includes the target environmental elements.

[0083] Among them, the building form may include, for example, the building appearance, and the interactive object specifically refers to an object that can interact with the user or the user can operate.

[0084] In a specific implementation, the terminal device may send the task type input by the user to the server, and the server determines a VR reference environment corresponding to the task type from multiple VR preset environments; alternatively, the terminal device directly determines a VR reference environment corresponding to the task type according to the task type input by the user. Specifically, each VR preset environment is set with basic environment information, which is used to represent information such as building types of the VR preset environment, such as residential buildings, commercial buildings, schools, etc., appearance information such as single-story or multi-story, and functional information such as elevators, stairways, and underground parking lots. The terminal device or the server can extract keywords from the task type and then determine a VR reference environment that matches the task. For example, when the task type is "go downstairs to pick up the express delivery", the VR preset environments of single-story buildings should be excluded.

[0085] In addition, in other embodiments, the task environment description information can be used not only to determine the target environmental elements but also to jointly determine the VR reference environment with the task type.

[0086] In a specific implementation, the VR reference environment is adjusted according to the target environmental elements to obtain a VR environment. Specifically, when the target environmental elements include the building form, the original form of the building in the VR reference environment is modified according to the building form. For example, if the original building in the VR reference environment is a 10-story building and the building form information indicates that the VR reference environment is a 40-story building, the building in the VR reference environment can be modified to a 10-story building. At this time, if the original building in the reference environment includes interactive objects whose interaction content is affected by the building form, such as an elevator, the interaction content of the interactive object also needs to be modified according to the building form information, such as adding floor buttons to the elevator. When the target environmental elements include the form of the interactive object, the form of the interactive object in the VR reference environment can also be modified according to the form of the interactive object. Taking the form of the interactive object included in the target environmental elements being a pedal trash can as an example, if there is no interactive object such as a trash can in the VR reference environment, a pedal trash can can be directly added. If there is a flip-top trash can in the VR reference environment, it can be modified to a pedal trash can.

[0087] It can be seen that in this example, the VR environment can specifically be obtained by determining the target environmental elements according to the task environment description information in the task configuration information and then adjusting the VR reference environment determined according to the task type in the task configuration information according to the target environmental elements. The target environmental elements specifically include at least one of the building form and the form of the interactive object. The VR environment adjusted by the target environmental elements is more in line with the task execution environment expected by the user. Based on this VR environment, task operation demonstrations are carried out, and the subtask sequence and atomic skill model parameters of the robot are generated, which can better adapt to the actual needs of the user and is conducive to further improving the diversity and flexibility of the robot's task execution.

[0088] In a possible example, the environmental description information includes a target image, and the form of the interactive object is determined through the following steps: identifying the target image to determine at least one reference interactive object included in the target image; determining the interaction method of the at least one reference interactive object; determining a scene interactive object from the at least one reference interactive object according to the interaction method and the atomic skills corresponding to the robot, where the atomic skills corresponding to the robot support the interaction method of the scene interactive object; and determining the form of the interactive object according to the scene interactive object.

[0089] In specific implementation, when the terminal device identifies the target image, it can first perform object recognition on the target image to determine the objects included in the target image, and further determine the object types of the target objects, so as to determine whether the object is a building or an interactive object. When it is determined that the object is a building, the building form can be obtained according to the building. If the object is an interactive object, it can be determined as a reference interactive object, and further the scene interactive object is screened out and the form of the interactive object is determined.

[0090] Among them, the atomic skills corresponding to the robot support the interaction method of the scene interactive object, that is, the robot can realize the interaction with the scene interactive object by using the atomic skills. For example, if the robot atomic skills do not support inserting a key into the keyhole and rotating to unlock, but only support swiping a card to open the door and pressing the doorbell, then if the reference interactive object in the target image includes a door that only supports key unlocking and has no doorbell, this door cannot be determined as a scene interactive object. If the reference interactive object in the target image includes a door with a doorbell or an intelligent lock that supports swiping a card to open the door, then this door can be determined as a scene interaction. Or, in other embodiments, for an interactive object in the reference interactive object that is not supported by the atomic skills corresponding to the robot, the terminal device can also modify the interaction method of the reference interactive object according to the atomic skills corresponding to the robot and this interactive object, modify it to an interaction method supported by the robot atomic skills, and further determine the form of the interactive object. For example, modify the traditional mechanical lock door to an intelligent lock door and add it to the VR reference environment to obtain the final VR environment.

[0091] In addition, in specific implementation, the task environmental description information can be determined not only by obtaining through the target image, but also by detecting the user's selection operation for the preset task environment or receiving the text information input by the user. It can be understood that the task environmental description information can be a combination of any one or more of the above target image, text information, and the preset task environment selected by the user.

[0092] For example, still taking Figure 4Taking the robot function configuration interface shown as an example, the robot function configuration interface may further include a setting control for the task environment. Specifically, the user can expand a preset task environment list through the drop-down control of the environment task bar. For example Figure 4 in Figure 4 , through the drop-down control, the user can select environmental description information from "Staircase Residence", "Elevator Residence", "Single-story Restaurant", or the user can also perform text input or image input operations in the text input box or image upload box under the task environment list.

[0093] It can be seen that in this example, the environmental description information includes a target image. The terminal device can identify the target image to determine at least one reference interaction object included in the target image, and determine the interaction method of at least one reference interaction object. Then, according to the interaction method and the atomic skills corresponding to the robot, the scene interaction objects whose interaction methods are supported by the robot atomic skills are determined from at least one reference interaction object. Furthermore, the form of the interactable object is determined according to the scene interaction object, which is beneficial to improving the matching degree between the VR environment and the robot, thereby improving the accuracy of the determined robot subtask sequence and the reliability of robot task execution.

[0094] In a possible example, the VR environment includes multiple VR sub-environments. After receiving the first pose data and the first video stream data from the VR wearable device, the method further includes: displaying the display content corresponding to the first video stream on the demo video editing interface; in response to a selection operation on a target segment in the display content, sending a fourth message to the VR wearable device, where the fourth message is used to instruct the VR wearable device to redisplay the target VR sub-environment corresponding to the target segment, and the multiple VR sub-environments include the target VR sub-environment; outputting a second prompt message, where the second prompt message is used to prompt the user to wear the VR wearable device to perform a task operation re-demonstration; receiving second pose data and a second video stream from the VR wearable device, where the second pose data is the pose data of the user collected during the task operation re-demonstration, and the second video stream is the video stream corresponding to the actual display content of the VR wearable device collected during the operation re-demonstration; updating the first pose data and the first video stream according to the second pose data and the second video stream to obtain third pose data and a third video stream, where the third pose data includes the second pose data, and the third video stream includes the second video stream; sending a fourth message carrying the task type, the third video stream, and the third pose data to the server.

[0095] In a specific implementation, the specific method for updating the first pose data and the first video stream according to the second pose data and the second video stream to obtain the third pose data and the third video stream may be as follows: determining the pose sub-data corresponding to the target segment from the first pose data according to the start time and the end time corresponding to the target segment, and determining the video stream sub-data corresponding to the target segment from the first video stream according to the start time and the end time corresponding to the target segment; replacing the pose sub-data with the second pose data to obtain the third pose data, and replacing the video stream sub-data with the second video stream to obtain the second video stream.

[0096] That is to say, the user can operate the terminal device to re-demonstrate a segment of the pre-generated video stream, i.e., the first video stream, and replace the segment of the pre-generated video stream with the re-demonstrated second pose data and the second video stream. Among them, the target VR sub-environment, i.e., the part of the VR environment actually displayed by the VR wearable device at the start time of the target segment. For example, if the VR environment includes a complete residence and the start time of the target segment selected by the user is the elevator door on the 10th floor in the complete residence, the terminal device can determine that the target VR sub-environment is the elevator door on the 10th floor and instruct the VR wearable device to re-display the elevator door on the 10th floor through the fourth information.

[0097] In a specific implementation, after receiving the fourth information, the server can use the third video stream as the first video stream and the third pose data as the first pose data, and then execute step S206 according to the task type, the third video stream, and the third pose data to determine the correspondence between the task subsequence and the task type, and determine at least one atomic skill model parameter.

[0098] In a specific implementation, taking Figure 8 and Figure 9 the shown demonstration video editing interface as an example, when the terminal detects a click operation by the user on the shaded part segment shown in the shaded part of the white progress bar of Figure 6 or Figure 7 , if it detects a click operation on the "replacement control", it can confirm that the selection of the target segment is detected, and determine the shaded part segment clicked by the user as the target segment, and then send the fourth information to the VR wearable device. In particular, after the terminal device detects a click operation on the reference segment, it can also distinguish the selected reference segment from other reference segments by changing the display color of the clicked reference segment to prompt the user of the position of the selected reference segment. For example, Figure 8 as shown in

[0099] when the terminal device detects that the user clicks on the reference segment on the left side of the progress bar, the color of the reference segment is changed to a different color from the unselected reference segment on the right side for display.Figure 8 Or Figure 9 After a click operation at any time point in the white progress bar, if a click operation on the "insert control" is detected, the VR sub - environment corresponding to this time point can also be determined, and the sixth information is sent to the VR wearable device, instructing the VR wearable device to display the VR sub - environment corresponding to this time point and continue to collect the user's fourth posture data and the fourth video stream corresponding to the display content of the VR wearable device until a click operation on the stop control is detected. Then, the VR wearable device is notified to stop collecting, and the fourth posture data and the fourth video stream collected this time are sent to the terminal device. At this time, the terminal can also insert the fourth posture data into the first posture data according to the time point selected by the user, splice the first posture data and the fourth posture data to obtain the third posture data, insert the fourth video into the first video stream, and splice the fourth video stream and the first video stream to obtain the first video stream. That is, the user can also add new operation demonstration content to the previous demonstration content.

[0100] Correspondingly, Figure 8 And Figure 9 The "delete control" in can also support the user to delete content from the demonstration video. For example, the user can view the image frames corresponding to the redundant actions in the deleted VR image, and then send the deleted video stream and posture data to the server.

[0101] It can be seen that in this example, the terminal device can respond to the user's selection operation for the target segment, trigger the VR wearable device to redisplay the target VR sub - environment corresponding to the target segment, and re - collect the user's posture data and video stream. And the terminal device can update the original posture data and video stream according to the re - collected posture data and video stream, and send the updated posture data and video stream to the server. The user can adjust the original data based on requirements, so that the data for determining the sub - task sequence and the atomic skill model parameters by the server is more adapted to the user's needs, which is conducive to further improving the flexibility of the robot to execute tasks.

[0102] In a possible example, the displaying the display content corresponding to the first video stream on the demonstration video editing interface includes: receiving the fifth information from the robot, where the fifth information is used to indicate the sub - task that the robot fails to execute successfully when executing the sub - task sequence; determining a reference segment corresponding to the sub - task that fails to execute successfully from the first video stream; determining the reference segment as the display content; and displaying the display content on the demonstration video editing interface.

[0103] In a specific implementation, the server can determine the corresponding end condition for each sub - task. If the robot fails to meet the end condition corresponding to the sub - task within the preset time when executing the sub - task, it can be determined that the robot's execution fails.

[0104] In a specific implementation, still taking Figure 8 and Figure 9 as an example, after the terminal device receives the fifth piece of information from the robot and determines the reference segment, the way to display the reference segment in the demonstration video editing interface can be, for example: changing the display color of the position of the reference segment in the white progress bar, such as Figure 8 and Figure 9 the two shaded segment parts shown in the shaded part.

[0105] In addition, in other embodiments, the reference segment can also be set by the user's own marking. For example, Figure 9 as shown, the user can also select the time point to be located by dragging the time point positioning bar perpendicular to the white progress bar, and when detecting the user's selection operation for the "mark" control, display the mark type table through the mark type drop-down box, so as to select the mark type of this time point, and can output a segment confirmation prompt to prompt the user whether to determine the start frame and end frame of two adjacent segments as a reference video stream segment. After detecting the user's confirmation operation, the reference segment can be displayed in the white progress bar.

[0106] Particularly, the terminal device can also adjust the VR image displayed above the video boundary interface according to the position of the time point positioning bar, so that the display screen of the VR image is synchronized with the time point corresponding to the time positioning bar.

[0107] In practical applications, the terminal device can send both the first pose data and the first video stream to the server, and also send the third pose data and the third video stream to the server. Furthermore, the server can determine a sub-task sequence and atomic skill model parameters of a task type based on the first pose data and the first video stream, determine another sub-task sequence and atomic skill model parameters of this task type based on the third pose data and the third video stream, and can obtain the priorities set by the user for the first video stream and the third video stream based on the terminal device. Then, determine the sub-task sequence corresponding to the video stream with a higher priority as the default task sequence of this task type, and determine the alternative sub-tasks corresponding to each sub-task in the default task sequence from the sub-task sequence corresponding to the video stream with a lower priority. If there are alternative sub-tasks for the sub-tasks that are not successfully executed when the robot executes the task based on the default task sequence, the robot can execute the alternative sub-tasks.

[0108] For example, when a robot is performing a garbage-throwing task, after pressing the elevator button, the robot monitors the content displayed on the elevator panel in real time, analyzes the content displayed on the elevator panel, and determines whether the elevator is operating normally. If it is operating normally, it determines that the current subtask is successfully executed. If the elevator shows a fault, it determines that the current subtask fails. At this time, the robot determines whether there is an alternative solution. For example, if the alternative solution is that there is a freight elevator in the current building, the robot modifies the current subtask to take the freight elevator. At this time, the robot executes the subtask, retrieves the location of the freight elevator, plans a movement path based on the current location, and moves to the freight elevator based on the planned movement path to execute the subtask. At the same time, when the robot is performing the garbage-throwing task, it can also judge the capacity of the current trash can and the capacity of the current garbage to be thrown. Based on this, it judges whether it is feasible to throw the garbage into the current trash can. If it is not feasible, the robot can also generate a reason why the task cannot be executed and send the reason why the task cannot be executed to the terminal device. Then the robot abandons the task. If it receives a task execution instruction from the terminal device again based on the user's operation, it can execute the task again. Specifically, the terminal device can also prompt the user with the reason why the task cannot be executed, such as through a pop-up window, etc., to improve the richness of information display.

[0109] It can be seen that in this example, the terminal device can also determine, according to the fifth information from the robot, a reference segment corresponding to the subtask that has not been successfully executed from the first video stream, and determine the reference segment as the display content for display, which is beneficial to quickly and accurately prompt the user with the reference segment corresponding to the subtask that the robot has not successfully executed, and then accurately re-demonstrate the subtask that has not been successfully executed.

[0110] In a possible example, the subtask sequence further includes start conditions and end conditions for each of the multiple subtasks. The multiple subtasks include a first subtask, and the end condition corresponding to the first subtask is determined according to the target interaction state. The target interaction state is determined by identifying an end image frame corresponding to a video stream segment of the first subtask, and the target interaction state is the interaction state between the user and the target interaction object in the VR environment in the end image frame.

[0111] In specific implementation, the end condition can be used as reference information for determining whether the robot has successfully executed the subtask. If the robot cannot reach the end condition corresponding to the subtask within the preset time when executing the subtask, it can be determined that the robot has failed to execute.

[0112] For example, taking the first subtask of opening a door as an example, the server can identify based on the end image frame of the video stream segment corresponding to the first subtask, determine the door as the target interaction object, rotate the doorknob and push it outward so that the opening angle of the door is greater than the preset angle state as the target interaction state. Then, when the robot executes the first subtask, the robot's dexterous hand rotates the doorknob and pushes it outward so that the opening angle of the door is greater than the preset angle, and it is considered that the first subtask has been completed.

[0113] In a specific implementation, the server can split the first video stream through video stream splitting and image recognition processing to obtain video stream segments corresponding to each subtask. Alternatively, the server can also receive the video stream segments corresponding to each subtask determined by the terminal device through detecting the start image frame and end image frame operations of the user.

[0114] In a specific implementation, the start condition of each subtask can be determined by performing image recognition on the start image frame of the video stream segment corresponding to each subtask. For example, the start condition can also be determined according to the interaction state of the interaction object in the start image frame. Alternatively, the end condition of a subtask whose execution order is before the current subtask can also be determined as the start condition of the current subtask.

[0115] It can be seen that in this example, the subtask sequence also includes the start condition and end condition of each subtask in the multiple subtasks. The end condition corresponding to the first subtask is determined according to the target interaction state of the user and the target interaction object in the VR environment in the end image frame corresponding to the video stream segment corresponding to the first subtask. Determining the end condition and start condition of the subtask based on the detection of the image frames of the video stream segment corresponding to the subtask and sending them to the humanoid robot is beneficial to improving the accuracy and reliability of the humanoid robot in executing tasks.

[0116] In a possible example, the third information also carries the timestamp information of the end image frame, and the timestamp information is determined in the following manner: displaying the display content corresponding to the first video stream through the demonstration video editing interface; in response to a selection operation on the target image frame in the display content, determining the target image frame as the end image frame of the video stream segment; determining the timestamp information corresponding to the end image frame.

[0117] Among them, still taking Figure 9 the shown demonstration video editing interface as an example, the selection operation for the target image frame can be, for example, detecting a selection operation for the "mark" control and detecting a selection operation for "fragment end frame" in the "mark type". At this time, the image frame corresponding to the time point positioning bar is the target image frame, and the time corresponding to the time point positioning bar is the timestamp information.

[0118] In a specific implementation, the server can specifically determine the end image frame corresponding to the subtask according to the timestamp information sent by the terminal device.

[0119] It can be seen that in this example, the terminal device can determine the target image frame as the end image frame of the video stream segment according to the selection operation for the target image frame in the first video stream, and carry the timestamp information corresponding to the end image frame in the third information and send it to the server, making the setting of the end condition more in line with the actual needs of users and further improving the flexibility of the robot task execution.

[0120] Please refer to Figure 10 , Figure 10 which is a functional unit composition block diagram of a task configuration device for a humanoid robot provided by an embodiment of the present application, and can be applied to a terminal device in a task configuration system of a humanoid robot as shown in Figure 1 . The humanoid robot control system includes: a terminal device, a virtual reality (VR) wearable device, a server, and a humanoid robot. The task configuration device 30 of the humanoid robot includes:

[0121] An acquisition unit 301, configured to acquire task configuration information in response to an input operation based on a robot function configuration interface, where the task configuration information includes a task type;

[0122] An output unit 302, configured to output a first prompt message, where the first prompt message is used to prompt the user to wear the VR wearable device to perform a task operation demonstration;

[0123] A first sending unit 303, configured to send a first message to the VR wearable device, where the first message is used to instruct the VR wearable device to display a VR environment corresponding to the task type, and the VR environment is used to simulate an operation scenario of the task operation demonstration;

[0124] A second sending unit 304, configured to send a second message to the VR wearable device in response to a selection operation for a target control in the robot function configuration interface, where the second message is used to instruct the VR wearable device to send first attitude data of the user collected during the task operation demonstration and a first video stream corresponding to the actual display content of the VR wearable device;

[0125] A receiving unit 305, configured to receive the first attitude data and the first video stream from the VR wearable device;

[0126] A third sending unit 306 is configured to send third information carrying the task type, the first video stream, and the first pose data to the server. The third information is used to instruct the server to send the association relationship between the task type and the subtask sequence, and at least one atomic skill model parameter to the robot. The subtask sequence is obtained by performing task splitting processing on the first video stream, and includes a plurality of subtasks and the execution order of the plurality of subtasks. The atomic skill model parameter is determined according to the first pose data, and is used to update at least one atomic skill corresponding to the robot and the plurality of subtasks.

[0127] In a possible example, the task configuration information further includes task environment description information. The VR environment is determined through the following steps: determining a VR reference environment corresponding to the task type from a plurality of preset VR basic environments; determining target environmental elements according to the task environment description information, where the target environmental elements include at least one of the following: building form, form of interactive object; adjusting the VR reference environment according to the target environmental elements to obtain the VR environment, and the VR environment includes the target environmental elements.

[0128] In a possible example, the environment description information includes a target image. The form of the interactive object is determined through the following steps: identifying the target image to determine at least one reference interactive object included in the target image; determining the interaction method of the at least one reference interactive object; determining a scene interactive object from the at least one reference interactive object according to the interaction method and the atomic skill corresponding to the robot, where the atomic skill corresponding to the robot supports the interaction method of the scene interactive object; determining the form of the interactive object according to the scene interactive object.

[0129] In a possible example, the task configuration device 30 of the humanoid robot is further configured to: when the VR environment includes multiple VR sub - environments, after receiving the first pose data and the first video stream data from the VR wearable device, display the display content corresponding to the first video stream on the demonstration video editing interface; in response to a selection operation on a target segment in the display content, send fourth information to the VR wearable device, where the fourth information is used to instruct the VR wearable device to redisplay the target VR sub - environment corresponding to the target segment, and the multiple VR sub - environments include the target VR sub - environment; output a second prompt message, where the second prompt message is used to prompt the user to wear the VR wearable device to perform a task operation redemonstration; receive second pose data and a second video stream from the VR wearable device, where the second pose data is the pose data of the user collected during the task operation redemonstration, and the second video stream is the video stream corresponding to the actual display content of the VR wearable device collected during the operation redemonstration; update the first pose data and the first video stream according to the second pose data and the second video stream to obtain third pose data and a third video stream, where the third pose data includes the second pose data and the third video stream includes the second video stream; send fourth information carrying the task type, the third video stream, and the third pose data to the server.

[0130] In a possible example, in terms of displaying the display content corresponding to the first video stream on the demonstration video editing interface, the task configuration device 30 of the humanoid robot is specifically configured to: receive fifth information from the robot, where the fifth information is used to indicate the subtasks that were not successfully executed when the robot executed the subtask sequence; determine a reference segment corresponding to the subtasks that were not successfully executed from the first video stream; determine the reference segment as the display content; and display the display content on the demonstration video editing interface.

[0131] In a possible example, the subtask sequence further includes the start condition and the end condition of each subtask in the multiple subtasks. The multiple subtasks include a first subtask, and the end condition corresponding to the first subtask is determined according to the target interaction state. The target interaction state is determined by identifying the end image frame of the video stream segment corresponding to the first subtask, and the target interaction state is the interaction state between the user and the target interaction object in the VR environment in the end image frame.

[0132] In a possible example, the third information further carries the timestamp information of the end image frame, and the timestamp information is determined as follows: displaying the display content corresponding to the first video stream through a demonstration video editing interface; in response to a selection operation on a target image frame in the display content, determining the target image frame as the end image frame of the video stream segment; and determining the timestamp information corresponding to the end image frame.

[0133] In the case of adopting an integrated unit, the functional unit composition block diagram of another task configuration device for a humanoid robot provided by an embodiment of the present application is as Figure 11 shown. In Figure 11 , the task configuration device of the humanoid robot includes: a processing module 310 and a communication module 311. The processing module 310 is used to control and manage the actions of the task configuration device of the humanoid robot. For example, it performs the steps executed by the acquisition unit 301, the output unit 302, the first sending unit 303, the second sending unit 304, the receiving unit 305, and the third sending unit 306, and / or is used to execute other processes of the technology described in the present application. The communication module 311 is used to support the interaction between the task configuration device of the humanoid robot and other devices. As Figure 11 shown, the task configuration device of the humanoid robot may further include a storage module 312, and the storage module 312 is used to store the program code and data of the task configuration device of the humanoid robot.

[0134] Among them, the processing module 310 may be a processor or a controller. For example, it may be a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an ASIC, an FPGA, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in combination with the disclosure of the present application. The processor may also be a combination that realizes a computing function, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, and so on. The communication module 311 may be a transceiver, an RF circuit, or a communication interface, etc. The storage module 312 may be a memory.

[0135] Among them, all relevant contents of each scenario involved in the above method embodiment can be cited in the function description of the corresponding functional module, and will not be repeated here. The above task configuration device of the humanoid robot can execute the steps executed by the terminal device in the task configuration method of the humanoid robot shown in the above Figure 3 .

[0136] An embodiment of the present application further provides a computer-readable storage medium, on which executable program code is stored. The executable program code includes execution instructions for executing some or all of the steps of the task configuration method of any humanoid robot described in the foregoing method embodiments. The computer includes a terminal device.

[0137] An embodiment of the present application further provides a computer program product, which includes a computer program that can be operated to cause a computer to execute some or all of the steps of the task configuration method of any humanoid robot described in the foregoing method embodiments. The computer program product can be a software installation package, and the computer includes a terminal device.

[0138] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and units involved are not necessarily essential to the present application.

[0139] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0140] In several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.

[0141] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0142] In addition, in each embodiment of the present application, each functional unit may be integrated into a processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0143] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable memory. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above-mentioned methods in each embodiment of the present application. The aforementioned memory includes: USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical disks and other media that can store program codes.

[0144] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable memory, and the memory may include: flash drives, read-only memories (abbreviation: ROM), random access memories (abbreviation: RAM), magnetic disks, or optical disks, etc.

[0145] The above has introduced the embodiments of the present application in detail. Specific examples are used in the present application to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A task configuration method for a humanoid robot, characterized in that: A terminal device applied to a humanoid robot control system, the humanoid robot control system comprising: the terminal device, a virtual reality VR wearable device, a server and the humanoid robot, the method comprising: In response to an input operation based on the robot function configuration interface, acquiring task configuration information, wherein the task configuration information includes a task type; Outputting first prompt information, where the first prompt information is used to prompt a user to wear the VR wearable device to perform a task operation demonstration; Sending first information to the VR wearable device, where the first information is used to instruct the VR wearable device to display a VR environment corresponding to the task type, where the VR environment is used to simulate an operation scenario of the task operation demonstration; In response to a selection operation on a target control in the robot function configuration interface, second information is sent to the VR wearable device, where the second information is used to instruct the VR wearable device to send the first posture data of the user collected during the task operation demonstration and a first video stream corresponding to the actual display content of the VR wearable device; Receiving the first posture data and the first video stream from the VR wearable device; Sending third information carrying the task type, the first video stream and the first posture data to the server, the third information is used to instruct the server to send the association between the task type and the subtask sequence, and at least one atomic skill model parameter to the robot, the subtask sequence is obtained by task splitting processing based on the first video stream, the subtask sequence includes multiple subtasks and the execution order of the multiple subtasks, the atomic skill model parameter is determined based on the first posture data, and the atomic skill model parameter is used to update at least one atomic skill corresponding to the robot and the multiple subtasks.

2. The method according to claim 1, characterized in that The task configuration information also includes task environment description information, and the VR environment is determined by the following steps: Determine a VR reference environment corresponding to the task type from a plurality of preset VR basic environments; Determine a target environment element according to the task environment description information, wherein the target environment element includes at least one of the following: a building form, an interactive object form; The VR reference environment is adjusted according to the target environment element to obtain the VR environment, and the VR environment includes the target environment element.

3. The method according to claim 2, characterized in that The environment description information includes a target image, and the interactive object form is determined by the following steps: Recognizing the target image to determine at least one reference interactive object included in the target image; determining an interaction mode of the at least one reference interaction object; Determine, according to the interaction mode and the atomic skill corresponding to the robot, a scene interaction object from the at least one reference interaction object, wherein the atomic skill corresponding to the robot supports the interaction mode of the scene interaction object; The interactive object form is determined according to the scene interactive object.

4. The method according to claim 1, characterized in that: The VR environment includes a plurality of VR sub-environments. After receiving the first posture data and the first video stream data from the VR wearable device, the method further includes: Displaying the display content corresponding to the first video stream on the demonstration video editing interface; In response to a selection operation on a target segment in the displayed content, fourth information is sent to the VR wearable device, where the fourth information is used to instruct the VR wearable device to redisplay a target VR sub-environment corresponding to the target segment, where the multiple VR sub-environments include the target VR sub-environment; Outputting second prompt information, where the second prompt information is used to prompt the user to wear the VR wearable device to perform a task operation re-demonstration; Receiving second posture data and a second video stream from the VR wearable device, wherein the second posture data is posture data of the user collected during the task operation re-demonstration process, and the second video stream is a video stream corresponding to actual display content of the VR wearable device collected during the operation re-demonstration process; According to the second posture data and the second video stream, the first posture data and the first video stream are updated to obtain third posture data and a third video stream, wherein the third posture data includes the second posture data, and the third video stream includes the second video stream; Sending fourth information carrying the task type, the third video stream and the third posture data to the server.

5. The method according to claim 4, characterized in that The display content corresponding to the first video stream displayed on the demonstration video editing interface includes: receiving fifth information from the robot, wherein the fifth information is used to indicate a subtask that was not successfully executed when the robot executes the subtask sequence; Determining, from the first video stream, a reference segment corresponding to the subtask that was not successfully executed; determining the reference segment as the display content; The display content is displayed in the demonstration video editing interface.

6. The method according to claim 1, characterized in that The subtask sequence also includes start conditions and end conditions for each of the multiple subtasks, the multiple subtasks include a first subtask, the end condition corresponding to the first subtask is determined according to a target interaction state, the target interaction state is determined by identifying an end image frame corresponding to a video stream segment corresponding to the first subtask, and the target interaction state is an interaction state between the user and the target interaction object in the VR environment in the end image frame.

7. The method according to claim 6, characterized in that The third information also carries the timestamp information of the end image frame, and the timestamp information is determined in the following manner: Displaying the display content corresponding to the first video stream through a demonstration video editing interface; In response to a selection operation on a target image frame in the display content, determining the target image frame as the end image frame of the video stream segment; The timestamp information corresponding to the end image frame is determined.

8. A task configuration device for a humanoid robot, characterized in that: A terminal device applied to a humanoid robot control system, the humanoid robot control system comprising: the terminal device, a virtual reality VR wearable device, a server and the humanoid robot, the device comprising: An acquisition unit, configured to acquire task configuration information in response to an input operation based on the robot function configuration interface, wherein the task configuration information includes a task type; An output unit, configured to output first prompt information, wherein the first prompt information is used to prompt a user to wear the VR wearable device to perform a task operation demonstration; A first sending unit, configured to send first information to the VR wearable device, wherein the first information is used to instruct the VR wearable device to display a VR environment corresponding to the task type, wherein the VR environment is used to simulate an operation scenario of the task operation demonstration; A second sending unit is used to send second information to the VR wearable device in response to a selection operation on a target control in the robot function configuration interface, wherein the second information is used to instruct the VR wearable device to send the first posture data of the user collected during the task operation demonstration and a first video stream corresponding to the actual display content of the VR wearable device; A receiving unit, configured to receive the first posture data and the first video stream from the VR wearable device; A third sending unit is used to send third information carrying the task type, the first video stream and the first posture data to the server, the third information is used to instruct the server to send the association relationship between the task type and the subtask sequence, and at least one atomic skill model parameter to the robot, the subtask sequence is obtained by task splitting processing based on the first video stream, the subtask sequence includes multiple subtasks and the execution order of the multiple subtasks, the atomic skill model parameter is determined based on the first posture data, and the atomic skill model parameter is used to update at least one atomic skill corresponding to the robot and the multiple subtasks.

9. A terminal device, characterized in that: The terminal device comprises: a processor, a memory and a communication interface, wherein the processor, the memory and the communication interface are interconnected, wherein the communication interface is used to receive or send data, the memory is used to store application code for the terminal device to execute the method as described in any one of claims 1-7, and the processor is configured to execute the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program for electronic data exchange, wherein the computer program enables a computer to execute the steps in the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Remote control system, information processing method, and program

    CN112154047A

  • Method for remotely operating humanoid robot by identifying hand postures through virtual reality glasses

    CN117746494A