Task configuration method of humanoid robot and related device

By introducing task configuration methods for terminal devices, VR devices and servers into the humanoid robot control system, the problem of insufficient flexibility in task execution of humanoid robots is solved, and more abundant and flexible task execution capabilities are achieved.

CN119910668AActive Publication Date: 2025-05-02SHANGHAI FOURIER INTELLIGENCE CO LTD

Patent Information

Application Number
CN202510412752.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-05-02
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

When existing humanoid robots imitate human behavior to perform tasks, they usually can only perform pre-set fixed types of tasks, and cannot flexibly configure tasks according to the actual usage needs of users, resulting in insufficient flexibility in task execution.

Method used

The task configuration method is realized by introducing terminal devices, virtual reality VR wearable devices and servers into the humanoid robot control system. The method includes obtaining task configuration information, prompting the user to wear VR equipment for task operation demonstration, collecting user posture data and video streams, and sending them to the server, generating sub-task sequences and atomic skill model parameters, and updating the robot's atomic skills.

Benefits of technology

The diversity and flexibility of humanoid robot task execution is improved, allowing the robot to customize task types according to the actual needs of the user and more accurately match the user's operational needs during task execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119910668A_ABST
    Figure CN119910668A_ABST
Patent Text Reader

Abstract

The invention discloses a task configuration method of a humanoid robot and a related device, which are applied to humanoid robot control system terminal equipment, and the method comprises the following steps: responding to an input operation based on a robot function configuration interface, and obtaining task configuration information; outputting first prompt information, and prompting a user to wear VR wearable equipment to perform task operation demonstration; first information is sent to the VR wearable device, and the VR wearable device is indicated to display a VR environment corresponding to the task type; in response to a selection operation for a target control of the robot function configuration interface, sending second information to the VR wearable device; receiving first posture data and a first video stream of the VR wearable device; and third information carrying the task type, the first video stream and the first attitude data is sent to the server, and the server is indicated to send an association relationship between the task type and the subtask sequence and at least one atomic skill model parameter to the robot. The task execution diversity and flexibility of the humanoid robot can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of humanoid robot control, and in particular to a task configuration method and related devices for a humanoid robot. Background Art

[0002] With the continuous development of robotics and intelligent technology, robots have been widely used in various fields such as service industry, education, medical care, entertainment and research. Among them, humanoid robots have become an important branch in the development of robots because of their similar appearance features to humans and their ability to imitate human behavior to perform tasks.

[0003] However, at present, when humanoid robots imitate human behavior to perform tasks, they can usually only perform pre-set fixed types of tasks, and cannot flexibly configure tasks according to the actual needs of users, resulting in the problem of insufficient flexibility in the task execution of humanoid robots. Summary of the invention

[0004] The embodiments of the present application provide a task configuration method and related devices for a humanoid robot, in order to improve the diversity and flexibility of the tasks performed by the humanoid robot.

[0005] In a first aspect, an embodiment of the present application provides a task configuration method for a humanoid robot, which is applied to a terminal device in a humanoid robot control system, wherein the humanoid robot control system includes: the terminal device, a virtual reality VR wearable device, a server, and the humanoid robot, and the method includes: In response to an input operation based on the robot function configuration interface, acquiring task configuration information, wherein the task configuration information includes a task type; Outputting first prompt information, where the first prompt information is used to prompt a user to wear the VR wearable device to perform a task operation demonstration; Sending first information to the VR wearable device, where the first information is used to instruct the VR wearable device to display a VR environment corresponding to the task type, where the VR environment is used to simulate an operation scenario of the task operation demonstration; In response to a selection operation on a target control in the robot function configuration interface, second information is sent to the VR wearable device, where the second information is used to instruct the VR wearable device to send the first posture data of the user collected during the task operation demonstration and a first video stream corresponding to the actual display content of the VR wearable device; Receiving the first posture data and the first video stream from the VR wearable device; Sending third information carrying the task type, the first video stream and the first posture data to the server, the third information is used to instruct the server to send the association between the task type and the subtask sequence, and at least one atomic skill model parameter to the robot, the subtask sequence is obtained by task splitting processing based on the first video stream, the subtask sequence includes multiple subtasks and the execution order of the multiple subtasks, the atomic skill model parameter is determined based on the first posture data, and the atomic skill model parameter is used to update at least one atomic skill corresponding to the robot and the multiple subtasks.

[0006] In a second aspect, an embodiment of the present application provides a task configuration device for a humanoid robot, which is applied to a terminal device in a humanoid robot control system. The humanoid robot control system includes: the terminal device, a virtual reality VR wearable device, a server and the humanoid robot. The device includes: An acquisition unit, configured to acquire task configuration information in response to an input operation based on the robot function configuration interface, wherein the task configuration information includes a task type; An output unit, configured to output first prompt information, wherein the first prompt information is used to prompt a user to wear the VR wearable device to perform a task operation demonstration; A first sending unit, configured to send first information to the VR wearable device, wherein the first information is used to instruct the VR wearable device to display a VR environment corresponding to the task type, wherein the VR environment is used to simulate an operation scenario of the task operation demonstration; A second sending unit is used to send second information to the VR wearable device in response to a selection operation on a target control in the robot function configuration interface, wherein the second information is used to instruct the VR wearable device to send the first posture data of the user collected during the task operation demonstration and a first video stream corresponding to the actual display content of the VR wearable device; A receiving unit, configured to receive the first posture data and the first video stream from the VR wearable device; A third sending unit is used to send third information carrying the task type, the first video stream and the first posture data to the server, the third information is used to instruct the server to send the association relationship between the task type and the subtask sequence, and at least one atomic skill model parameter to the robot, the subtask sequence is obtained by task splitting processing based on the first video stream, the subtask sequence includes multiple subtasks and the execution order of the multiple subtasks, the atomic skill model parameter is determined based on the first posture data, and the atomic skill model parameter is used to update at least one atomic skill corresponding to the robot and the multiple subtasks.

[0007] In a third aspect, an embodiment of the present application provides a terminal device, wherein the terminal device includes: A processor, a memory and a communication interface, wherein the processor, the memory and the communication interface are interconnected, wherein the communication interface is used to receive or send data, the memory is used to store application code for a terminal device to execute the method of the first aspect, and the processor is configured to execute the method of the first aspect.

[0008] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program for electronic data exchange, wherein the computer program enables a computer to execute the steps of the task configuration method for a humanoid robot as described in the first aspect.

[0009] In a fifth aspect, an embodiment of the present application provides a computer program product, wherein the computer program product includes a computer program, and the computer program is operable to cause a computer to execute some or all of the steps described in the task configuration method for a humanoid robot as described in the first aspect of the embodiment of the present application. The computer program product can be a software installation package.

[0010] It can be seen that in the embodiment of the present application, the terminal can obtain task configuration information based on the interface input operation, and output prompt information to prompt the user to wear the VR wearable device to perform the task operation demonstration, and send information to the VR wearable device to instruct it to display the VR environment corresponding to the task type of the task configuration information to simulate the operation scenario of the task operation demonstration. Then the terminal can also respond to the selection operation for the target control, send information to the VR wearable device, instruct it to send the user posture data collected during the task operation demonstration and the video stream corresponding to the actual display content of the VR wearable device, and send the posture data, video stream and acquired task type received from the VR wearable device to the server, and the server sends the association relationship between the task type and the subtask sequence and at least one atomic skill model parameter to the humanoid robot. The subtask sequence is obtained by task splitting processing according to the video stream, including multiple subtasks and the execution order of multiple subtasks. The atomic skill model parameter is determined according to the posture data and is used to update the robot at least one atomic skill corresponding to the multiple subtasks. In this way, the terminal can receive the task type of the humanoid robot customized by the user, and can generate a sub-task sequence of the humanoid robot in the task type by obtaining the posture data obtained by the user in the actual operation demonstration of the customized task type, and the video stream of the actual display screen of the VR wearable device, and update the robot atomic skills corresponding to the sub-tasks in the sub-task sequence based on the user's actual posture data, so that the robot can perform more tasks. The execution of tasks can better match the actual needs of users, which is conducive to improving the diversity and flexibility of the task execution of humanoid robots. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0012] Figure 1 It is a schematic diagram of the architecture of a task configuration system for a humanoid robot provided in an embodiment of the present application; Figure 2 It is a structural diagram of a terminal device provided in an embodiment of the present application; Figure 3 It is a flowchart of a method for configuring a task of a humanoid robot provided in an embodiment of the present application; Figure 4 is a schematic diagram of a robot function configuration interface provided in an embodiment of the present application; Figure 5 is a schematic diagram of a robot function selection interface provided in an embodiment of the present application; Figure 6 is a schematic diagram of a prompt information display interface provided in an embodiment of the present application; Figure 7 is a schematic diagram of another prompt information display interface provided in an embodiment of the present application; Figure 8 is a schematic diagram of a demonstration video editing interface provided in an embodiment of the present application; Fig. 9 is a schematic diagram of another demonstration video editing interface provided in an embodiment of the present application; Fig.10 It is a block diagram of the functional units of a task configuration device for a humanoid robot provided in an embodiment of the present application; Fig.11 This is a block diagram of the functional units of another task configuration device for a humanoid robot provided in an embodiment of the present application. DETAILED DESCRIPTION

[0013] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0014] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices.

[0015] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0016] The embodiment of the present application provides a task configuration method and related devices for a humanoid robot. The embodiment of the present application is described in detail below in conjunction with the accompanying drawings.

[0017] Please refer to Figure 1 , Figure 1 1 is a schematic diagram of the architecture of a task configuration system for a humanoid robot provided in an embodiment of the present application. The task configuration method and the task configuration device of the humanoid robot in the embodiment of the present application can be applied to the task configuration system of the humanoid robot. The task configuration system 10 of the humanoid robot may include a terminal device 101, a server 102, a virtual reality VR wearable device 103 and a humanoid robot 104. Any two devices in the task configuration system 10 of the humanoid robot may be directly or indirectly connected by wired or wireless communication, which is not limited in the embodiment of the present application.

[0018] In the embodiment of the present application, the terminal device 101 is mainly responsible for acquiring the interaction with the user and the VR wearable device, and then acquiring the task configuration information and the user posture data during the task operation demonstration and the video stream corresponding to the actual display content of the VR wearable device, and sending it to the server for processing. The user can interact with the system through the terminal device, such as inputting and selecting operations in the terminal device display interface, thereby triggering the terminal device to execute the steps of acquiring task configuration information and sending information to the VR wearable device. The VR wearable device is mainly responsible for displaying the VR environment used to simulate the task operation demonstration scene, and collecting the posture data when the user performs the task operation demonstration based on the VR environment, and collecting the video stream corresponding to the actual display content of the VR wearable device during the task operation demonstration, and sending it to the terminal device. The server is mainly responsible for processing the video stream and posture data from the terminal device, such as performing task splitting processing on the video stream, generating the corresponding subtask sequence, and determining the atomic skill model parameters according to the posture data, which is used to update the atomic skills corresponding to the subtasks in the humanoid robot and the subtask sequence.

[0019] In a specific implementation, after receiving the association between the task type and the subtask sequence and at least one atomic skill model parameter from the server 102, the humanoid robot 104 can update the atomic skill corresponding to the subtask associated with the subtask sequence according to the atomic skill model parameter, and can save the association between the subtask sequence and the task type. The user can subsequently select the configured task type through the display interface of the terminal device 101. After the terminal device 101 confirms the task type to be executed based on the user's operation, it can send a task execution request carrying the task type to the humanoid robot 104. After receiving the task execution request, the humanoid robot 104 can execute the user-defined task type.

[0020] Based on this system, the terminal can receive the task type of the humanoid robot customized by the user, and can generate a sub-task sequence of the humanoid robot in the task type by obtaining the posture data obtained by the user in the actual operation demonstration of the customized task type and the video stream of the actual display screen of the VR wearable device, and update the robot atomic skills corresponding to the sub-tasks of the sub-task sequence based on the user's actual posture data, so that the robot can perform more tasks. The execution of tasks can better match the actual needs of users, which is conducive to improving the diversity and flexibility of robot task execution.

[0021] Of course, in other embodiments, the terminal device 101 and the server 102 may also be the same device, that is, a single computer can support receiving user input operations and selection operations to obtain task configuration information and trigger interaction with VR wearable devices, as well as support processing of posture data and video stream data collected by VR wearable devices to generate subtask sequences and atomic skill model parameters, and send them to the robot.

[0022] Understandably, Figure 1 The forms and quantities of the terminal device 101, server 102, VR wearable device 103 and humanoid robot 104 shown in the figure are for example only and do not constitute a limitation on the implementation manner of the present application.

[0023] In the present application, the terminal can be various types of terminals such as laptops, tablet computers, desktop computers, set-top boxes, smart phones, smart speakers, smart watches, smart TVs, and vehicle-mounted terminals. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. VR wearable devices may include accessories such as VR head-mounted displays, VR handles and controllers, VR gloves, VR full-body tracking devices, and headphones to support the VR wearable device to display the VR environment, and to support the VR wearable device to collect the user's posture data (such as the posture data of the user's hands, head, wrists, body and other joints or parts). A humanoid robot may be a robot with a body structure similar to that of a human being, capable of performing functional operations similar to human behaviors such as walking, grasping, and turning the head, and may also be equipped with a variety of sensor systems such as cameras, microphones, and touch sensors for sensing the environment.

[0024] Please refer to Figure 2 The terminal device in this application may be structured as follows: Figure 2As shown, the terminal device may include a processor 110, a memory 120, a communication interface 130, and one or more programs 121, wherein the one or more programs 121 are stored in the above-mentioned memory 120 and are configured to be executed by the above-mentioned processor 110, and the one or more programs 121 include instructions for executing any step performed by the terminal device in the following method embodiment. Among them, the communication interface 130 is used to support the communication between the terminal device and other devices. In a specific implementation, the processor 110 is used to execute any step performed by the terminal device in the following method embodiment, and when performing data transmission such as sending, the communication interface 130 can be selectively called to complete the corresponding operation. It should be noted that the structural diagram of the above-mentioned terminal device is only an example, and the specific devices included may be more or less, and are not uniquely limited here.

[0025] See also Figure 3 , Figure 3 is a flow chart of a task configuration method for a humanoid robot provided in an embodiment of the present application. The task configuration method for a humanoid robot can be applied to Figure 1 The humanoid robot task configuration system shown in the figure is a terminal device, and the humanoid robot control system includes: a terminal device, a virtual reality VR wearable device, a server and a humanoid robot, such as Figure 3 As shown, the task configuration method of the humanoid robot includes the following steps: S201, the terminal device obtains task configuration information in response to an input operation based on the robot function configuration interface.

[0026] The task configuration information includes the task type.

[0027] In a specific implementation, a humanoid robot control application may be run on the terminal device, and the user may control the robot and configure the robot's functions through the humanoid robot control application.

[0028] For example, the robot function configuration interface can be as follows Figure 4 As shown, an input box corresponding to the task type may be displayed in the robot function configuration interface, and the user may complete the setting of the task type by entering the task name in the input box.

[0029] In a specific implementation, the display of the robot function configuration interface can also be triggered based on the detection of operations on controls of other interfaces. For example, before step S201, the electronic device can also display a robot function selection interface, which can display the robot's system preset functions and user-defined functions. The "New Task" control can also be displayed on the user-defined function interface. The terminal device can also jump to the robot function configuration interface in S201 when a preset operation (such as a click operation) on the "New Task" control is detected.

[0030] In addition, Figure 5 Taking the robot function selection interface shown in the figure as an example, the top of the function selection interface can also display the "system preset function" control and the "custom function" control. If the terminal device detects that the user clicks the "system preset function" control, it can display the system preset humanoid robot function list. If the terminal device detects that the user clicks the "custom function" control, it can switch to display the list of task types created by the user and the "new task" control. Specifically, Figure 5 Taking the case where it is detected that a user clicks on the "custom function" control as an example, the function selection page can display a list of task types that the user has created, such as Figure 5 The function list includes multiple created task types such as "go downstairs to pick up a parcel", "clean the living room", and "clean the kitchen". When the terminal device detects a preset operation (such as a click operation) for any created task type, it can send a task execution request carrying the selected created task type information to the corresponding humanoid robot, and the selected created task type information can be executed by the humanoid robot.

[0031] It can be understood that the interface diagrams provided in the embodiments of the present application are all illustrative descriptions. In actual applications, the positions, quantities and forms of expression of controls in various functional interfaces (such as the aforementioned robot function configuration interface and robot function selection interface) can be set differently from the interface diagrams provided in the embodiments of the present application. For example, the "New Task" control in the robot function selection interface can also be set in the upper right corner of the interface, or at any other position, or the "New Task" control can also be directly displayed as a "+" control. No specific restrictions are made here.

[0032] Furthermore, task configuration prompt information can also be displayed in the robot function configuration interface. The task configuration prompt information can be generated based on the atomic skills of the humanoid robot. Specifically, the terminal device can send the identification information of the humanoid robot, such as the model, to the server, and the server queries the task type of the created function corresponding to the identification information of the humanoid robot uploaded by the user as a reference task type, and sends the obtained reference task type to the terminal device. The terminal device can then generate task configuration prompt information based on the reference task type.

[0033] In particular, the terminal device can obtain a task configuration prompt information template, use the reference task type as a keyword, fill it into the information template, and finally generate the task configuration prompt information. Figure 4 Taking the robot function configuration page shown as an example, in other embodiments, in the input box corresponding to the task type, in addition to displaying the prompt information of "Please enter the task name", it can also display task configuration prompt information such as "Try to create a task and let the robot help you collect the express delivery".

[0034] In particular, please continue to refer to Figure 4 When obtaining task configuration information, the terminal device can also display a robot list for the user. The robot list includes robots bound to the user's login account or the terminal device. The terminal device can establish an association between the robot and the user-created task type based on the user's selection operation on the robot in the robot list. In scenarios such as sending robot identification information to the server, the terminal device can determine the robot identification information that needs to be sent from multiple robot identification information based on the association relationship.

[0035] Correspondingly, in other embodiments, when the terminal device displays the created task list and the system preset task list, it can also generate a task sub-list corresponding to each robot based on the association between the robot and each task, and display the tasks corresponding to different robots in partitions to facilitate user viewing and operation.

[0036] S202, the terminal device outputs the first prompt information.

[0037] Among them, the first prompt information is used to prompt the user to wear the VR wearable device to perform a task operation demonstration.

[0038] The first prompt information may be output through a display interface, or the first prompt information may be other information such as audio information or vibration information.

[0039] For example, taking the first prompt information output to the user through the display interface as an example, in S202, the terminal device may display the following Figure 6The interface shown implements the output of the first prompt information, and prompts the user to wear the VR wearable device to perform the task operation demonstration by displaying a text prompt information such as "VR device is connected, please wear the VR device to start the task operation demonstration". Alternatively, in other embodiments, the terminal device may also prompt by outputting a voice such as "Please wear the VR device to start the task operation demonstration". It is understandable that the specific text content and voice content of the prompt information here are only exemplary, and other forms of prompt information can also be selected in actual applications. Regardless of the form of prompt information used, the specific content of the prompt information can be set as needed, and no specific restrictions are made here.

[0040] In a specific implementation, before step S202, the terminal device may first detect whether a connection has been established with the VR wearable device. If it is detected that no connection has been established with the VR wearable device, the terminal device may jump to display the VR wearable device connection setting interface to prompt the user to perform the VR wearable device connection setting. After the terminal device detects that the user operation has established a connection with the VR wearable device, the first prompt information may be output.

[0041] In a specific implementation, after step S201 and before step S202, the terminal device may also display, for example, Figure 7 The interface shown in the figure prompts the user whether to start the task operation demonstration. Specifically, the terminal device can display the confirmation operation on the task configuration information when it detects the confirmation operation on the task configuration information. Figure 7 The interface shown is still based on Figure 4 For example, the robot function configuration interface of , the confirmation operation can be a selection operation for the "Confirm" control. Figure 7 As shown, the terminal device may display information such as "Task creation completed, do you want to start the task operation demonstration immediately?" If a user selection operation for the "Start Now" control is detected, step S202 may be executed, such as detecting whether the VR wearable device is connected, and jumping to display when connected. Figure 6 If the terminal device detects the user's selection operation on the "wait" control, there is no need to continue to perform steps S202 to S206, and the terminal device can jump to display the following Figure 5 The robot function selection interface shown in the figure may also display a list of tasks to be demonstrated, which includes task types in which the user has only completed task creation in the robot function configuration interface but has not completed task demonstration operations, such as Figure 5 When the terminal device detects a selection operation for a task type in the task list to be demonstrated again, steps S201 to S206 may be executed again.

[0042] S203: The terminal device sends first information to the VR wearable device.

[0043] Among them, the first information is used to instruct the VR wearable device to display a VR environment corresponding to the task type, and the VR environment is used to simulate the operation scenario of the task operation demonstration.

[0044] In a specific implementation, the VR environment, for example, extracts keywords from the task type in the task configuration information input by the user, and sends the keywords to the server, which determines the matching VR environment based on the keywords and feeds back the environmental data of the VR environment or the environmental identification of the VR environment, and then the electronic device carries the environmental data or environmental identification in the first information and sends it to the VR wearable device for display. Alternatively, the VR wearable device can save the VR general environment, and when the environmental data or environmental identification from the electronic device is not received, the VR general environment can be directly displayed as the VR environment corresponding to the task type.

[0045] S204: The terminal device sends second information to the VR wearable device in response to a selection operation on a target control in the robot function configuration interface.

[0046] Among them, the second information is used to instruct the VR wearable device to send the first posture data of the user collected during the task operation demonstration and the first video stream corresponding to the actual display content of the VR wearable device.

[0047] In a specific implementation, the user's posture data may specifically be posture data of the user's hands, head, wrists, body and other joints or parts collected by the VR wearable device. The video stream corresponding to the actual display content of the VR wearable device, that is, when the user wears the VR wearable device to perform a task operation demonstration, the video stream corresponding to the display screen of the VR wearable device when the VR wearable device interacts with the VR environment through the VR wearable device.

[0048] In a specific implementation, the selection operation may be, for example, a click operation, or a long press operation for a preset time, etc. The target control may be, for example, Figure 6 The “Complete Demonstration” control shown in , after the terminal device detects that the “Complete Demonstration” control is selected, it can send a second message to the VR wearable device, and the second message can also instruct the VR wearable device to stop collecting posture data and video streams.

[0049] also, Figure 6 The "Exit Demonstration" control can also be displayed in the terminal device. After the terminal device detects the "Exit Demonstration" control, it can also output a save prompt message to prompt the user to choose whether to save the previously collected user posture information and video stream. If a selection operation for the save control is detected, the second information can also be sent to the VR wearable device, and the received user posture information and video stream can be saved, and the display can also be jumped. Figure 5The robot function selection interface is opened, and the task type information in the task list to be demonstrated in the robot function selection interface is updated. When a selection operation for the task type in the task list to be demonstrated is subsequently detected and steps S201 to S206 are restarted, the screen that the VR wearable device currently needs to display is determined based on the actual display image of the VR wearable device recorded when the demonstration was previously exited, and the interaction status of each interactive object in the VR environment is confirmed, so that the user can continue the operation demonstration through the VR wearable device based on the display screen when the demonstration was last exited, and the interactive objects in each interaction status when the operation demonstration is started are also consistent with those when the demonstration was last exited, thereby improving the efficiency of task configuration and saving system resources.

[0050] In a specific implementation, the VR wearable device can stop collecting the posture data and the video stream when receiving the second information, and start displaying the VR environment when detecting that a connection is established with the terminal device or receiving a start request from the terminal device. The start request can be, for example, sent by the terminal device when detecting that the user is directed to Figure 7 Emitted when a "Start Now" control in is selected.

[0051] S205, the terminal device receives the first posture data and the first video stream from the VR wearable device.

[0052] S206, the terminal device sends third information carrying the task type, the first video stream and the first posture data to the server.

[0053] Among them, the third information is used to instruct the server to send the association relationship between the task type and the subtask sequence, and at least one atomic skill model parameter to the robot. The subtask sequence is obtained by task splitting processing based on the first video stream. The subtask sequence includes multiple subtasks and the execution order of the multiple subtasks. The atomic skill model parameter is determined based on the first posture data. The atomic skill model parameter is used to update at least one atomic skill corresponding to the robot and the multiple subtasks.

[0054] In a specific implementation, the server can split the first video stream into subtask sequences, each subtask sequence corresponds to a skill group, each skill group corresponds to an atomic skill set that needs to be mapped to a humanoid robot, and the atomic skill set includes at least one atomic skill corresponding to the multiple subtasks.

[0055] Taking the task type of "taking out the garbage" as an example, the multiple subtasks included in the subtask sequence may specifically include: opening the entrance door, going out, grabbing the garbage bag, closing the door, walking from the entrance door to the elevator entrance, pressing the elevator down button, waiting for the elevator door to open and entering the elevator, pressing the 1st floor button, waiting for the elevator door to open and the elevator display screen to display the 1st floor, exiting the elevator door, walking to the building gate, clicking the door opening button and pushing the door to exit, walking to the garbage placement point, putting the garbage bag into the garbage entrance, walking to the building gate, waiting for the door to open, entering the building, walking to the elevator entrance, detecting and pressing the up button, waiting for the elevator door to open and enter, pressing the target floor button such as the 6th floor, waiting for the elevator door to open and the display screen to display the 6th floor, exiting, walking to the entrance door, waiting for the door to open, and entering. The execution sequence relationship between adjacent subtasks is constructed through the subtask sequence and synchronized to the humanoid robot.

[0056] Among them, the atomic skills corresponding to the above-mentioned subtasks, that is, the robot atomic skills needed to be used in each subtask can be pre-trained. The server can create a fine-tuning data set based on the first posture data, and train the atomic skills needed to be used in the current task type of the humanoid robot according to the fine-tuning data set, that is, at least one atomic skill corresponding to multiple subtasks, so that each atomic skill can be more adaptable to the actual needs of the user. After the server trains the humanoid robot and at least one atomic skill corresponding to the multiple subtasks based on the fine-tuning data set, the new model parameters of each atomic skill after training, that is, at least one atomic skill model parameter in the embodiment of the present application, can be obtained, and sent to the humanoid robot, so that the humanoid robot can subsequently complete the user-defined task type.

[0057] It can be seen that in the embodiment of the present application, the terminal can obtain task configuration information based on the interface input operation, and output prompt information to prompt the user to wear the VR wearable device to perform the task operation demonstration, and send information to the VR wearable device to instruct it to display the VR environment corresponding to the task type of the task configuration information to simulate the operation scenario of the task operation demonstration. Then the terminal can also respond to the selection operation for the target control, send information to the VR wearable device, instruct it to send the user posture data collected during the task operation demonstration and the video stream corresponding to the actual display content of the VR wearable device, and send the posture data, video stream and acquired task type received from the VR wearable device to the server, and the server sends the association relationship between the task type and the subtask sequence and at least one atomic skill model parameter to the humanoid robot. The subtask sequence is obtained by task splitting processing according to the video stream, including multiple subtasks and the execution order of multiple subtasks. The atomic skill model parameter is determined according to the posture data and is used to update the robot at least one atomic skill corresponding to the multiple subtasks. In this way, the terminal can receive the task type of the humanoid robot customized by the user, and can generate a sub-task sequence of the humanoid robot in the task type by obtaining the posture data obtained by the user in the actual operation demonstration of the customized task type, and the video stream of the actual display screen of the VR wearable device, and update the robot atomic skills corresponding to the sub-tasks in the sub-task sequence based on the user's actual posture data, so that the robot can perform more tasks. The execution of tasks can better match the actual needs of users, which is conducive to improving the diversity and flexibility of the task execution of humanoid robots.

[0058] In a possible example, the task configuration information also includes task environment description information, and the VR environment is determined by the following steps: determining a VR reference environment corresponding to the task type from a plurality of preset VR basic environments; determining target environment elements according to the task environment description information, and the target environment elements include at least one of the following: building form, interactive object form; adjusting the VR reference environment according to the target environment elements to obtain the VR environment, and the VR environment includes the target environment elements.

[0059] The architectural form may include, for example, the appearance of the building, and the interactive object may specifically refer to an object that can interact with the user or can be operated by the user.

[0060] In a specific implementation, the terminal device may send the task type input by the user to the server, and the server determines the VR reference environment corresponding to the task type from multiple VR preset environments; or, the terminal device directly determines the VR reference environment corresponding to the task type based on the task type input by the user. Specifically, each VR preset environment is provided with basic environmental information, and the environmental information is used to characterize the VR preset environment, such as building type information such as residential, commercial buildings, and schools, appearance information such as single-story and multi-story, and functional information such as elevators, stairs, and underground parking lots. The terminal device or server can extract keywords based on the task type, and then determine the VR reference environment that matches the task. For example, when the task type is "go downstairs to pick up a parcel", the VR preset environment of a single-story building should be eliminated.

[0061] Furthermore, in other embodiments, the task environment description information may be used not only to determine the target environment elements, but also to determine the VR reference environment together with the task type.

[0062] In a specific implementation, the VR environment is obtained by adjusting the VR reference environment according to the target environment elements. Specifically, when the target environment elements include the building form, the original building form of the VR reference environment is modified according to the building form. For example, if the original building of the VR reference environment is a 10-story building, and the building form information indicates that the VR reference environment is a 40-story building, the building in the VR reference environment can be modified to a 10-story building. At this time, if the original building of the reference environment includes interactive objects such as elevators whose interactive content is affected by the building form, the interactive content of the interactive objects needs to be modified according to the building form information, such as adding elevator floor buttons. When the target environment elements include the interactive object form, the interactive object form in the VR reference environment can also be modified according to the interactive object form. For example, if the interactive object form included in the target environment elements includes a pedal-type trash can, if the VR reference environment does not include an interactive object such as a trash can, a pedal-type trash can can be directly added, and if the VR reference environment includes a flip-top trash can, it can be modified to a pedal-type trash can.

[0063] It can be seen that in this example, the VR environment can be specifically obtained by determining the target environmental elements according to the task environment description information in the task configuration information, and then adjusting the VR reference environment determined according to the task type in the task configuration information according to the target environmental elements. The target environmental elements specifically include at least one of a building form and an interactive object form. The VR environment after adjusting the VR reference environment through the target environmental elements is more in line with the task execution environment expected by the user. Task operation demonstration based on the VR environment and generation of the robot's sub-task sequence and atomic skill model parameters can better adapt to the actual needs of users, which is conducive to further improving the diversity and flexibility of the robot's task execution.

[0064] In one possible example, the environment description information includes a target image, and the form of the interactive object is determined by the following steps: identifying the target image to determine at least one reference interactive object included in the target image; determining an interaction mode of the at least one reference interactive object; determining a scene interaction object from the at least one reference interactive object based on the interaction mode and the atomic skill corresponding to the robot, the atomic skill corresponding to the robot supporting the interaction mode of the scene interaction object; and determining the form of the interactive object based on the scene interaction object.

[0065] In a specific implementation, when the terminal device recognizes the target image, it may also first perform object recognition on the target image to determine the object included in the target image, and further determine the object type of the target object, thereby determining whether the object is a building or an interactive object. When it is determined that the object is a building, the building shape can be obtained based on the building. If the object is an interactive object, it can be determined as a reference interactive object, and the scene interactive objects can be further screened out and the interactive object shape can be determined.

[0066] Among them, the atomic skill corresponding to the robot supports the interaction mode of the scene interaction object, that is, the robot can interact with the scene interaction object by using the atomic skill. For example, the robot atomic skill does not support the use of inserting a key into the keyhole and rotating to unlock, but only supports swiping a card to open the door and pressing the doorbell. If the reference interaction object of the target image includes a door that only supports key unlocking and has no doorbell, the door cannot be determined as a scene interaction object. If the reference interaction object of the target image includes a door with a doorbell or a smart lock that supports swiping a card to open the door, the door can be determined as a scene interaction. Alternatively, in other embodiments, for the interaction object in the reference interaction object that is not supported by the atomic skill corresponding to the robot, the terminal device can also modify the interaction mode of the reference interaction object according to the atomic skill corresponding to the robot and the interaction object, modify it to the interaction mode supported by the robot atomic skill, and further determine the interactive object form, such as modifying a traditional mechanical lock door to a smart lock door and adding it to the VR reference environment to obtain the final VR environment.

[0067] In addition, in a specific implementation, the task environment description information can be obtained through the target image, and can also be determined by detecting the user's selection operation for the preset task environment, or receiving text information input by the user. It can be understood that the task environment description information can be any one or more combinations of the above-mentioned target image, text information, and the preset task environment selected by the user.

[0068] For example, Figure 4Taking the robot function configuration interface shown in the figure as an example, the robot function configuration interface may also include a setting control for the task environment. Specifically, the user may expand a preset task environment list through the drop-down control of the environment task bar, for example Figure 4 Through the drop-down control, the user can select environment description information from "staircase residence", "elevator residence", and "single-story restaurant", or the user can also enter text or images in the text input box or image upload box under the task environment list.

[0069] It can be seen that in this example, the environment description information includes a target image. The terminal device can identify the target image to determine at least one reference interaction object included in the target image, and determine the interaction mode of at least one reference interaction object, and then determine the scene interaction object whose interaction mode is supported by the robot's atomic skills from at least one reference interaction object according to the interaction mode and the robot's corresponding atomic skills, and then determine the form of the interactive object according to the scene interaction object, which is conducive to improving the matching degree between the VR environment and the robot, and then improving the accuracy of the determined robot subtask sequence, and improving the reliability of the robot's task execution.

[0070] In a possible example, the VR environment includes multiple VR sub-environments, and after receiving the first posture data and the first video stream data from the VR wearable device, the method further includes: displaying the display content corresponding to the first video stream on a demonstration video editing interface; in response to a selection operation on a target segment in the display content, sending fourth information to the VR wearable device, the fourth information is used to instruct the target VR sub-environment corresponding to the target segment to be redisplayed through the VR wearable device, and the multiple VR sub-environments include the target VR sub-environment; outputting second prompt information, the second prompt information is used to prompt the user to wear the VR wearable device to perform task operation re-demonstration; receiving second posture data and a second video stream from the VR wearable device, the second posture data is the posture data of the user collected during the task operation re-demonstration process, and the second video stream is the video stream corresponding to the actual display content of the VR wearable device collected during the operation re-demonstration process; updating the first posture data and the first video stream according to the second posture data and the second video stream to obtain third posture data and a third video stream, the third posture data includes the second posture data, and the third video stream includes the second video stream; sending fourth information carrying the task type, the third video stream and the third posture data to the server.

[0071] In a specific implementation, the first posture data and the first video stream are updated according to the second posture data and the second video stream, and the specific method for obtaining the third posture data and the third video stream can be: according to the start time and end time corresponding to the target segment, the posture sub-data corresponding to the target segment is determined from the first posture data, and according to the start time and end time corresponding to the target segment, the video stream sub-data corresponding to the target segment is determined from the first video stream; the posture sub-data is replaced by the second posture data to obtain the third posture data, and the video stream sub-data is replaced by the second video stream to obtain the second video stream.

[0072] That is to say, the user can operate the terminal device to re-present the first video stream segment, and replace the first video stream segment with the re-presented second posture data and the second video stream. Among them, the target VR sub-environment is the part of the VR environment actually displayed by the VR wearable device at the target segment start time. For example, the VR environment includes a complete residential building, and the target segment start time selected by the user is the elevator door on the 10th floor of the user in the complete residential building. Then the terminal device can determine that the target VR sub-environment is the elevator door on the 10th floor, and instruct the VR wearable device to redisplay the elevator door on the 10th floor through the fourth information.

[0073] In a specific implementation, after receiving the fourth information, the server can use the third video stream as the first video stream and the third posture data as the first posture data, and then execute step S206 to determine the correspondence between the task subsequence and the task type according to the task type, the third video stream and the third posture data, and determine at least one atomic skill model parameter.

[0074] In the specific implementation, Figure 8 and Fig. 9 Take the demo video editing interface shown in the figure as an example. When the terminal detects that the user is targeting Figure 6 or Figure 7 After a click operation on the shaded portion of the white progress bar, if a click operation on the "replace control" is detected, it can be confirmed that the selection of the target segment is detected, and the shaded portion of the segment clicked by the user is determined as the target segment, and then the fourth information is sent to the VR wearable device. In particular, after detecting a click operation on a reference segment, the terminal device can also change the display color of the clicked reference segment to distinguish the selected reference segment from other reference segments, and prompt the user the location of the selected reference segment, for example Figure 8 As shown in , when the terminal device detects that the user clicks on the reference segment on the left side of the progress bar, the color of the reference segment is changed to a color different from the unselected reference segment on the right side for display.

[0075] In addition, the terminal detects that the user is targeting Figure 8 or Fig. 9 After a click operation at any time point in the white progress bar, if a click operation for "Insert Control" is detected, the VR sub-environment corresponding to the time point can also be determined, and the sixth information is sent to the VR wearable device, instructing the VR wearable device to display the VR sub-environment corresponding to the time point and continue to collect the user's fourth posture data and the fourth video stream corresponding to the content displayed by the VR wearable device, until the user's click operation for the stop control is detected, the VR wearable device is notified to stop collecting, and the fourth posture data and the fourth video stream collected this time are sent to the terminal device. At this time, the terminal can also insert the fourth posture data into the first posture data according to the time point selected by the user, splice the first posture data and the fourth posture data to obtain the third posture data, insert the fourth video into the first video stream, and splice the fourth video stream and the first video stream to obtain the first video stream. That is, the user can also add operation demonstration content to the previous demonstration content.

[0076] Accordingly, Figure 8 and Fig. 9 The "deletion control" in the demo can also support users to delete content in the demonstration video. For example, users can view and delete image frames corresponding to unnecessary actions in the VR image, and then send the deleted video stream and posture data to the server.

[0077] It can be seen that in this example, the terminal device can respond to the user's selection operation on the target segment, trigger the VR wearable device to re-display the target VR sub-environment corresponding to the target segment, and re-collect the user posture data and video stream, and the terminal device can update the original posture data and video stream according to the re-collected posture data and video stream, and send the updated posture data and video stream to the server. The user can adjust the original data based on demand, so that the data of the sub-task sequence and atomic skill model parameters determined by the server are more adapted to user needs, which is conducive to further improving the flexibility of the robot in performing tasks.

[0078] In one possible example, displaying the display content corresponding to the first video stream in the demonstration video editing interface includes: receiving fifth information from the robot, the fifth information being used to indicate a subtask that was not successfully executed when the robot executed the subtask sequence; determining a reference segment corresponding to the subtask that was not successfully executed from the first video stream; determining the reference segment as the display content; and displaying the display content in the demonstration video editing interface.

[0079] In a specific implementation, the server may determine a corresponding end condition for each subtask. If the robot cannot achieve the end condition corresponding to the subtask within a preset time when executing the subtask, the robot may be judged to have failed in execution.

[0080] In the specific implementation, Figure 8 and Fig. 9 For example, after the terminal device receives the fifth information from the robot and determines the reference segment, the reference segment is displayed in the demonstration video editing interface in a manner such as: changing the display color of the display position of the reference segment in the white progress bar, for example Figure 8 and Fig. 9 The two shaded segments are shown in the middle shaded area.

[0081] In addition, in other embodiments, the reference segment can also be marked and set by the user, for example Fig. 9 As shown, the user can also select the time point to be located by dragging the time point positioning bar perpendicular to the white progress bar, and when the user's selection operation on the "Mark" control is detected, the mark type table is displayed through the mark type drop-down box to select the mark type of the time point, and a fragment confirmation prompt can be output to prompt the user whether to determine the two adjacent fragment start frames and fragment end frames as a reference video stream fragment. After the user confirmation operation is detected, the reference fragment can be displayed in the white progress bar.

[0082] In particular, the terminal device can also adjust the VR image displayed above the video boundary interface according to the position of the time point positioning bar, so that the VR image display screen is synchronized with the time point corresponding to the time positioning bar.

[0083] In actual applications, the terminal device can send both the first posture data and the first video stream to the server and the third posture data and the third video stream to the server. The server can then determine a subtask sequence and atomic skill model parameters of the task type based on the first posture data and the first video stream, and determine another subtask sequence and atomic skill model parameters of the task type based on the third posture data and the third video stream. The server can also obtain the priority set by the user for the first video stream and the third video stream based on the terminal device, and then determine the subtask sequence corresponding to the video stream with high priority as the default task sequence of the task type. From the subtask sequence corresponding to the video stream with low priority, determine the alternative subtasks corresponding to each subtask in the default task sequence. If, when the robot performs a task based on the default task sequence, there is an alternative subtask for the subtask that was not successfully executed, the robot can execute the alternative subtask.

[0084] For example, when the robot is performing a garbage throwing task, after pressing the elevator button, the robot monitors the elevator panel display content in real time, analyzes the elevator panel display content, and determines whether the elevator is operating normally. If it is operating normally, it determines that the current subtask is executed successfully. If the elevator displays a fault, it determines that the current subtask fails. At this time, the robot determines whether there is an alternative solution. For example, if the alternative solution is that there is a freight elevator in the current building, the robot modifies the current subtask to take the freight elevator. At this time, the robot executes the subtask, retrieves the freight elevator position, plans a moving path based on the current position, and moves to the freight elevator based on the planned moving path to execute the subtask. At the same time, when the robot is performing the garbage throwing task, it can also determine the capacity of the current trash can and the capacity of the current garbage to be thrown. Based on this, it determines whether throwing garbage in the current trash can is executable. If it is not executable, the robot can also generate a reason for the task not being executable and send the reason for not being executable to the terminal device. Then the robot abandons the task. If the task execution instruction issued by the terminal device based on the user operation is received again, the task can be executed again. In particular, the terminal device can also prompt the user of the reason for not being executable by means such as a pop-up window, etc., to improve the richness of information display.

[0085] It can be seen that in this example, the terminal device can also determine the reference segment corresponding to the subtask that was not successfully executed from the first video stream according to the fifth information from the robot, and determine the reference segment as the display content for display, which is conducive to quickly and accurately prompting the user of the reference segment corresponding to the subtask that was not successfully executed by the robot, and then accurately re-demonstrating the subtask that was not successfully executed.

[0086] In a possible example, the subtask sequence also includes start conditions and end conditions for each of the multiple subtasks, and the multiple subtasks include a first subtask. The end condition corresponding to the first subtask is determined according to a target interaction state, and the target interaction state is determined by identifying an end image frame corresponding to a video stream segment corresponding to the first subtask, and the target interaction state is an interaction state between the user and the target interaction object in the VR environment in the end image frame.

[0087] In a specific implementation, the end condition can be used as reference information to determine whether the robot has successfully executed a subtask. If the robot cannot achieve the end condition corresponding to the subtask within the preset time when executing the subtask, it can be determined that the robot has failed to execute.

[0088] For example, taking the first subtask of opening the door as an example, the server can identify the end image frame of the video stream segment corresponding to the first subtask, determine that the door is the target interaction object, and rotate the door handle and push it outward so that the door opens at a greater angle than the preset angle as the target interaction state. When the robot executes the first subtask, it uses the robot's dexterous hand to rotate the door handle and push it outward so that the door opens at a greater angle than the preset angle, and the first subtask is considered to be completed.

[0089] In a specific implementation, the server may split the first video stream through video stream splitting and image recognition processing to obtain the video stream segment corresponding to each subtask, or the server may also receive the video stream segment corresponding to each subtask from the terminal device by detecting the user's start image frame and end image frame operations.

[0090] In a specific implementation, the start condition of each subtask can be determined by performing image recognition on the start image frame of the video stream segment corresponding to each subtask. For example, the start condition can also be determined based on the interaction state of the interactive object in the start image frame. Alternatively, the end condition of a subtask that is executed before the current subtask can be determined as the start condition of the current subtask.

[0091] It can be seen that in this example, the subtask sequence also includes the start conditions and end conditions of each subtask in the multiple subtasks. The end condition corresponding to the first subtask is determined according to the target interaction state between the user and the target interaction object in the VR environment in the end image frame corresponding to the video stream segment corresponding to the first subtask. The end condition and start condition of the subtask are determined based on the detection of the image frame of the video stream segment corresponding to the subtask, and are sent to the humanoid robot, which is conducive to improving the accuracy and reliability of the humanoid robot in performing tasks.

[0092] In one possible example, the third information also carries timestamp information of the end image frame, and the timestamp information is determined in the following manner: displaying the display content corresponding to the first video stream through a demonstration video editing interface; in response to a selection operation on a target image frame in the display content, determining the target image frame as the end image frame of the video stream segment; and determining the timestamp information corresponding to the end image frame.

[0093] Among them, Fig. 9 Taking the demonstration video editing interface shown as an example, the selection operation for the target image frame may be, for example, detecting the selection operation for the "Mark" control, and detecting the selection operation for the "Clip End Frame" in the "Mark Type". At this time, the image frame corresponding to the time point positioning bar is the target image frame, and the time corresponding to the time point positioning bar is the timestamp information.

[0094] In a specific implementation, the server may determine the end image frame corresponding to the subtask according to the timestamp information sent by the terminal device.

[0095] It can be seen that in this example, the terminal device can determine that the target image frame is the end image frame of the video stream segment based on the selection operation of the target image frame in the first video stream, and carry the timestamp information corresponding to the end image frame in the third information to the server, so that the setting of the end condition is more in line with the actual needs of the user, further improving the flexibility of the robot task execution.

[0096] See also Fig.10 , Fig.10 is a functional unit block diagram of a task configuration device for a humanoid robot provided in an embodiment of the present application, which can be applied to Figure 1 The terminal device in the task configuration system of the humanoid robot shown in the figure, the humanoid robot control system includes: a terminal device, a virtual reality VR wearable device, a server and a humanoid robot, and the task configuration device 30 of the humanoid robot includes: An acquisition unit 301 is used to acquire task configuration information in response to an input operation based on the robot function configuration interface, wherein the task configuration information includes a task type; An output unit 302 is used to output first prompt information, where the first prompt information is used to prompt the user to wear the VR wearable device to perform a task operation demonstration; A first sending unit 303 is used to send first information to the VR wearable device, where the first information is used to instruct the VR wearable device to display a VR environment corresponding to the task type, where the VR environment is used to simulate an operation scenario of the task operation demonstration; A second sending unit 304 is used to send second information to the VR wearable device in response to a selection operation on a target control in the robot function configuration interface, wherein the second information is used to instruct the VR wearable device to send the first posture data of the user collected during the task operation demonstration and a first video stream corresponding to the actual display content of the VR wearable device; A receiving unit 305 is configured to receive the first posture data and the first video stream from the VR wearable device; The third sending unit 306 is used to send the third information carrying the task type, the first video stream and the first posture data to the server, the third information is used to instruct the server to send the association relationship between the task type and the subtask sequence, and at least one atomic skill model parameter to the robot, the subtask sequence is obtained by task splitting processing based on the first video stream, the subtask sequence includes multiple subtasks and the execution order of the multiple subtasks, the atomic skill model parameter is determined based on the first posture data, and the atomic skill model parameter is used to update at least one atomic skill corresponding to the robot and the multiple subtasks.

[0097] In a possible example, the task configuration information also includes task environment description information, and the VR environment is determined by the following steps: determining a VR reference environment corresponding to the task type from a plurality of preset VR basic environments; determining target environment elements according to the task environment description information, and the target environment elements include at least one of the following: building form, interactive object form; adjusting the VR reference environment according to the target environment elements to obtain the VR environment, and the VR environment includes the target environment elements.

[0098] In one possible example, the environment description information includes a target image, and the form of the interactive object is determined by the following steps: identifying the target image to determine at least one reference interactive object included in the target image; determining an interaction mode of the at least one reference interactive object; determining a scene interaction object from the at least one reference interactive object based on the interaction mode and the atomic skill corresponding to the robot, the atomic skill corresponding to the robot supporting the interaction mode of the scene interaction object; and determining the form of the interactive object based on the scene interaction object.

[0099] In a possible example, the task configuration device 30 of the humanoid robot is also used for: when the VR environment includes multiple VR sub-environments, after receiving the first posture data and the first video stream data from the VR wearable device, displaying the display content corresponding to the first video stream on a demonstration video editing interface; in response to a selection operation on a target segment in the display content, sending fourth information to the VR wearable device, the fourth information is used to instruct the VR wearable device to redisplay the target VR sub-environment corresponding to the target segment, the multiple VR sub-environments including the target VR sub-environment; outputting second prompt information, the second prompt information is used to prompt the user to wear the VR wearable device Prepare to perform task operation re-demonstration; receive second posture data and a second video stream from the VR wearable device, the second posture data is the posture data of the user collected during the task operation re-demonstration, and the second video stream is the video stream corresponding to the actual display content of the VR wearable device collected during the operation re-demonstration; according to the second posture data and the second video stream, update the first posture data and the first video stream to obtain third posture data and a third video stream, the third posture data includes the second posture data, and the third video stream includes the second video stream; send fourth information carrying the task type, the third video stream and the third posture data to the server.

[0100] In one possible example, in terms of displaying the display content corresponding to the first video stream in the demonstration video editing interface, the task configuration device 30 of the humanoid robot is specifically used to: receive fifth information from the robot, the fifth information being used to indicate a subtask that was not successfully executed when the robot executed the subtask sequence; determine a reference segment corresponding to the subtask that was not successfully executed from the first video stream; determine the reference segment as the display content; and display the display content in the demonstration video editing interface.

[0101] In a possible example, the subtask sequence also includes start conditions and end conditions for each of the multiple subtasks, and the multiple subtasks include a first subtask. The end condition corresponding to the first subtask is determined according to a target interaction state, and the target interaction state is determined by identifying an end image frame corresponding to a video stream segment corresponding to the first subtask, and the target interaction state is an interaction state between the user and the target interaction object in the VR environment in the end image frame.

[0102] In one possible example, the third information also carries timestamp information of the end image frame, and the timestamp information is determined in the following manner: displaying the display content corresponding to the first video stream through a demonstration video editing interface; in response to a selection operation on a target image frame in the display content, determining the target image frame as the end image frame of the video stream segment; and determining the timestamp information corresponding to the end image frame.

[0103] In the case of using integrated units, the functional unit composition block diagram of another humanoid robot task configuration device provided by the embodiment of the present application is as follows: Fig.11 As shown. Fig.11 In the embodiment, the task configuration device of the humanoid robot includes: a processing module 310 and a communication module 311. The processing module 310 is used to control and manage the actions of the task configuration device of the humanoid robot, for example, the steps performed by the acquisition unit 301, the output unit 302, the first sending unit 303, the second sending unit 304, the receiving unit 305 and the third sending unit 306, and / or other processes for performing the technology described in this application. The communication module 311 is used to support the interaction between the task configuration device of the humanoid robot and other devices. Fig.11 As shown, the task configuration device of the humanoid robot may further include a storage module 312, and the storage module 312 is used to store program codes and data of the task configuration device of the humanoid robot.

[0104] Among them, the processing module 310 can be a processor or a controller, for example, a central processing unit (CPU), a general processor, a digital signal processor (DSP), an ASIC, an FPGA or other programmable logic device, a transistor logic device, a hardware component or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of this application. The processor can also be a combination that implements a computing function, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like. The communication module 311 can be a transceiver, an RF circuit or a communication interface, and the like. The storage module 312 can be a memory.

[0105] Among them, all relevant contents of each scene involved in the above method embodiment can be referred to the functional description of the corresponding functional module, and will not be repeated here. The task configuration device of the above humanoid robot can execute the above Figure 3 The steps performed by the terminal device in the task configuration method of the humanoid robot shown.

[0106] An embodiment of the present application also provides a computer-readable storage medium, wherein an executable program code is stored on the computer-readable storage medium, and the executable program code includes execution instructions, and the execution instructions are used to execute part or all of the steps of the task configuration method of any humanoid robot recorded in the above method embodiment, and the computer includes a terminal device.

[0107] The present application also provides a computer program product, which includes a computer program, which is operable to cause a computer to execute some or all of the steps of any humanoid robot task configuration method described in the above method embodiment. The computer program product can be a software installation package, and the computer includes a terminal device.

[0108] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and units involved are not necessarily required by the present application.

[0109] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0110] In the several embodiments provided in the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only schematic, such as the division of the above-mentioned units, which is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.

[0111] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0112] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0113] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a memory, including a number of instructions to enable a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the above-mentioned methods in each embodiment of the present application. The aforementioned memory includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk and other media that can store program codes.

[0114] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable memory, which may include a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0115] The embodiments of the present application are introduced in detail above. Specific examples are used in the present application to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of ​​the present application. At the same time, for general technical personnel in the field, according to the idea of ​​the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A task configuration method for a humanoid robot, characterized in that: A terminal device applied to a humanoid robot control system, the humanoid robot control system comprising: the terminal device, a virtual reality VR wearable device, a server and the humanoid robot, the method comprising: In response to an input operation based on the robot function configuration interface, acquiring task configuration information, wherein the task configuration information includes a task type; Outputting first prompt information, where the first prompt information is used to prompt a user to wear the VR wearable device to perform a task operation demonstration; Sending first information to the VR wearable device, where the first information is used to instruct the VR wearable device to display a VR environment corresponding to the task type, where the VR environment is used to simulate an operation scenario of the task operation demonstration; In response to a selection operation on a target control in the robot function configuration interface, second information is sent to the VR wearable device, where the second information is used to instruct the VR wearable device to send the first posture data of the user collected during the task operation demonstration and a first video stream corresponding to the actual display content of the VR wearable device; Receiving the first posture data and the first video stream from the VR wearable device; Sending third information carrying the task type, the first video stream and the first posture data to the server, the third information is used to instruct the server to send the association between the task type and the subtask sequence, and at least one atomic skill model parameter to the robot, the subtask sequence is obtained by task splitting processing based on the first video stream, the subtask sequence includes multiple subtasks and the execution order of the multiple subtasks, the atomic skill model parameter is determined based on the first posture data, and the atomic skill model parameter is used to update at least one atomic skill corresponding to the robot and the multiple subtasks.

2. The method according to claim 1, characterized in that The task configuration information also includes task environment description information, and the VR environment is determined by the following steps: Determine a VR reference environment corresponding to the task type from a plurality of preset VR basic environments; Determine a target environment element according to the task environment description information, wherein the target environment element includes at least one of the following: a building form, an interactive object form; The VR reference environment is adjusted according to the target environment element to obtain the VR environment, and the VR environment includes the target environment element.

3. The method according to claim 2, characterized in that The environment description information includes a target image, and the interactive object form is determined by the following steps: Recognizing the target image to determine at least one reference interactive object included in the target image; determining an interaction mode of the at least one reference interaction object; Determine, according to the interaction mode and the atomic skill corresponding to the robot, a scene interaction object from the at least one reference interaction object, wherein the atomic skill corresponding to the robot supports the interaction mode of the scene interaction object; The interactive object form is determined according to the scene interactive object.

4. The method according to claim 1, characterized in that: The VR environment includes a plurality of VR sub-environments. After receiving the first posture data and the first video stream data from the VR wearable device, the method further includes: Displaying the display content corresponding to the first video stream on the demonstration video editing interface; In response to a selection operation on a target segment in the displayed content, fourth information is sent to the VR wearable device, where the fourth information is used to instruct the VR wearable device to redisplay a target VR sub-environment corresponding to the target segment, where the multiple VR sub-environments include the target VR sub-environment; Outputting second prompt information, where the second prompt information is used to prompt the user to wear the VR wearable device to perform a task operation re-demonstration; Receiving second posture data and a second video stream from the VR wearable device, wherein the second posture data is posture data of the user collected during the task operation re-demonstration process, and the second video stream is a video stream corresponding to actual display content of the VR wearable device collected during the operation re-demonstration process; According to the second posture data and the second video stream, the first posture data and the first video stream are updated to obtain third posture data and a third video stream, wherein the third posture data includes the second posture data, and the third video stream includes the second video stream; Sending fourth information carrying the task type, the third video stream and the third posture data to the server.

5. The method according to claim 4, characterized in that The display content corresponding to the first video stream displayed on the demonstration video editing interface includes: receiving fifth information from the robot, wherein the fifth information is used to indicate a subtask that was not successfully executed when the robot executes the subtask sequence; Determining, from the first video stream, a reference segment corresponding to the subtask that was not successfully executed; determining the reference segment as the display content; The display content is displayed in the demonstration video editing interface.

6. The method according to claim 1, characterized in that The subtask sequence also includes start conditions and end conditions for each of the multiple subtasks, the multiple subtasks include a first subtask, the end condition corresponding to the first subtask is determined according to a target interaction state, the target interaction state is determined by identifying an end image frame corresponding to a video stream segment corresponding to the first subtask, and the target interaction state is an interaction state between the user and the target interaction object in the VR environment in the end image frame.

7. The method according to claim 6, characterized in that The third information also carries the timestamp information of the end image frame, and the timestamp information is determined in the following manner: Displaying the display content corresponding to the first video stream through a demonstration video editing interface; In response to a selection operation on a target image frame in the display content, determining the target image frame as the end image frame of the video stream segment; The timestamp information corresponding to the end image frame is determined.

8. A task configuration device for a humanoid robot, characterized in that: A terminal device applied to a humanoid robot control system, the humanoid robot control system comprising: the terminal device, a virtual reality VR wearable device, a server and the humanoid robot, the device comprising: An acquisition unit, configured to acquire task configuration information in response to an input operation based on the robot function configuration interface, wherein the task configuration information includes a task type; An output unit, configured to output first prompt information, wherein the first prompt information is used to prompt a user to wear the VR wearable device to perform a task operation demonstration; A first sending unit, configured to send first information to the VR wearable device, wherein the first information is used to instruct the VR wearable device to display a VR environment corresponding to the task type, wherein the VR environment is used to simulate an operation scenario of the task operation demonstration; A second sending unit is used to send second information to the VR wearable device in response to a selection operation on a target control in the robot function configuration interface, wherein the second information is used to instruct the VR wearable device to send the first posture data of the user collected during the task operation demonstration and a first video stream corresponding to the actual display content of the VR wearable device; A receiving unit, configured to receive the first posture data and the first video stream from the VR wearable device; A third sending unit is used to send third information carrying the task type, the first video stream and the first posture data to the server, the third information is used to instruct the server to send the association relationship between the task type and the subtask sequence, and at least one atomic skill model parameter to the robot, the subtask sequence is obtained by task splitting processing based on the first video stream, the subtask sequence includes multiple subtasks and the execution order of the multiple subtasks, the atomic skill model parameter is determined based on the first posture data, and the atomic skill model parameter is used to update at least one atomic skill corresponding to the robot and the multiple subtasks.

9. A terminal device, characterized in that: The terminal device comprises: a processor, a memory and a communication interface, wherein the processor, the memory and the communication interface are interconnected, wherein the communication interface is used to receive or send data, the memory is used to store application code for the terminal device to execute the method as described in any one of claims 1-7, and the processor is configured to execute the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program for electronic data exchange, wherein the computer program enables a computer to execute the steps in the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Remote control system, information processing method, and program

    CN112154047A

  • Method for remotely operating humanoid robot by identifying hand postures through virtual reality glasses

    CN117746494A

  • Human-machine-data three-element integrated robot teleoperation and data acquisition system and method

    CN119304909A

  • Multi-mode wearable humanoid robot data acquisition and remote operation system

    CN119416153A

  • Robot teleoperation method and system based on virtual reality and digital twinning

    CN119501929A

Cited By

  • Unstructured environment supervised personal task data processing method and device

    CN120080326A

  • Robot control method and device, equipment and computer readable storage medium

    CN120503204A

  • Operation control method and device based on object flow sequence, equipment and medium

    CN120928733A

  • Method and device for constructing data set, equipment and storage medium

    CN121290363A

  • Task generation method, storage medium, electronic equipment and program product

    CN121541948A