Man-machine cooperation task processing method, device and equipment, medium and program product
By using belief models and advanced planners in a simulation environment of human-computer collaboration, the problem of conflict between robots and assisted persons in collaborative tasks is solved, and the task completion effect and robot collaboration capabilities are improved.
Patent Information
- Application Number
- CN202411941323.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-26
AI Technical Summary
The existing technology cannot effectively train or test the robot's active collaboration capabilities, resulting in conflicts when the robot and the assisted want to complete the same task at the same time, reducing the task completion effect.
By creating a simulation environment for human-computer collaboration, generate motion sequences using belief models and advanced planners, guide robots and virtual humans to complete tasks in collaboration and avoid conflicts.
It improves the effectiveness of robots and assisted persons completing tasks, and enhances the robot's active collaboration ability in collaborative scenarios.
Smart Images

Figure CN119937376A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of robotics technology, and in particular to a method, device, equipment, medium and program product for human-machine collaborative task processing. Background Art
[0002] With the development of robots and information technology, robots have expanded from the industrial field to fields such as home services, making home service robots gradually become a part of our daily lives. In order to effectively assist humans in the home environment, home service robots can complete a variety of complex tasks such as navigation, object recognition, and robotic arm operation. Combined with deep learning technology, using a large amount of training data to train robots can improve the robot's perception and recognition capabilities, thereby assisting in completing more tasks. However, deep learning requires a large amount of training data, and training robots directly in a real environment is costly and poses safety risks. As a branch of virtual reality technology, robot simulation can achieve robot testing and training in a computer simulation environment by creating a digital simulation environment.
[0003] In the prior art, humanoid agents observe the behaviors of the assisted persons, infer the tasks that the assisted persons want to complete, and assist the assisted persons in completing the corresponding tasks.
[0004] However, the solutions of the prior art cannot train or test the robot's active collaboration capabilities. When the robot and the assisted person want to complete the same task at the same time, conflicts will arise, resulting in reduced effectiveness of the robot and the assisted person in completing the task. Summary of the invention
[0005] The human-machine collaborative task processing method, device, equipment, medium and program product provided in the embodiments of the present application are used to improve the effectiveness of the robot and the assisted person in completing the task.
[0006] In a first aspect, an embodiment of the present application provides a human-machine collaborative task processing method, which is applied to a computer device, comprising:
[0007] Creating a human-machine collaboration simulation environment, wherein the simulation environment includes a robot and a virtual human;
[0008] Obtaining the task objectives and collecting the observable scene state diagram of the simulation environment to determine the completion progress of the task objectives;
[0009] If the task goal is not completed, inputting the observable scene state diagram into the belief model to output global scene estimation information of the robot and the virtual person;
[0010] Inputting global scene estimation information of the robot and the virtual human and the task goal into a high-level planner to output sub-goal instructions;
[0011] generating a motion sequence according to the sub-goal instructions;
[0012] The robot and the virtual human execute the motion sequence to generate an updated simulation environment;
[0013] Collecting the observable scene state diagram of the updated simulation environment to obtain a new observable scene state diagram;
[0014] Identify whether the task goal in the new observable scenario state diagram is completed;
[0015] If the task objective in the new observable scene state diagram is not completed, the step of "inputting the observable scene state diagram into the belief model to output global scene estimation information of the robot and the virtual person" is re-executed until the motion sequence of the robot and the virtual person reaches the task objective.
[0016] In a possible implementation, the inputting the observable scene state diagram into the belief model to output global scene estimation information of the robot and the virtual person includes: generating scene observation information of the robot and the virtual person according to the observable scene state diagram; the robot updates the belief model according to the scene observation information of the robot; the robot samples through the updated belief model to obtain the global scene estimation information of the robot; the virtual person updates the belief model according to the scene observation information of the virtual person; the virtual person samples through the updated belief model to obtain the global scene estimation information of the virtual person.
[0017] In a possible implementation, the global scene estimation information of the robot and the virtual person and the task goal are input into a high-level planner to output sub-goal instructions, including: parsing the task goal through the high-level planner to obtain multiple sub-goals; determining the type of the sub-goal; if the type of the sub-goal requires the robot and the virtual person to cooperate to complete, then inputting the global scene estimation information of the robot and the virtual person and the sub-goal into a Monte Carlo tree search model to output sub-goal instructions.
[0018] In a possible implementation, the inputting of the global scene estimation information of the robot and the virtual person and the sub-goal into a Monte Carlo tree search model to output a sub-goal instruction includes: creating a Monte Carlo tree search model of the robot and the virtual person, wherein the Monte Carlo tree search model includes multiple layers of nodes, and the multiple layers of nodes record the task rounds of the virtual person or the robot; setting the task round of the root node in the Monte Carlo tree search model as the task round of the virtual person, and setting the task round of the next layer of the task round of the virtual person as the task round of the robot; when the task round of the virtual person is During a task round, the virtual human derives and generates the sub-goals to be executed by the virtual human and the sub-goals to be executed by the robot in the next round based on the global scene estimation information of the virtual human and the sub-goals; when in the robot's task round, the robot derives and generates the sub-goals to be executed by the robot and the sub-goals to be executed by the virtual human in the next round based on the robot's global scene estimation information and the sub-goals; obtains the sub-goals to be executed by the robot in each task round and generates sub-goal instructions for the robot; obtains the sub-goals to be executed by the virtual human in each task round and generates sub-goal instructions for the virtual human.
[0019] In a possible implementation, generating a motion sequence according to the sub-goal instruction includes: extracting text information in the sub-goal instruction according to a regular expression through a low-level planner; identifying the text information to generate action information and information about an object to be interacted with; and generating a motion sequence according to the action information and the information about an object to be interacted with.
[0020] In a second aspect, an embodiment of the present application provides a human-machine collaborative task processing device, which is applied to a computer device, including:
[0021] A creation module, used to create a simulation environment for human-machine collaboration, wherein the simulation environment includes a robot and a virtual human;
[0022] An acquisition module is used to acquire the task target and collect the observable scene state diagram of the simulation environment to determine the completion progress of the task target;
[0023] A first output module, configured to input the observable scene state diagram into a belief model to output global scene estimation information of the robot and the virtual person if the task goal is not completed;
[0024] A second output module, for inputting the global scene estimation information of the robot and the virtual human and the task goal into a high-level planner to output a sub-goal instruction;
[0025] A first generating module, used for generating a motion sequence according to the sub-target instruction;
[0026] A second generating module, configured for the robot and the virtual human to execute the motion sequence to generate an updated simulation environment;
[0027] A collection module, used for collecting the observable scene state diagram of the updated simulation environment to obtain a new observable scene state diagram;
[0028] An identification module, used to identify whether the task goal in the new observable scene state diagram is completed;
[0029] The third output module is used to re-execute the step of "inputting the observable scene state diagram into the belief model to output the global scene estimation information of the robot and the virtual person" if the task goal in the new observable scene state diagram is not completed, until the motion sequence of the robot and the virtual person reaches the completion of the task goal.
[0030] In a possible implementation, the first output module includes: a first generation unit, used to generate scene observation information of the robot and scene observation information of the virtual person according to the observable scene state diagram; a first updating unit, used by the robot to update the belief model according to the scene observation information of the robot; a first sampling unit, used by the robot to sample through the updated belief model to obtain the robot's global scene estimation information; a second updating unit, used by the virtual person to update the belief model according to the scene observation information of the virtual person; a second sampling unit, used by the virtual person to sample through the updated belief model to obtain the virtual person's global scene estimation information.
[0031] In a third aspect, an embodiment of the present application provides a human-machine collaborative task processing device, including: a memory, a processor;
[0032] The memory stores computer-executable instructions;
[0033] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementations of the first aspect.
[0034] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the first aspect above and / or various possible implementations of the first aspect.
[0035] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above first aspect and / or various possible implementation methods of the first aspect.
[0036] The human-machine collaborative task processing method, device, equipment, medium and program product provided in the embodiments of the present application create a simulation environment for human-machine collaboration, obtain task goals and collect observable scene state diagrams, and judge the completion progress of the task goals. If the task goals are not completed, the observable scene state diagram is input into the belief model, and the belief model outputs global scene estimation information of the robot and the virtual person. The global scene estimation information is used to guide the current behavior of the robot and the virtual person and infer the behavior goals of each other. The high-level planner is used to generate sub-goal instructions and motion sequences. The robot and the virtual person execute the motion sequence to generate an updated simulation environment, and the observable scene state diagram of the updated simulation environment is collected. If the task goals in the updated observable scene state diagram are not completed, the steps of inputting the observable scene state diagram into the belief model and outputting the global scene estimation information are re-executed until the task goals are completed. The scene estimation information output by the belief model guides the current behavior of the intelligent agent and derives the behavior goals of the other party, thereby improving the effect of the robot and the assisted person in completing the task. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0038] Figure 1 A schematic diagram of the system structure of a computer device provided in an embodiment of the present application;
[0039] Figure 2 A flowchart of the human-machine collaborative task processing method provided in this application;
[0040] Figure 3 A flowchart of active assisted task planning execution based on a robot simulation environment provided in an embodiment of the present application;
[0041] Figure 4 A schematic diagram of a hierarchical planner provided in an embodiment of the present application;
[0042] Figure 5 A schematic diagram of a human-machine collaboration task target provided in an embodiment of the present application;
[0043] Figure 6 A schematic diagram of a task of moving a utility table provided in an embodiment of the present application;
[0044] Figure 7 A visual schematic diagram randomly generated for the scene arrangement of the four basic tasks provided in the embodiment of the present application;
[0045] Figure 8 A schematic diagram of the output processing flow of the advanced planner provided in an embodiment of the present application;
[0046] Fig. 9 A schematic diagram of a Monte Carlo tree search model provided in an embodiment of the present application;
[0047] Fig.10 A schematic diagram of the instruction execution process provided in the embodiment of the present application;
[0048] Fig.11 A schematic diagram of the structure of the human-machine collaborative task processing device provided in this application;
[0049] Fig.12 A schematic diagram of the structure of a human-machine collaborative task processing device provided in an embodiment of the present application.
[0050] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0051] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0052] With the development of robots and information technology, robots have expanded from the industrial field to fields such as home services, making home service robots gradually become a part of our daily life. In order to effectively assist humans in the home environment, home service robots can complete a variety of complex tasks such as navigation, object recognition and robotic arm operation. Combined with deep learning technology, using a large amount of training data to train robots can improve the robot's perception and recognition capabilities, thereby assisting in completing more tasks. However, deep learning requires a large amount of training data, and training robots directly in a real environment is costly and poses safety risks. Robot simulation, as a branch of virtual reality technology, can achieve robot testing and training in a computer simulation environment by creating a digital simulation environment. In the prior art, humanoid agents observe the behavior of the assisted person, infer the task that the assisted person wants to complete, and assist the assistant to complete the corresponding task. However, the prior art solutions cannot train or test the robot's active collaboration ability. When the robot and the assisted person want to complete the same task at the same time, conflicts will occur, resulting in reduced effectiveness of the robot and the assisted person in completing the task.
[0053] In order to solve the above technical problems, the embodiment of the present application proposes the following technical ideas: the inventor considers creating a simulation environment for human-machine collaboration, sets virtual humans and robots in the simulation environment, and judges the completion progress of the task goal by obtaining the observable scene state diagram and task goal of the simulation environment. If the task goal is not completed, consider creating a belief model, and input the observable scene state diagram into the belief model. The belief model outputs the global scene estimation information of the robot and the virtual human, and guides the current behavior of the robot and the virtual human through the global scene estimation information and infers the behavior goal of the other party. The global scene estimation information is input into the high-level planner, the sub-goal instructions are output, and the motion sequence is generated. The robot and the virtual human execute the motion sequence to generate an updated simulation environment. By collecting the observable scene state diagram of the updated simulation environment, it is judged whether the task goal in the new observable scene state diagram is completed. If the task goal is not completed, the observable scene state diagram is input into the belief model again, and the step of outputting the global scene estimation information is re-executed until the task goal is completed. The following detailed embodiments are used for detailed description.
[0054] Figure 1 A schematic diagram of the system structure of a computer device provided in an embodiment of the present application. Figure 1 As shown, the computer device includes: a receiving device 101, a processor 102 and a display device 103.
[0055] It is understandable that the structure illustrated in the embodiment of the present application does not constitute a specific limitation on the object identification method. In other feasible implementations of the present application, the above architecture may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or arrange the components differently, which can be determined according to the actual application scenario and is not limited here. Figure 1 The components shown may be implemented in hardware, software, or a combination of software and hardware.
[0056] In a specific implementation process, the receiving device 101 may be an input / output interface or a communication interface, and may acquire a task target.
[0057] The processor 102 may generate a motion sequence and generate an observable scene state diagram.
[0058] The display device 103 can be used to display the above-mentioned observable scene state diagram, etc.
[0059] The display device may also be a touch display screen, which is used to receive user instructions while displaying the above-mentioned content, so as to realize operational interaction with the user.
[0060] It should be understood that the above-mentioned processor can be implemented by the processor reading instructions in the memory and executing the instructions, or it can be implemented by a chip circuit.
[0061] In addition, the network architecture and business scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Ordinary technicians in this field can know that with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0062] Figure 2 A flowchart of the human-machine collaborative task processing method provided in this application is as follows: Figure 2 As shown, the method includes:
[0063] S201: Create a simulation environment for human-machine collaboration, where the simulation environment includes a robot and a virtual human.
[0064] Figure 3 A flowchart of the execution of active assisted task planning based on a robot simulation environment provided in an embodiment of the present application.
[0065] In this embodiment, the simulation environment includes a high-level planner and a low-level planner.
[0066] Figure 4 A schematic diagram of a hierarchical planner provided in an embodiment of the present application.
[0067] like Figure 4 As shown, the high-level planner is used to plan which actions need to be performed, and the low-level planner focuses on executing specific actions.
[0068] For example, for the task of setting a table, it is necessary to place several plates, cups and other tableware. The high-level planner will plan instructions according to the task objectives, such as placing plate 1 on the table and cup 2 on the table. The low-level planner will parse the actions and the objects that need to be interacted with. According to the actions and the objects to be interacted with, the low-level planner will execute the behavior.
[0069] During the training and testing process, the high-level planner and low-level planner of the virtual human share the task process with the high-level planner and low-level planner of the robot. During the testing process, developers can manually control the interface of the virtual human to collaborate on tasks.
[0070] S202: Acquire the task target and collect the observable scene state diagram of the simulation environment to determine the completion progress of the task target.
[0071] Specifically, an observable scene state diagram of the simulation environment is obtained, and whether the task goal is completed is determined based on the status of the virtual person, robot, and objects involved in the assistance task in the state diagram.
[0072] In this embodiment, there are three types of task objectives: objectives that do not require assistance, objectives that can be completed independently, and objectives that must be completed in cooperation with others.
[0073] For example, a goal that does not require assistance might be for a virtual person to read on the couch; a goal that can be completed independently might be for a virtual person to set the dining table; and a goal that must be completed in collaboration with others might be for a virtual person to move the table.
[0074] Figure 5 A schematic diagram of the human-machine collaboration task objectives provided in an embodiment of the present application.
[0075] like Figure 5 As shown, Figure 5 Scene A in the video is that the virtual person is sitting on the sofa watching TV, which is a goal that does not require help. Scene B is that the virtual person picks up items on the table, which is a goal that can be completed independently. Scene C is that the virtual person and the robot lift and move the table together, which is a goal that can be completed in cooperation with others.
[0076] Specifically, four tasks and corresponding sub-goals are set according to the type of task objectives in the simulation environment. According to different task objectives, the robot provides assistance to the virtual person. Among them, setting the dining table, recycling parameters and reading and leisure can be completed independently by the robot or the virtual person, and moving the utility table requires the robot and the virtual person to work together.
[0077] Among them, Table 1 is a record table of four human-machine collaborative tasks and their sub-goal types.
[0078] Table 1 Four human-machine collaborative tasks and their sub-goal types
[0079]
[0080] S203: If the task goal is not completed, the observable scene state diagram is input into the belief model to output the global scene estimation information of the robot and the virtual human.
[0081] Specifically, the robot and the virtual human update the belief model according to the scene observation information, and obtain the global scene estimation information of the robot and the virtual human through sampling of the updated belief model.
[0082] S204: Input the global scene estimation information and task objectives of the robot and virtual human into the high-level planner to output sub-goal instructions.
[0083] Specifically, the high-level planner parses the task goal to obtain multiple sub-goals, and inputs the global scene estimation information and sub-goals into the Monte Carlo tree search model to output sub-goal instructions.
[0084] S205: Generate a motion sequence according to the sub-target instruction.
[0085] Specifically, a low-level planner extracts text information according to regular expressions, identifies text information to generate action information and information about items to be interacted with, and generates a motion sequence.
[0086] S206: The robot and the virtual human execute a motion sequence to generate an updated simulation environment.
[0087] For example, Figure 6 A schematic diagram of the task of executing a moving utility table provided in an embodiment of the present application.
[0088] like Figure 6 As shown, Figure 6 In scene A, the robot found the debris on the table and took away the cup on the table; in scene B, the robot took away the small box on the table; in scene C, the virtual person found and took away the last debris cup on the table and placed the cup in the living room; in scene D, the virtual person lifted the table and prepared to move it, but the robot was not in place; in scene E, the robot was in place and lifted and moved the table together with the virtual person; in scene F, the robot and the virtual person put down the table together.
[0089] S207: Collect the observable scene state diagram of the updated simulation environment to obtain a new observable scene state diagram.
[0090] In this embodiment, the actions of the virtual human or robot will continuously change the state of the simulation environment and generate a new observable scene state diagram.
[0091] S208: Identify whether the task objectives in the new observable scenario state diagram are completed.
[0092] Specifically, the states of the robot, the virtual person, and the objects in the task target in the new observable scene state diagram are judged to determine whether the task target is completed.
[0093] S209: If the task goal in the new observable scene state diagram is not completed, re-execute the step of "inputting the observable scene state diagram into the belief model to output the global scene estimation information of the robot and the virtual person" until the movement sequence of the robot and the virtual person reaches the task goal.
[0094] In this embodiment, the training methods for human-machine collaborative tasks in the prior art include a watch-and-help method and a NOPA (Neurally-guided Online Probabilistic Assistance, referred to as NOPA) method.
[0095] In this embodiment, seven basic test scenarios are set. Table 2 lists all tasks to be tested and their generated data, and displays the index numbers of the basic test scenarios that can be used to generate task data.
[0096] Table 2 Detailed information on scene layout generation
[0097]
[0098] Figure 7 A visual schematic diagram is randomly generated for the scene arrangement of the four basic tasks provided in the embodiments of the present application.
[0099] like Figure 7 As shown, Figure 7 Scene A is about setting up the dining table. Only fruit platters and condiments are left on the table, and the tableware is scattered in containers such as the refrigerator and cabinets in the kitchen. Scene B is about recycling tableware. The tableware is placed on the dining table, refrigerator, etc. in the kitchen and needs to be placed in the dishwasher. Scene C is about reading and leisure. Snacks and drinks are scattered in the scene, and the virtual person will eventually sit on the sofa to read. Scene D is about moving a utility table. The utility table is randomly placed in the room, and a table needs to be moved to the next room.
[0100] Table 3 shows the execution results of the four tasks. The efficiency is measured by the average number of steps to complete the task. In the task of moving the utility table, the sub-goal of cleaning the table is additionally considered when calculating the number of steps. The results in the table are divided into no robot intervention, random reinforcement learning method, watch-and-help method, NOPA method and the method of this scheme. The results show that the random reinforcement learning method (watch-and-help and NOPA) method cannot converge in the task of moving the utility table due to the large action space and order sensitivity, and often fails due to timeout. The behavior of the robot under the reinforcement learning method may interfere with the completion of the task, such as mistakenly taking books that should not be taken.
[0101] Table 3 Average number of single-target execution steps of active assistance methods on four types of combined tasks
[0102]
[0103] It can be seen from the above embodiments that by creating a simulation environment for human-machine collaboration, obtaining task goals and collecting observable scene state diagrams, the completion progress of the task goals is judged. If the task goals are not completed, the observable scene state diagram is input into the belief model, and the belief model outputs global scene estimation information of the robot and the virtual person. The global scene estimation information is used to guide the current behavior of the robot and the virtual person and infer the behavior goals of each other. The high-level planner is used to generate sub-goal instructions and motion sequences. The robot and the virtual person execute the motion sequence, generate an updated simulation environment, and collect the observable scene state diagram of the updated simulation environment. If the task goals in the updated observable scene state diagram are not completed, the steps of inputting the observable scene state diagram into the belief model and outputting the global scene estimation information are re-executed until the task goals are completed. The scene estimation information output by the belief model guides the current behavior of the intelligent agent and derives the behavior goals of the other party, thereby improving the effect of the robot and the assisted person in completing the task.
[0104] In one embodiment of the present application, step S203 includes:
[0105] S2031: Generate scene observation information of the robot and scene observation information of the virtual person according to the observable scene state diagram.
[0106] Figure 8 A schematic diagram of the output processing flow of the advanced planner provided in an embodiment of the present application.
[0107] In this embodiment, the scene observation information observed by the robot and the virtual person at a set time is part of the information in the entire simulation environment.
[0108] S2032: The robot updates the belief model based on the robot's scene observation information.
[0109] In this embodiment, the belief model is the agent's probability estimate of the distribution of objects in the scene.
[0110] For example, the robot believes that the probability that the book is in the living room is 30%, in the kitchen is 20%, and in the bedroom is 50%.
[0111] S2033: The robot samples through the updated belief model to obtain the robot's global scene estimation information.
[0112] Specifically, the robot updates the scene by executing sub-goals, obtains scene observation information in the new scene, updates the belief model through the scene observation information in the new scene, and samples from the updated belief model when executing the next sub-goal to obtain global scene estimation information.
[0113] S2034: The virtual human updates the belief model according to the scene observation information of the virtual human.
[0114] In this embodiment, unobserved items are assigned an average estimate, and when an item is observed, its position in the belief model is updated to the exact position. When an item is out of sight, its estimate gradually regresses to the initial average estimate.
[0115] The belief model is initialized based on the global state graph, which provides information about the existence of all objects in the scene, but does not specify their locations. In order to use this belief information, it is necessary to obtain a global state estimate of the scene through sampling, including the possible locations of objects, the on / off states of electrical appliances, and the on / off states of furniture, so as to generate a detailed global state graph.
[0116] S2035: The virtual human samples through the updated belief model to obtain global scene estimation information of the virtual human.
[0117] Specifically, the virtual human updates the scene by executing sub-goals, obtains scene observation information in the new scene, updates the belief model through the scene observation information in the new scene, and samples from the updated belief model when executing the next sub-goal to obtain global scene estimation information.
[0118] It can be seen from the above embodiments that the belief model is updated through the scene observation information of the robot and the scene observation information of the virtual person. The robot and the virtual person sample from the updated belief model to obtain their respective corresponding global scene estimation information. The global scene estimation information is used to guide their respective current behaviors and infer the behavior goals of the other party, thereby realizing collaboration between the robot and the virtual person.
[0119] In one embodiment of the present application, step S204 includes:
[0120] S2041: Analyze the task goal through the high-level planner to obtain multiple sub-goals.
[0121] In this embodiment, the task goal includes multiple sub-goals.
[0122] Exemplarily, the task goal of moving the utility table includes the sub-goals of picking up the mug on the utility table, moving the utility table to a target location, and placing milk on the target utility table.
[0123] S2042: Determine the type of sub-goal.
[0124] In this embodiment, the types of sub-goals include goals that do not require assistance, goals that can be completed independently, and goals that must be completed in cooperation with others.
[0125] S2043: If the type of the sub-goal is that it requires cooperation between the robot and the virtual person to complete, the global scene estimation information of the robot and the virtual person and the sub-goal are input into the Monte Carlo tree search model to output the sub-goal instruction.
[0126] In this embodiment, during the collaborative task, the robot and the virtual person need to update their beliefs based on observations and sample from their beliefs to obtain a global state estimate of the scene. This estimate information is used to guide the current behavior of the agent and infer the behavior goals of the other agent.
[0127] It can be seen from the above embodiments that by dividing the task goal into sub-goals and judging the type of sub-goals, if the sub-goal is a goal that needs to be completed in cooperation, the Monte Carlo tree search model is used to generate sub-goal instructions for the virtual person and the robot based on the global scene estimation information and sub-goals, so that the robot and the virtual person can collaborate to complete the task.
[0128] In one embodiment of the present application, step S2043 includes:
[0129] S301: Create a Monte Carlo tree search model for the robot and the virtual person, wherein the Monte Carlo tree search model includes multiple layers of nodes, and the multiple layers of nodes record the task rounds of the virtual person or the robot.
[0130] Fig. 9 A schematic diagram of a Monte Carlo tree search model provided in an embodiment of the present application.
[0131] S302: setting the task round of the root node in the Monte Carlo tree search model as the task round of the virtual person, and setting the task round of the next layer of the task round of the virtual person as the task round of the robot.
[0132] In this embodiment, the virtual human is defined as agent A and the robot is defined as agent B.
[0133] For example, in the task of moving a table, the cups and remote control on the table need to be cleared before the table can be lifted, and then the task of moving the table to the target location can be performed.
[0134] Specifically, the root node is set to the turn of agent A, the second-layer node is set to the turn of agent B, and so on, and the nodes of the tree are expanded in turn through agent A executing sub-goals.
[0135] S303: When the virtual person is in the task round, the virtual person derives and generates the sub-goals to be executed by the virtual person and the sub-goals to be executed by the robot in the next round based on the virtual person's global scene estimation information and sub-goals.
[0136] Specifically, the virtual human derives its own sub-goals based on its own beliefs and task objectives, and also infers the possible sub-goals of agent B.
[0137] S304: When the robot is in its task round, the robot derives and generates sub-goals to be executed by the robot and sub-goals to be executed by the virtual person in the next round based on the robot's global scene estimation information and sub-goals.
[0138] Specifically, the robot derives its own sub-goals based on its own beliefs and task objectives, and also infers the possible sub-goals of agent A.
[0139] S305: Obtain the sub-goals to be executed by the robot in each task round, and generate sub-goal instructions for the robot.
[0140] Specifically, the nodes with the optimal UCT (Upper Confidence Bound Apply to Trees, UCT for short) value in the Monte Carlo Tree Search model and belonging to the B round of the intelligent agent are traversed, the sub-goals executed by the nodes are recorded, and the sub-goal instructions are generated.
[0141] S306: Obtain the sub-goals to be executed by the virtual person in each task round, and generate sub-goal instructions for the virtual person.
[0142] Specifically, the nodes with the optimal UCT value in the Monte Carlo tree search model and belonging to round A of the intelligent agent are traversed, the sub-goals executed by the nodes are recorded, and sub-goal instructions are generated.
[0143] It can be seen from the above embodiments that by creating a Monte Carlo tree search model, setting the root node in the model as the virtual person's turn, and the virtual person's lower-level task turn as the robot's turn, the loop setting is executed in sequence. When in the virtual person's task turn, the virtual person derives the sub-goals to be executed by the virtual person and the sub-goals to be executed by the robot in the next turn, and the robot derives the sub-goals to be executed by the robot and the sub-goals to be executed by the virtual person in the next turn. The sub-goals of all turns of the robot and the virtual person are generated into corresponding sub-goal instructions. The robot and the virtual person derive the sub-goals that each other may execute and generate target instructions, thereby avoiding task conflicts between the robot and the virtual person when performing tasks and improving the efficiency of collaborative task execution.
[0144] In one embodiment of the present application, step S205 includes:
[0145] S2051: Extract text information in sub-goal instructions according to regular expressions through a low-level planner.
[0146] Specifically, the low-level planner receives and parses instructions, and uses regular expressions to extract key text information from the text of the instructions.
[0147] The text information includes but is not limited to the action to be performed, the category of the item to be interacted with, and the instance ID.
[0148] S2052: Identify text information, generate action information and information about items to be interacted with.
[0149] Specifically, the low-level planner identifies specific interactive items from the constructed state graph and sends the action information and item information to the action executor.
[0150] S2053: Generate a motion sequence according to the action information and the information of the object to be interacted with.
[0151] Specifically, the action executor generates motion sequences based on the basic modules in the simulation environment.
[0152] Fig.10 A schematic diagram of the instruction execution process provided in an embodiment of the present application.
[0153] It can be seen from the above embodiments that the text information in the instruction is recognized by the low-level planner, the actions recorded in the text information and the objects to be interacted are parsed, and a motion sequence is generated according to the actions and the objects to be interacted, so as to avoid conflicts in the execution actions of the robot and the intelligent agent.
[0154] Fig.11 A schematic diagram of the structure of the human-machine collaborative task processing device provided in this application, such as Fig.11 As shown, the human-machine collaborative task processing device 11 provided in this embodiment includes: a creation module 111, an acquisition module 112, a first output module 113, a second output module 114, a first generation module 115, a second generation module 116, a collection module 117, an identification module 118 and a third output module 119.
[0155] The creation module 111 is used to create a simulation environment for human-machine collaboration, wherein the simulation environment includes a robot and a virtual human.
[0156] The acquisition module 112 is used to acquire the task target and collect the observable scene state diagram of the simulation environment to determine the completion progress of the task target.
[0157] The first output module 113 is used to input the observable scene state diagram into the belief model to output global scene estimation information of the robot and the virtual person if the task goal is not completed.
[0158] The second output module 114 is used to input the global scene estimation information and task objectives of the robot and the virtual human into the high-level planner to output sub-goal instructions.
[0159] The first generating module 115 is used to generate a motion sequence according to the sub-target instruction.
[0160] The second generation module 116 is used for the robot and the virtual human to execute a motion sequence to generate an updated simulation environment.
[0161] The acquisition module 117 is used to acquire the observable scene state diagram of the updated simulation environment to obtain a new observable scene state diagram.
[0162] The identification module 118 is used to identify whether the task objectives in the new observable scene state diagram are completed.
[0163] The third output module 119 is used to re-execute the step of "inputting the observable scene state diagram into the belief model to output the global scene estimation information of the robot and the virtual person" if the task goal in the new observable scene state diagram is not completed, until the motion sequence of the robot and the virtual person reaches the task goal.
[0164] In a possible implementation, the first output module 113 includes:
[0165] The first generating unit 131 is used to generate scene observation information of the robot and scene observation information of the virtual person according to the observable scene state diagram.
[0166] The first updating unit 132 is used for the robot to update the belief model according to the scene observation information of the robot.
[0167] The first sampling unit 133 is used for the robot to sample through the updated belief model to obtain the global scene estimation information of the robot.
[0168] The second updating unit 134 is used for the virtual human to update the belief model according to the scene observation information of the virtual human.
[0169] The second sampling unit 135 is used for the virtual human to sample through the updated belief model to obtain the global scene estimation information of the virtual human.
[0170] In a possible implementation, the second output module 114 includes:
[0171] The parsing unit 141 is used to parse the task goal through the high-level planner to obtain multiple sub-goals.
[0172] The judging unit 142 is used to judge the type of the sub-goal.
[0173] The output unit 143 is used to input the global scene estimation information of the robot and the virtual human and the sub-goal into the Monte Carlo tree search model to output the sub-goal instruction if the type of the sub-goal requires the cooperation of the robot and the virtual human to complete.
[0174] In a possible implementation, the output unit 143 includes:
[0175] The creation subunit 1431 is used to create a Monte Carlo tree search model for the robot and the virtual person, wherein the Monte Carlo tree search model includes multiple layers of nodes, and the multiple layers of nodes record the task rounds of the virtual person or the robot.
[0176] The setting subunit 1432 is used to set the task round of the root node in the Monte Carlo tree search model as the task round of the virtual person, and set the next layer of task round of the virtual person as the task round of the robot.
[0177] The first generation subunit 1433 is used to derive and generate the sub-goals to be executed by the virtual human and the sub-goals to be executed by the robot in the next round based on the virtual human's global scene estimation information and sub-goals when the virtual human is in the task round.
[0178] The second generation subunit 1434 is used to derive and generate the sub-goals to be executed by the robot and the sub-goals to be executed by the virtual human in the next round according to the robot's global scene estimation information and sub-goals when the robot is in its task round.
[0179] The first acquisition subunit 1435 is used to acquire the sub-goals to be executed by the robot in each task round and generate sub-goal instructions for the robot.
[0180] The second acquisition subunit 1436 is used to acquire the sub-goals to be executed by the virtual person in each task round and generate sub-goal instructions for the virtual person.
[0181] In a possible implementation, the first generating module 115 includes:
[0182] The extraction unit 151 is used to extract text information in the sub-goal instruction according to the regular expression through the low-level planner.
[0183] The recognition unit 152 is used to recognize text information and generate action information and information about items to be interacted with.
[0184] The second generating unit 153 is used to generate a motion sequence according to the action information and the information of the object to be interacted with.
[0185] The human-machine collaborative task processing device provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and this embodiment will not be described in detail here.
[0186] Fig.12 This is a schematic diagram of the structure of the human-machine collaborative task processing device provided in the embodiment of the present application. Fig.12As shown, the human-machine collaborative task processing device 120 provided in this embodiment includes: at least one processor 1201 and a memory 1202. Optionally, the human-machine collaborative task processing device 120 also includes a communication component 1203. The processor 1201, the memory 1202 and the communication component 1203 are connected via a bus 1204.
[0187] In the specific implementation process, at least one processor 1201 executes the computer execution instructions stored in the memory 1202, so that at least one processor 1201 executes the above-mentioned human-machine collaborative task processing method.
[0188] The specific implementation process of the processor 1201 can be found in the above method embodiment, and its implementation principle and technical effect are similar, so this embodiment will not be repeated here.
[0189] In the above embodiments, it should be understood that the processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the invention can be directly implemented as a hardware processor, or can be implemented by a combination of hardware and software modules in the processor.
[0190] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (NVM), such as at least one disk storage.
[0191] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.
[0192] The present application also provides a computer program product, including a computer program, which implements the above-mentioned human-machine collaborative task processing method when executed by a processor.
[0193] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the above-mentioned human-machine collaborative task processing method is implemented.
[0194] The above-mentioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general or special-purpose computer.
[0195] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (Application Specific Integrated Circuits, referred to as: ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.
[0196] The division of units is only a logical function division, and there may be other divisions in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.
[0197] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0198] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0199] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0200] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.
[0201] Finally, it should be noted that those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses or adaptations of the present invention, which follow the general principles of the present invention and include common knowledge or customary technical means in the art not disclosed by the present invention, are not limited to the precise structure described above and shown in the drawings, and may be modified and changed in various ways without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.
Claims
1. A method for processing human-machine collaborative tasks, characterized in that: Applicable to computer equipment, including: Creating a human-machine collaboration simulation environment, wherein the simulation environment includes a robot and a virtual human; Obtaining the task objectives and collecting the observable scene state diagram of the simulation environment to determine the completion progress of the task objectives; If the task goal is not completed, inputting the observable scene state diagram into the belief model to output global scene estimation information of the robot and the virtual person; Inputting global scene estimation information of the robot and the virtual human and the task goal into a high-level planner to output sub-goal instructions; generating a motion sequence according to the sub-goal instructions; The robot and the virtual human execute the motion sequence to generate an updated simulation environment; Collecting the observable scene state diagram of the updated simulation environment to obtain a new observable scene state diagram; Identify whether the task goal in the new observable scenario state diagram is completed; If the task objective in the new observable scene state diagram is not completed, the step of "inputting the observable scene state diagram into the belief model to output global scene estimation information of the robot and the virtual person" is re-executed until the motion sequence of the robot and the virtual person reaches the task objective.
2. The method according to claim 1, characterized in that The step of inputting the observable scene state diagram into a belief model to output global scene estimation information of the robot and the virtual person includes: Generate scene observation information of the robot and scene observation information of the virtual person according to the observable scene state diagram; The robot updates the belief model according to scene observation information of the robot; The robot samples through the updated belief model to obtain global scene estimation information of the robot; The virtual person updates the belief model according to the scene observation information of the virtual person; The virtual human is sampled through the updated belief model to obtain global scene estimation information of the virtual human.
3. The method according to claim 1, characterized in that The step of inputting the global scene estimation information of the robot and the virtual person and the task goal into a high-level planner to output sub-goal instructions comprises: Parsing the task goal by the high-level planner to obtain a plurality of sub-goals; Determining the type of the sub-goal; If the type of the sub-goal is that it requires the cooperation of the robot and the virtual person to complete, the global scene estimation information of the robot and the virtual person and the sub-goal are input into a Monte Carlo tree search model to output a sub-goal instruction.
4. The method according to claim 3, characterized in that The step of inputting the global scene estimation information of the robot and the virtual person and the sub-goal into a Monte Carlo tree search model to output a sub-goal instruction comprises: Creating a Monte Carlo tree search model of the robot and the virtual person, wherein the Monte Carlo tree search model includes multiple layers of nodes, and the multiple layers of nodes record the task rounds of the virtual person or the robot; The task round of the root node in the Monte Carlo tree search model is set as the task round of the virtual person, and the task round of the next layer of the task round of the virtual person is set as the task round of the robot; When the virtual person is in the task round, the virtual person derives and generates the sub-goal to be executed by the virtual person and the sub-goal to be executed by the robot in the next round according to the global scene estimation information of the virtual person and the sub-goal; When the robot is in a task round, the robot derives and generates a sub-goal to be executed by the robot and a sub-goal to be executed by the virtual person in the next round according to the robot's global scene estimation information and the sub-goal; Obtaining the sub-goals to be executed by the robot in each task round, and generating sub-goal instructions for the robot; Obtain the sub-goals to be executed by the virtual person in each task round, and generate sub-goal instructions for the virtual person.
5. The method according to any one of claims 1 to 4, characterized in that: The step of generating a motion sequence according to the sub-target instruction comprises: Extracting text information in the sub-goal instructions according to regular expressions using a low-level planner; Recognize the text information, and generate action information and information of an item to be interacted with; A motion sequence is generated according to the action information and the information of the object to be interacted with.
6. A human-machine collaborative task processing device, characterized in that: Applicable to computer equipment, including: A creation module, used to create a simulation environment for human-machine collaboration, wherein the simulation environment includes a robot and a virtual human; An acquisition module is used to acquire the task target and collect the observable scene state diagram of the simulation environment to determine the completion progress of the task target; A first output module, configured to input the observable scene state diagram into a belief model to output global scene estimation information of the robot and the virtual person if the task goal is not completed; A second output module, for inputting the global scene estimation information of the robot and the virtual human and the task goal into a high-level planner to output a sub-goal instruction; A first generating module, used for generating a motion sequence according to the sub-target instruction; A second generating module, configured for the robot and the virtual human to execute the motion sequence to generate an updated simulation environment; A collection module, used for collecting the observable scene state diagram of the updated simulation environment to obtain a new observable scene state diagram; An identification module, used to identify whether the task goal in the new observable scene state diagram is completed; The third output module is used to re-execute the step of "inputting the observable scene state diagram into the belief model to output the global scene estimation information of the robot and the virtual person" if the task goal in the new observable scene state diagram is not completed, until the motion sequence of the robot and the virtual person reaches the task goal.
7. The device according to claim 6, characterized in that The first output module comprises: A first generating unit, configured to generate scene observation information of the robot and scene observation information of the virtual person according to the observable scene state diagram; A first updating unit, configured for the robot to update the belief model according to scene observation information of the robot; A first sampling unit, configured for the robot to sample through the updated belief model to obtain global scene estimation information of the robot; A second updating unit, configured for the virtual human to update the belief model according to the scene observation information of the virtual human; The second sampling unit is used for sampling the virtual person through the updated belief model to obtain global scene estimation information of the virtual person.
8. A human-machine collaborative task processing device, characterized in that: include: at least one processor and memory; The memory stores computer-executable instructions; The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the human-machine collaborative task processing method according to any one of claims 1 to 5.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the human-machine collaborative task processing method according to any one of claims 1 to 5.
10. A computer program product, characterized in that It includes a computer program, which, when executed by a processor, implements the human-machine collaborative task processing method described in any one of claims 1 to 5.
Citation Information
Patent Citations
Intelligent manufacturing method and system based on man-machine collaboration
CN112936267A
Man-machine cooperation method and system based on scene perception and robot
CN114092791A
Distributed multi-agent cooperation method, system, medium and equipment
CN116578636A
Control method and system for cooperative work of multiple robots and robot
CN117957500A
Resilient multi-robot system with social learning for smart factories
US20230311312A1