Robot-based task execution method and device, computer equipment, readable storage medium and program product

By disassembling complex tasks into subtasks and generating task execution diagrams, the problem of reduced processing performance of robots when executing complex tasks is solved, and efficient and accurate execution of tasks is achieved.

CN119974026AActive Publication Date: 2025-05-13SHANGHAI FOURIER INTELLIGENCE CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510473545.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

Existing robots cannot execute smoothly due to lack of definition when performing complex tasks, resulting in reduced processing performance.

Method used

Through a large model, complex tasks are disassembled into subtasks, and a task execution diagram is generated based on the front and back relationships of the subtasks, the front and back relationships of actions and skills, the state and pose of the robot. Finally, the task execution target path is generated based on the task execution diagram to achieve accurate execution of the task.

Benefits of technology

Improve task processing efficiency, ensure accurate execution of tasks, and effectively handle complex tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119974026A_ABST
    Figure CN119974026A_ABST
Patent Text Reader

Abstract

The invention relates to a task execution method and device based on a robot, computer equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: under the condition that a task type is determined to be a complex task type based on a received voice signal, disassembling a complex task through a large model to obtain sub-tasks and a front-back relationship corresponding to the sub-tasks; disassembling each sub-task to obtain an action and / or a skill corresponding to each sub-task; a task execution graph is generated according to the front-back relation of the subtasks, the actions and skills obtained through disassembly, the front-back relation of the actions and skills and the state and pose of the robot; and generating a task execution target path based on the task execution graph, and performing task execution based on actions and skills on the target path to obtain a task execution result. By adopting the method, the task processing efficiency can be improved, and accurate execution of the task is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of robot instruction technology, and in particular to a robot-based task execution method, device, computer equipment, computer-readable storage medium and computer program product. Background Art

[0002] With the rapid development of modern technology, the application scope of intelligent robots is becoming more and more extensive. Intelligent robots can be seen in homes, shopping malls, banks and other public places.

[0003] However, in the process of executing tasks, the robot in the wear technology can smoothly execute the defined tasks, but cannot smoothly execute complex tasks because they are not defined, which leads to the reduction of the robot's processing performance. Summary of the invention

[0004] Based on this, it is necessary to provide a robot-based task execution method, device, computer equipment, computer-readable storage medium and computer program product that can improve task processing efficiency and ensure accurate execution of tasks in response to the above technical problems.

[0005] In a first aspect, the present application provides a robot-based task execution method, the method comprising:

[0006] When the task type is determined to be a complex task type based on the received speech signal, the complex task is decomposed into subtasks and the corresponding context relationships of the subtasks through the large model;

[0007] Decompose each subtask to obtain the corresponding actions and / or skills for each subtask;

[0008] Generate a task execution graph based on the context of subtasks, the disassembled actions and skills, the context of actions and skills, and the robot's state and posture;

[0009] A task execution target path is generated based on the task execution graph, and the task is executed based on the actions and skills on the target path to obtain a task execution result.

[0010] In one embodiment, generating a task execution graph according to the context of subtasks, the disassembled actions and skills, the context of actions and skills, and the state and posture of the robot includes:

[0011] The disassembled actions and skills are used as graph nodes;

[0012] Determine the first connection line of each of the graph nodes according to the contextual relationship of the subtasks and the decomposed actions and skills;

[0013] Based on the context of actions and skills and the state and position of the robot, a second connection line of each of the graph nodes is determined.

[0014] In one embodiment, determining the second connection line of each of the graph nodes based on the contextual relationship between the action and the skill and the state and position of the robot includes:

[0015] Based on the context of actions and skills, determine the virtual connections of each graph node;

[0016] When the states and positions of the robots of the two graph nodes connected by the virtual line are the same, the virtual line is used as the second line;

[0017] When the states and positions of the robots of two graph nodes connected by the virtual line are different, the virtual line is deleted.

[0018] In one embodiment, generating a task execution target path based on the task execution graph includes:

[0019] Determine the key graph nodes corresponding to each subtask, and determine the start nodes based on the context of the subtasks.

[0020] Taking each starting node as the starting point and traversing all key nodes as the goal, we get the execution path of each task;

[0021] The task execution path that meets the requirements is selected as the task execution target path.

[0022] In one of the embodiments, before determining that the task type is a complex task type based on the received voice signal, the method includes:

[0023] Receive user voice signals, convert the user voice signals into text information, input the text information into the big model, and obtain the user's intention;

[0024] Identify the task type based on the user's intention and obtain the task type;

[0025] When the task type is an action in an action library corresponding to a task or a skill in a skill library corresponding to a task, obtaining an instruction corresponding to the task;

[0026] Control the robot to execute the instructions.

[0027] In one embodiment, the step of identifying the task type based on the user intention to obtain the task type includes:

[0028] Based on a pre-trained classification model, the user intention is identified to obtain a task category;

[0029] The pre-trained classification model is trained based on a small data sample set.

[0030] In a second aspect, the present application further provides a robot-based task execution device, the device comprising:

[0031] A first disassembly module is used to disassemble the complex task into subtasks and the corresponding context relationships of the subtasks by using a large model when the task type is determined to be a complex task type based on the received voice signal;

[0032] The second disassembly module is used to disassemble each subtask to obtain the action and / or skill corresponding to each subtask;

[0033] A task execution graph generation module is used to generate a task execution graph based on the context of subtasks, the disassembled actions and skills, the context of actions and skills, and the state and posture of the robot;

[0034] The task execution module is used to generate a task execution target path based on the task execution graph, and to execute the task based on the actions and skills on the target path to obtain a task execution result.

[0035] In a third aspect, the present application further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method in any one of the above-mentioned embodiments when executing the computer program.

[0036] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method in any one of the above-mentioned embodiments.

[0037] In a fifth aspect, the present application also provides a computer program product, including a computer program, which implements the steps of the method in any one of the above embodiments when executed by a processor.

[0038] The above-mentioned robot-based task execution method, apparatus, computer equipment, computer-readable storage medium and computer program product, when determining that the task type is a complex task type, can decompose the complex task into sub-tasks and the context of sub-tasks, decompose the sub-tasks into actions and / or skills corresponding to each sub-task, and finally generate a task execution graph based on the context of sub-tasks, the decomposed actions and skills, the context of actions and skills, and the state and position of the robot, and then generate the task execution target path based on the task execution graph, so that some nodes can be filtered out, the task processing efficiency can be improved, and the accurate execution of the task can be ensured. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0040] Figure 1 is a flowchart of a robot-based task execution method in one embodiment;

[0041] Figure 2 A flowchart of the steps of generating a task execution graph in one embodiment;

[0042] Figure 3 A flowchart of a task execution target path generation step in one embodiment;

[0043] Figure 4 is a structural block diagram of a robot-based task execution device in one embodiment;

[0044] Figure 5 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0046] In one embodiment, Figure 1 As shown, a robot-based task execution method is provided. This embodiment uses the method applied to a robot terminal as an example. It can be understood that the method can also be applied to a server corresponding to the robot terminal, and can also be applied to a system including a robot terminal and a server, and is implemented through the interaction between the robot terminal and the server. In this embodiment, the method includes the following steps:

[0047] S102: When it is determined based on the received voice signal that the task type is a complex task type, the complex task is decomposed into subtasks and corresponding context relationships of the subtasks using a large model.

[0048] After receiving the voice signal, the robot can determine the task type based on the voice signal, and optionally, the voice signal can be processed based on the large model to obtain the task type. In the present application, the task type can include the action corresponding to the task type in the action library, the skill corresponding to the task type in the skill library, and the complex task type, wherein the task type that is not the action corresponding to the task type in the action library or the skill corresponding to the task type in the skill library is a complex task type.

[0049] When it is determined that the task type is a complex type, the complex task is decomposed by using a large model, so that each subtask and the corresponding front-and-back relationship of the subtasks can be obtained. The subtask is an abstract task. For example, if the complex task is "exchange the water cups on two tables", it can be decomposed into the subtasks of "putting water cup A1 on table 1 on table 2", "putting water cup A2 on table 1 on table 2", "putting water cup B1 on table 2 on table 1", and "putting water cup B2 on table 2 on table 1". In addition, it should be noted that the splitting here is based on the grammar of natural language.

[0050] The context of subtasks is the relationship between subtasks, where a certain subtask must be executed before another subtask, or a certain subtask must be executed after another subtask, or there is no context between two subtasks, etc. In one optional embodiment, the context of subtasks can be represented by a tree structure.

[0051] S104: Decompose each subtask to obtain the action and / or skill corresponding to each subtask.

[0052] In order to realize the execution of the subtask, the subtask needs to be decomposed into actions in the action library and / or skills in the skill library, wherein the action library includes pre-configured basic actions of the robot, such as "take" and "put", etc. The skill library includes encapsulated skills, and each skill can include at least two actions in the action library.

[0053] In this application, since the subtasks are expressed in an abstract natural language, which the robot cannot recognize, each subtask is decomposed into actions in the action library and / or skills in the skill library, thereby obtaining an execution sequence corresponding to the robot. In addition, it should be noted that the decomposition here is based on the actions in the action library and the skills in the skill library.

[0054] For ease of understanding, the subtask of "putting cup A1 on table 1 onto table 2" can be divided into the actions of "picking up cup A1 on table 1", "moving it above table 2" and "putting down cup A1".

[0055] S106: Generate a task execution graph according to the context of the subtasks, the disassembled actions and skills, the context of the actions and skills, and the state and posture of the robot.

[0056] The context of actions and skills is that one action or skill must be performed before another action or skill, or another action or skill must be performed after another action or skill, or there is no context between two actions or skills, etc. In one optional embodiment, the context of actions or skills can be configured as an attribute of an action or an attribute of a skill when generating an action library and a skill library.

[0057] The robot status refers to the robot's working state, including working and non-working. The working state can be subdivided into whether it is holding other objects, etc. The specific configuration needs to be combined with the specific application scenario.

[0058] Posture refers to the position and posture of the robot, including the position and posture of the robot.

[0059] The task execution graph is an execution flow chart of a complex task. In the task execution graph, graph nodes are actions or skills, and the lines between graph nodes are execution relationships. The lines can have directions, that is, a graph node after the line can be executed only after a graph node is executed.

[0060] Based on the contextual relationships of subtasks, the contextual relationships of actions and skills, and the state and posture of the robot, the connections between the various graph nodes can be determined, thereby generating a task execution graph.

[0061] S108: Generate a task execution target path based on the task execution graph, and execute the task based on the actions and skills on the target path to obtain a task execution result.

[0062] The task execution target path is the optimal path for task execution. The optimal path may be a path with the shortest path, or a path with the least number of turns, etc., which is not specifically limited here.

[0063] In this application, the task execution graph can be traversed by the A* algorithm to obtain the optimal path. Then, the robot can call the instructions corresponding to the actions in the action library and the instructions corresponding to the skills in the skill library based on the various graph nodes on the optimal path to execute the corresponding tasks and obtain the task execution results.

[0064] The above-mentioned robot-based task execution method, when determining that the task type is a complex task type, can decompose the complex task to obtain sub-tasks and the context of sub-tasks, decompose the sub-tasks to obtain the actions and / or skills corresponding to each sub-task, and finally generate a task execution graph based on the context of sub-tasks, the decomposed actions and skills, the context of actions and skills, and the state and position of the robot, and then generate the task execution target path based on the task execution graph, so that some nodes can be filtered out, the task processing efficiency can be improved, and the accurate execution of the task can be ensured.

[0065] In one of the optional embodiments, see Figure 2 As shown, Figure 2 The present invention is a flowchart of a task execution graph generation step in an embodiment, wherein the task execution graph generation step generates a task execution graph according to the context of subtasks, disassembled actions and skills, the context of actions and skills, and the state and posture of the robot, and includes: using the disassembled actions and skills as graph nodes; determining the first connection line of each graph node according to the context of subtasks and the disassembled actions and skills; and determining the second connection line of each graph node based on the context of actions and skills and the state and position of the robot.

[0066] This embodiment involves a process of generating a task execution graph, and the task execution graph includes graph nodes and connections between graph nodes.

[0067] The decomposed actions and skills are used as graph nodes. One thing that needs to be explained is that different subtasks may include the same actions and skills. Therefore, different graph nodes are generated for the same actions and skills of different subtasks. The index of each graph node can be the identifier of the subtask, action and skill.

[0068] Since there is a contextual relationship among subtasks, the first connection line of each action and skill can be determined based on the contextual relationship among subtasks, thus completing the initialization of the task execution graph.

[0069] Subsequently, since there is a causal relationship between actions and skills, and each graph node corresponds to the state and position of the robot, the second connection line of each graph node can be determined based on the state and position of the robot and the causal relationship between actions and skills, thereby enriching the task execution graph.

[0070] For the sake of ease of understanding, the above example is still used for explanation. The subtask "Put cup A1 on table 1 on table 2" can be broken down into the actions of "pick up cup A1 on table 1", "move above table 2" and "put down cup A1". Similarly, the subtask of "put cup A2 on table 1 on table 2", the subtask of "put cup B1 on table 2 on table 1" and the subtask of "put cup B2 on table 2 on table 1" can also be broken down into corresponding actions or skills.

[0071] Therefore, various graph nodes can be generated: the action of "picking up cup A1 on table 1", the action of "moving to the top of table 2", and "putting down cup A1", etc., and the first connection of the graph nodes can be established. Since there is no before-and-after relationship between the subtasks here, the starting point of each subtask can be connected to the end point of any subtask.

[0072] The second connection can be determined later based on the robot's state and position and the contextual relationship between actions and skills. For example, the action of "putting down cup A1" must be after the action of "picking up cup A1 on table 1", and the action of "picking up cup B1 on table 2" must be after "moving to the top of table 2" and the robot is not currently holding other items. Therefore, the second connection can be determined based on these contextual relationships.

[0073] In the above embodiment, the connection between the various graph nodes in the task execution graph is determined not only by the context of the subtasks but also by the context of the actions and skills, so that multiple paths for task execution can be given.

[0074] In one of the embodiments, the second connection line of each graph node is determined based on the contextual relationship between actions and skills and the state and position of the robot, including: determining the virtual connection line of each graph node based on the contextual relationship between actions and skills; when the state and position of the robots of two graph nodes connected by the virtual connection line are the same, using the virtual connection line as the second connection line; when the state and position of the robots of two graph nodes connected by the virtual connection line are different, deleting the virtual connection line.

[0075] The virtual lines of each graph node are determined based on the context of actions and skills. These virtual lines do not consider the robot's state and position, which makes some virtual lines inappropriate. Therefore, the virtual lines are filtered based on the robot's state and position, and only the virtual lines with the same robot state and position are retained as the second lines.

[0076] Specifically, in the present application, the virtual lines can be first filtered based on the position of the robot to delete a part of the virtual lines, and then a second filtering can be performed based on the state of the robot. In combination with the above example, the first filtering can be performed based on whether the robot is located above table 2, and then a second filtering can be performed based on whether the robot holds other objects, thereby ensuring the accuracy of the second line.

[0077] In the above embodiment, the connection lines are filtered according to the position and state of the robot, thereby reducing the number of connection lines in the task execution graph and ensuring the accuracy of the task execution graph.

[0078] In one of the optional embodiments, in combination Figure 3 As shown, Figure 3 The present invention is a flowchart of a task execution target path generation step in an embodiment. The task execution target path generation step, i.e., generating a task execution target path based on a task execution graph, includes: determining a key graph node corresponding to each subtask, and determining each starting node based on the context relationship corresponding to the subtask; taking each starting node as a starting point and traversing each key node as a goal to obtain each task execution path; and selecting a task execution path that meets the requirements as the task execution target path.

[0079] The key graph nodes may include the end graph node of each subtask and the start node determined based on the subtask. The start node is the first node of each subtask determined after considering the context of the subtasks. If a subtask does not need to execute any other subtasks before, the first node of the subtask can be the start node, and the obtained start node is stored in the start node set. The key graph node is the end graph node of each subtask, and the end graph nodes of these subtasks can be stored in the key graph node set as key graph nodes.

[0080] Then, each task execution path is generated with the start node as the starting point and each end graph node in the key graph node as the goal. In this application, only each end graph node in the key graph node is traversed as the goal, because the end graph node must rely on the complete execution of the subtask to end. For example, before the action of "putting down water cup A1" in the above text, there must be an action of "picking up water cup A1 on table 1".

[0081] The requirements to be met may include the shortest path, the least number of turns, etc. In this application, corresponding requirements can be set based on scenario needs, and then the task execution path can be screened based on the requirements to obtain the task execution target path.

[0082] Still taking the above example of picking up cups as an example, if path optimization is not performed, the final task execution plan is the subtask of "putting cup A1 on table 1 onto table 2", the subtask of "putting cup A2 on table 1 onto table 2", the subtask of "putting cup B1 on table 2 onto table 1", and the subtask of "putting cup B2 on table 2 onto table 1"; in this application, optimization is performed, and the task execution becomes the subtask of "putting cup A1 on table 1 onto table 2", the subtask of "putting cup B2 on table 2 onto table 1", the subtask of "putting cup A2 on table 1 onto table 2", and the subtask of "putting cup B1 on table 2 onto table 1", thereby achieving path optimization.

[0083] In the above embodiment, each starting point and key graph node are first determined, and then each task execution path is generated with the start node as the starting point and with the goal of traversing each end graph node in the key graph node. The paths corresponding to the task execution can be obtained, laying the foundation for subsequent path screening.

[0084] In one of the optional embodiments, before determining that the task type is a complex task type based on the received voice signal, it includes: receiving a user voice signal, converting the user voice signal into text information, inputting the text information into a large model, and obtaining the user intention; identifying the task type based on the user intention to obtain the task type; when the task type is an action in the action library corresponding to a task or a skill in the skill library corresponding to a task, obtaining instructions corresponding to the task; and controlling the robot to execute the instructions.

[0085] The voice signal is input by the user, and after the robot collects the user's voice signal, it can convert the voice signal into text information. Specifically, the text conversion can be performed through voice recognition.

[0086] The big model is pre-trained to obtain user intent. The big model can process text information to obtain user intent, where the user intent can be expressed in natural language.

[0087] Task types include action-corresponding task types in the action library, skill-corresponding task types in the skill library, and complex task types, among which task types that are not action-corresponding task types in the action library or skill-corresponding task types in the skill library are all complex task types.

[0088] Among them, the actions in the action library corresponding to the tasks or the skills in the skill library corresponding to the tasks can be executed directly, for example, directly obtaining the robot executable instructions corresponding to the actions in the action library or the robot executable instructions corresponding to the skills in the skill library, and then controlling the robot to execute these executable instructions.

[0089] In one of the optional embodiments, task type identification is performed based on user intent to obtain the task type, including: based on a pre-trained classification large model, user intent is identified to obtain the task category; the pre-trained classification large model is trained based on a small data sample set.

[0090] The identification of task types can also be performed through a large model, which can be trained based on a small sample data set, thereby improving the accuracy of training while ensuring the training speed. Subsequently, the user's intention is input into the large model to obtain the corresponding task category.

[0091] In each of the above embodiments, user intention recognition is first performed based on a large model, and the task type (actions in the action library correspond to tasks, skills in the skill library correspond to tasks, and complex tasks) is determined. For the first two types of tasks, the action library and the skill library can be directly called for execution, and for complex tasks, the tasks are split. Secondly, the execution of complex tasks includes: first, the complex tasks are disassembled into subtasks through a large model, and then the subtasks are disassembled to obtain the actions and / or skills corresponding to each subtask. A task execution graph is generated based on the tree structure relationship of the subtasks, the disassembled actions and skills, the action attributes of the action library and the skill attributes of the skill library, and the state and position of the robot. The key nodes (end nodes) corresponding to each subtask are identified, and each initial node is determined based on the tree structure relationship of the subtask. The initial node is used as the starting point, and the goal is to traverse all the key nodes. The optimal path for task execution is generated based on the task execution graph, and the task is executed based on the optimal path. Thus, the execution of complex tasks can be achieved in the above manner, and the accuracy of task execution is guaranteed.

[0092] It should be understood that, although the steps in the flowcharts involved in the above embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0093] Based on the same inventive concept, the embodiment of the present application also provides a robot-based task execution device for implementing the robot-based task execution method involved above. The implementation solution provided by the device to solve the problem is similar to the implementation solution recorded in the above method, so the specific limitations in one or more robot-based task execution device embodiments provided below can refer to the limitations of the robot-based task execution method above, and will not be repeated here.

[0094] In an exemplary embodiment, Figure 4 As shown, a robot-based task execution device is provided, comprising: a first disassembly module 401, a second disassembly module 402, a task execution graph generation module 403 and a task execution module 404, wherein:

[0095] A first decomposition module 401 is used to decompose the complex task into subtasks and the corresponding context relationships of the subtasks by using a large model when the task type is determined to be a complex task type based on the received speech signal;

[0096] The second decomposition module 402 is used to decompose each subtask to obtain the action and / or skill corresponding to each subtask;

[0097] A task execution graph generation module 403 is used to generate a task execution graph according to the context of subtasks, the decomposed actions and skills, the context of actions and skills, and the state and posture of the robot;

[0098] The task execution module 404 is used to generate a task execution target path based on the task execution graph, and execute the task based on the actions and skills on the target path to obtain a task execution result.

[0099] In one of the optional embodiments, the above-mentioned task execution graph generation module 403 is specifically used to use the disassembled actions and skills as graph nodes; determine the first connection line of each graph node based on the contextual relationship of the subtasks and the disassembled actions and skills; determine the second connection line of each graph node based on the contextual relationship of the actions and skills and the state and position of the robot.

[0100] In one of the optional embodiments, the above-mentioned task execution graph generation module 403 is specifically used to determine the virtual connection of each graph node based on the contextual relationship between actions and skills; when the states and positions of the robots of the two graph nodes connected by the virtual connection are the same, the virtual connection is used as the second connection; when the states and positions of the robots of the two graph nodes connected by the virtual connection are different, the virtual connection is deleted.

[0101] In one of the optional embodiments, the above-mentioned task execution module 404 is specifically used to determine the key graph node corresponding to each subtask, and determine each starting node based on the corresponding front-end relationship of the subtask; take each starting node as the starting point, and traverse all key nodes as the goal to obtain each task execution path; select the task execution path that meets the requirements as the task execution target path.

[0102] In one of the optional embodiments, the above-mentioned device also includes: an instruction acquisition module, used to receive user voice signals, convert the user voice signals into text information, input the text information into the large model, and obtain the user intention; identify the task type based on the user intention to obtain the task type; when the task type is an action in the action library corresponding to the task or a skill in the skill library corresponds to the task, obtain the instruction corresponding to the task; control the robot to execute the instruction.

[0103] In one of the optional embodiments, the above-mentioned instruction acquisition module is also used to identify the user intention based on a pre-trained classification model to obtain the task category; the pre-trained classification model is trained based on a small data sample set.

[0104] Each module in the above robot-based task execution device can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each module.

[0105] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 5As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC) or other technologies. When the computer program is executed by the processor, a robot-based task execution method is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse.

[0106] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0107] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the following steps when executing the computer program: when it is determined that the task type is a complex task type based on a received voice signal, the complex task is decomposed into subtasks and the corresponding context relationships of the subtasks through a large model; each subtask is decomposed to obtain the actions and / or skills corresponding to each subtask; a task execution graph is generated according to the context relationships of the subtasks, the decomposed actions and skills, the context relationships between the actions and skills, and the state and posture of the robot; a task execution target path is generated based on the task execution graph, and the task is executed based on the actions and skills on the target path to obtain a task execution result.

[0108] In one embodiment, when a processor executes a computer program, a task execution graph is generated based on the context of subtasks, the disassembled actions and skills, the context of actions and skills, and the state and posture of the robot, including: using the disassembled actions and skills as graph nodes; determining the first connection line of each graph node based on the context of subtasks and the disassembled actions and skills; determining the second connection line of each graph node based on the context of actions and skills and the state and position of the robot.

[0109] In one embodiment, the processor determines the second line of each graph node based on the contextual relationship of actions and skills and the state and position of the robot when executing a computer program, including: determining the virtual line of each graph node based on the contextual relationship of actions and skills; when the states and positions of the robots of two graph nodes connected by the virtual line are the same, using the virtual line as the second line; when the states and positions of the robots of two graph nodes connected by the virtual line are different, deleting the virtual line.

[0110] In one embodiment, the generation of a task execution target path based on a task execution graph implemented by a processor when executing a computer program includes: determining a key graph node corresponding to each subtask, and determining each start node based on the contextual relationship corresponding to the subtask; taking each start node as a starting point and traversing each key node as a goal to obtain each task execution path; and selecting a task execution path that meets the requirements as the task execution target path.

[0111] In one embodiment, the steps implemented when the processor executes a computer program include: receiving a user voice signal, converting the user voice signal into text information, inputting the text information into a large model, and obtaining the user intention; identifying the task type based on the user intention to obtain the task type; obtaining instructions corresponding to the task when the task type is an action in the action library corresponding to a task or a skill in the skill library corresponding to a task; and controlling the robot to execute the instructions.

[0112] In one embodiment, the task type recognition based on user intent implemented by the processor when executing a computer program to obtain the task type includes: based on a pre-trained classification large model, identifying the user intent to obtain the task category; the pre-trained classification large model is trained based on a small data sample set.

[0113] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: when the task type is determined to be a complex task type based on a received voice signal, the complex task is decomposed into subtasks and the corresponding context relationships of the subtasks through a large model; each subtask is decomposed to obtain the actions and / or skills corresponding to each subtask; a task execution graph is generated according to the context relationships of the subtasks, the decomposed actions and skills, the context relationships between the actions and skills, and the state and posture of the robot; a task execution target path is generated based on the task execution graph, and the task is executed based on the actions and skills on the target path to obtain a task execution result.

[0114] In one embodiment, the computer program implemented when executed by the processor generates a task execution graph based on the context of subtasks, the disassembled actions and skills, the context of actions and skills, and the state and posture of the robot, including: using the disassembled actions and skills as graph nodes; determining the first connection line of each graph node based on the context of subtasks and the disassembled actions and skills; determining the second connection line of each graph node based on the context of actions and skills and the state and position of the robot.

[0115] In one embodiment, the computer program implemented when executed by the processor determines the second connection line of each graph node based on the contextual relationship of actions and skills and the state and position of the robot, including: determining the virtual connection line of each graph node based on the contextual relationship of actions and skills; when the state and position of the robots of two graph nodes connected by the virtual connection are the same, using the virtual connection line as the second connection line; when the state and position of the robots of two graph nodes connected by the virtual connection are different, deleting the virtual connection line.

[0116] In one embodiment, the computer program generated a task execution target path based on a task execution graph when executed by a processor, including: determining the key graph node corresponding to each subtask, and determining each start node based on the context relationship corresponding to the subtask; taking each start node as the starting point and traversing each key node as the goal to obtain each task execution path; selecting the task execution path that meets the requirements as the task execution target path.

[0117] In one embodiment, the computer program implemented when executed by the processor includes: receiving a user voice signal, converting the user voice signal into text information, inputting the text information into a large model, and obtaining the user intention; identifying the task type based on the user intention to obtain the task type; when the task type is an action in the action library corresponding to the task or a skill in the skill library corresponding to the task, obtaining the instructions corresponding to the task; and controlling the robot to execute the instructions.

[0118] In one embodiment, when a computer program is executed by a processor, task type recognition based on user intent is implemented to obtain the task type, including: based on a pre-trained classification model, user intent is recognized to obtain the task category; the pre-trained classification model is trained based on a small data sample set.

[0119] In one embodiment, a computer program product is provided, including a computer program, which implements the following steps when executed by a processor: when the task type is determined to be a complex task type based on a received voice signal, the complex task is decomposed into subtasks and the corresponding context relationships of the subtasks through a large model; each subtask is decomposed to obtain the actions and / or skills corresponding to each subtask; a task execution graph is generated according to the context relationships of the subtasks, the decomposed actions and skills, the context relationships between the actions and skills, and the state and posture of the robot; a task execution target path is generated based on the task execution graph, and the task is executed based on the actions and skills on the target path to obtain a task execution result.

[0120] In one embodiment, the computer program implemented when executed by the processor generates a task execution graph based on the context of subtasks, the disassembled actions and skills, the context of actions and skills, and the state and posture of the robot, including: using the disassembled actions and skills as graph nodes; determining the first connection line of each graph node based on the context of subtasks and the disassembled actions and skills; determining the second connection line of each graph node based on the context of actions and skills and the state and position of the robot.

[0121] In one embodiment, the computer program implemented when executed by the processor determines the second connection line of each graph node based on the contextual relationship of actions and skills and the state and position of the robot, including: determining the virtual connection line of each graph node based on the contextual relationship of actions and skills; when the state and position of the robots of two graph nodes connected by the virtual connection are the same, using the virtual connection line as the second connection line; when the state and position of the robots of two graph nodes connected by the virtual connection are different, deleting the virtual connection line.

[0122] In one embodiment, the computer program generated a task execution target path based on a task execution graph when executed by a processor, including: determining the key graph node corresponding to each subtask, and determining each start node based on the context relationship corresponding to the subtask; taking each start node as the starting point and traversing each key node as the goal to obtain each task execution path; selecting the task execution path that meets the requirements as the task execution target path.

[0123] In one embodiment, the computer program implemented when executed by the processor includes: receiving a user voice signal, converting the user voice signal into text information, inputting the text information into a large model, and obtaining the user intention; identifying the task type based on the user intention to obtain the task type; when the task type is an action in the action library corresponding to the task or a skill in the skill library corresponding to the task, obtaining the instructions corresponding to the task; and controlling the robot to execute the instructions.

[0124] In one embodiment, when a computer program is executed by a processor, task type recognition based on user intent is implemented to obtain the task type, including: based on a pre-trained classification model, user intent is recognized to obtain the task category; the pre-trained classification model is trained based on a small data sample set.

[0125] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0126] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.

[0127] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0128] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A robot-based task execution method, characterized in that: The method comprises: When the task type is determined to be a complex task type based on the received speech signal, the complex task is decomposed into subtasks and the corresponding context relationships of the subtasks through the large model; Decompose each subtask to obtain the corresponding actions and / or skills for each subtask; A task execution graph is generated according to the contextual relationship of subtasks, the decomposed actions and skills, the contextual relationship between actions and skills, and the state and posture of the robot; the task execution graph includes graph nodes and lines between the graph nodes; the graph nodes are actions or skills, the lines between the graph nodes are execution relationships, and the lines have directions, and the directions are used to indicate that a graph node after the line can be executed only after a graph node is executed; Generate a task execution target path based on the task execution graph, and execute the task based on the actions and skills on the target path to obtain a task execution result; The generating of the task execution graph according to the context of the subtasks, the disassembled actions and skills, the context of the actions and skills, and the state and posture of the robot includes: The disassembled actions and skills are used as graph nodes; Determine the first connection line of each of the graph nodes according to the contextual relationship of the subtasks and the decomposed actions and skills; Based on the context of actions and skills and the state and position of the robot, a second connection line of each of the graph nodes is determined.

2. The method according to claim 1, characterized in that The determining of the second connection line of each of the graph nodes based on the contextual relationship between the action and the skill and the state and position of the robot comprises: Based on the context of actions and skills, determine the virtual connections of each graph node; When the states and positions of the robots of the two graph nodes connected by the virtual line are the same, the virtual line is used as the second line; When the states and positions of the robots of two graph nodes connected by the virtual line are different, the virtual line is deleted.

3. The method according to claim 1, characterized in that The generating a task execution target path based on the task execution graph includes: Determine the key graph nodes corresponding to each subtask, and determine the start nodes based on the context of the subtasks. Taking each starting node as the starting point and traversing all key nodes as the goal, we get the execution path of each task; The task execution path that meets the requirements is selected as the task execution target path.

4. The method according to any one of claims 1 to 3, characterized in that: Before determining that the task type is a complex task type based on the received voice signal, the method includes: Receive user voice signals, convert the user voice signals into text information, input the text information into the big model, and obtain the user's intention; Identify the task type based on the user's intention and obtain the task type; When the task type is an action in an action library corresponding to a task or a skill in a skill library corresponding to a task, obtaining an instruction corresponding to the task; Control the robot to execute the instructions.

5. The method according to claim 4, characterized in that The task type is identified based on the user intention to obtain the task type, including: Based on a pre-trained classification model, the user intention is identified to obtain a task category; The pre-trained classification model is trained based on a small data sample set.

6. A robot-based task execution device, characterized in that: The device comprises: A first disassembly module is used to disassemble the complex task into subtasks and the corresponding context relationships of the subtasks by using a large model when the task type is determined to be a complex task type based on the received voice signal; The second disassembly module is used to disassemble each subtask to obtain the action and / or skill corresponding to each subtask; A task execution graph generation module is used to generate a task execution graph according to the context of subtasks, the decomposed actions and skills, the context of actions and skills, and the state and posture of the robot; the task execution graph includes graph nodes and lines between the graph nodes; the graph nodes are actions or skills, the lines between the graph nodes are execution relationships, and the lines have directions, which are used to indicate that a graph node after the line can be executed only after one graph node is executed; A task execution module, used to generate a task execution target path based on the task execution graph, and to execute the task based on the actions and skills on the target path to obtain a task execution result; The task execution graph generation module is specifically used to use the disassembled actions and skills as graph nodes; determine the first connection line of each graph node according to the contextual relationship of the subtasks and the disassembled actions and skills; and determine the second connection line of each graph node based on the contextual relationship of the actions and skills and the state and position of the robot.

7. The device according to claim 6, characterized in that The task execution graph generation module is specifically used to determine the key graph nodes corresponding to each subtask, and determine each starting node based on the context relationship corresponding to the subtask; taking each starting node as the starting point and traversing each key node as the goal, obtain each task execution path; The task execution path that meets the requirements is selected as the task execution target path.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Robot action sequence generation method and device

    CN110297697A

  • Multi-stage task processing method and device based on intelligent Agent model and medium

    CN118656196A

  • Man-machine interaction assembly method and system based on multi-modal large model and reinforcement learning

    CN118744426A

  • Task processing method, device and equipment based on large model agent arrangement, storage medium and program product

    CN118819778A

  • Robot control method, device and equipment, robot, medium and product

    CN119610098A