An interaction method and device for collaborative control of the upper limbs of a humanoid robot

Through the large language model combined with the interaction method of visual and voice modules, the problem of humanoid robots lacking a perception system is solved, real-time perception and task planning of the external environment are realized, and the operation ability and human-computer interaction effect of complex tasks are improved.

CN119927906BActive Publication Date: 2025-08-01江淮前沿技术协同创新中心
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510103239.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-08-01
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

Existing humanoid robots lack perception systems and cannot perform real-time perception and timely adjustments to the external environment, resulting in insufficient operation capabilities of complex tasks.

Method used

A large language model is used to combine visual and voice modules to realize task analysis, target recognition, motion planning and operation control, and real-time adjustments are made through voice interactive feedback.

Benefits of technology

It improves the operation ability of humanoid robots to perform complex tasks on the upper limbs, realizes multimodal information perception fusion and voice interaction feedback, and enhances the ease of use and emotional connection of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119927906B_ABST
    Figure CN119927906B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides an interaction method for collaborative control of the upper limbs of a humanoid robot, including: performing task analysis on the received voice request; if there is a target object to be executed, performing target object recognition processing according to the task analysis result; based on the target recognition information and target pose information output by the recognition processing, as well as the task analysis result, using a large language model to perform task planning processing according to the principle of dual-arm division of labor and cooperation, generating a task sequence and a text instruction corresponding to each task; performing the following operations on each target task in the task sequence in turn: based on the target pose information and target recognition information, controlling the upper limbs of the humanoid robot to perform an operation action corresponding to the target task on the target object; thereby, this embodiment is based on a large language model, in cooperation with a vision module and a voice module, not only enabling users to communicate with the robot more intuitively; but also improving the operation ability of the upper limbs of the humanoid robot to perform complex tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of humanoid robots, and particularly relates to an interaction method and device for upper limb collaborative control of humanoid robots. Background Art

[0002] The robotic arm is an important component structure of a humanoid robot and is the direct execution structure for the robot to complete tasks. A dual-arm collaborative robot can have two arms like a human and perform coordinated work within a certain range. This design enables the robot to greatly enhance its adaptability to complex assembly tasks and improve the utilization efficiency of the working space. However, the upper limb collaboration of a humanoid robot is not a simple combination of two robotic arms + grippers, but a multi-structure combination of dual robotic arms + dexterous hands. In addition to the control objectives to be achieved by each arm, it is also necessary to satisfy the coordinated control between the two arms, between the arm and the hand, as well as the adaptability to the environment, which greatly increases the complexity of the upper limb operation of the humanoid robot and requires more advanced integrated systems, efficient collaborative control algorithms, accurate environmental perception and positioning, and adjustable control methods, etc. In addition, a humanoid robot not only needs basic limb movement capabilities, but also should have a "smart brain" capable of tasks such as instruction disassembly, task planning, and human-machine interaction to reflect the superiority of the humanoid robot over other conventional industrial robots in terms of intelligence and autonomy. Therefore, based on the language understanding ability of large models and various perception information such as speech and vision, combined with the arm-hand joint system, it is possible to achieve real-time perception of the external environment by the robot, interaction, and real-time adjustment of motion planning.

[0003] Currently, the coordinated control of dual-arm robots is usually based on kinematics and dynamics control. The coordinated control based on kinematics focuses on the research of the redundant characteristics of the robotic arm, and the coordinated control based on dynamics focuses on the research of the end force control of the robotic arm. However, due to the significant differences in the redundant degrees of freedom of the dual-arm system and the single-arm kinematics in the solution of the inverse kinematics, the different dual-arm configurations also determine the differences in the inverse solution methods, and the uncertainty of the inverse solution also affects the establishment method of the dynamics model, the changes in the movement speed and acceleration of the dual arms, and the differences in the end relative forces. In an unstructured environment, the dual-arm robot system cannot obtain all the information of kinematics and dynamics. How the robot perceives the external environment, including parameters such as the position, speed, and contact force of the operation object, and how the external environment information participates in the coordinated control of the dual arms has become a new research hotspot.

[0004] The Chinese patent specification CN 118143954 A discloses a compliant control method and device for the upper limb dual robotic arms of a humanoid robot. A closed-chain kinematic model is constructed based on the closed motion chain formed by the upper limb dual robotic arms and the operating object, and a joint dynamics model is constructed for the upper limb dual robotic arms. A dual-arm collaborative controller for the dual robotic arms is set up to achieve dual-arm collaborative motion control. However, this method is completely based on traditional strategies and cannot achieve precise environmental perception and real-time self-adjustment, so the operation of complex tasks cannot be achieved in practical applications.

[0005] The Chinese patent specification CN 117697763 A discloses a learning method and system for dual-arm operation tasks based on a large model. The large model receives a task environment map and task description information for dual-arm operation, determines a reward function, and obtains a dual-arm operation strategy through an imitation learning algorithm. This method designs the reinforcement learning reward function using the large model. However, the setting and improvement of the reward function are of a certain degree of difficulty. Usually, the robot hardware can only demonstrate one skill at a time. At the same time, this method focuses on dual-arm operation and does not emphasize multi-modal perception and interaction, which is essentially different from the upper limb operation of a humanoid robot in practical applications. Summary of the Invention

[0006] In view of the above problems existing in the prior art, an embodiment of the present invention provides an interaction method and device for the collaborative control of the upper limbs of a humanoid robot; this method can solve the problem that the existing humanoid robot cannot perform real-time perception of the external environment and timely adjustment due to the lack of a perception system; and further improves the operation ability of the upper limbs of the humanoid robot to execute complex tasks.

[0007] According to the first aspect of the embodiment of the present invention, an interaction method for the collaborative control of the upper limbs of a humanoid robot is provided. The method includes: using a large language model to perform task analysis on the text information corresponding to the received voice request and output a task analysis result; if the task analysis result indicates that there is a target object to be executed, perform target object recognition processing according to the task analysis result and output target recognition information and target pose information; based on the task analysis result, target recognition information, and the target pose information, use the large language model to perform task planning processing according to the dual-arm division of labor and cooperation principle, generate a task sequence and a text instruction corresponding to each task; for any target task in the task sequence: broadcast the task status in voice form for the text instruction corresponding to the target task; when receiving the feedback voice from the user for the task status as executing the task, control the upper limbs of the humanoid robot to perform an operation action corresponding to the feedback voice for the target object based on the target pose information and target recognition information corresponding to the target task; according to the time sequence of the task sequence, control the upper limbs of the humanoid robot to perform corresponding operation actions for each target task in turn to generate an interaction result.

[0008] Optionally, the method further includes: if the task analysis result indicates that there is no target object to be executed, using the large language model to perform task planning processing on the task analysis result to generate a text instruction corresponding to the task analysis result; broadcasting the task status in the form of voice for the text instruction corresponding to the task analysis result; when the feedback voice instruction received from the user for the task status is to modify the task, using the feedback voice as a voice request and continuing to use the large language model to perform task analysis on the text information corresponding to the voice request.

[0009] Optionally, controlling the upper limb of the humanoid robot to perform an operation action corresponding to the feedback voice for the target object based on the target pose information and target recognition information corresponding to the target task includes: generating a robotic arm motion trajectory based on the target pose information corresponding to the target task; generating a grasping gesture based on the target recognition information and target pose information corresponding to the target task; in response to the execution instruction corresponding to the feedback voice, controlling the robotic arm of the humanoid robot to move to a preset position according to the robotic arm motion trajectory to generate a two-handed execution instruction; in response to the two-handed execution instruction, controlling the two hands of the humanoid robot to perform a grasping operation on the target object according to the grasping gesture.

[0010] Optionally, generating a robotic arm motion trajectory based on the target pose information corresponding to the target task includes: converting the target pose information corresponding to the target task into the end pose of the robotic arm; using the robotic arm motion model to perform motion analysis on the end pose of the robotic arm according to the orocos KDL library; performing motion trajectory planning on the motion analysis result based on DMP to generate a robotic arm motion trajectory.

[0011] Optionally, generating a grasping gesture based on the target recognition information and target pose information corresponding to the target task includes: generating a 3D point cloud model of the target object and a target object contact point based on the target pose information and the target recognition information; performing feature extraction processing on the 3D point cloud model of the target object and the target object contact point to generate a contact map; matching the current finger positions of the two hands of the humanoid robot with the contact map to generate a grasping gesture.

[0012] Optionally, using the large language model to perform task analysis on the text information corresponding to the received voice request and output a task analysis result; converting the received voice request into text information: performing feature extraction on the text information to generate key information; performing prompt word matching on the key information to generate a matching result; performing analysis and reasoning based on the matching result to output a task analysis result.

[0013] Optionally, performing object recognition processing according to the task analysis result, and outputting object recognition information and object pose information, including: obtaining key information of the object from the task analysis result; obtaining object recognition information corresponding to the key information; performing object segmentation based on the object recognition information to generate an object segmentation result; performing pose detection of the object based on the object segmentation result to generate object pose information; wherein the object pose information includes object position information and object attitude information.

[0014] Optionally, in chronological order of the task sequence, successively controlling the upper limbs of the humanoid robot to execute corresponding operation actions according to each target task, and generating an interaction result, including: in chronological order of the task sequence, successively controlling the upper limbs of the humanoid robot to execute corresponding operation actions according to each target task to generate a completion instruction; converting the completion instruction into voice information, and feeding back the task status corresponding to the voice information to the user.

[0015] According to a second aspect of an embodiment of the present invention, there is also provided an interaction device for collaborative control of the upper limbs of a humanoid robot. The device includes: a first task analysis module, configured to perform task analysis on the text information corresponding to the received voice request by using the large language model, and output a task analysis result; a first instruction disassembling module, configured to, if the task analysis result indicates that there is an object to be executed, perform object recognition processing according to the task analysis result, and output object recognition information and object pose information; performing task planning processing by using the large language model according to the principle of dual-arm division of labor and cooperation based on the task analysis result, the object recognition information, and the object pose information, to generate a task sequence and a text instruction corresponding to each task; a control module, configured to, for any target task in the task sequence: broadcasting the text instruction corresponding to the target task in a voice form as a task status; when receiving the feedback voice of the user for the task status as executing the task, controlling the upper limbs of the humanoid robot to execute the operation action corresponding to the feedback voice for the object based on the object pose information and the object recognition information corresponding to the target task; a first generating module, configured to, in chronological order of the task sequence, successively control the upper limbs of the humanoid robot to execute corresponding operation actions according to each target task, and generate an interaction result.

[0016] According to a third aspect of an embodiment of the present invention, there is also provided an electronic device, which includes: a processor; a memory for storing executable instructions of the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the method as described in the first aspect.

[0017] According to a fourth aspect of the embodiments of the present invention, there is also provided a computer-readable medium having a computer program stored thereon, and when the program is executed by a processor, the method described in the first aspect is implemented.

[0018] The embodiments of the present invention provide an interaction method and device for upper limb collaborative control of a humanoid robot. The method includes: First, using a large language model to perform task analysis on the text information corresponding to the received voice request, and outputting a task analysis result; Second, if the task analysis result indicates that there is a target object to be executed, then perform target object recognition processing according to the task analysis result, and output target recognition information and target pose information; Based on the task analysis result, target recognition information, and the target pose information, use the large language model to perform task planning processing according to the principle of dual-arm division of labor and cooperation, generate a task sequence and a text instruction corresponding to each task; After that, for any target task in the task sequence: broadcast the task status of the text instruction corresponding to the target task in a voice form; When receiving the feedback voice from the user for the task status as executing the task, then based on the target pose information and target recognition information corresponding to the target task, control the upper limbs of the humanoid robot to perform an operation action corresponding to the feedback voice for the target object; Finally, in the time sequence of the task sequence, control the upper limbs of the humanoid robot to perform corresponding operation actions according to each target task in turn, and generate an interaction result. Thus, the method of this embodiment is based on a large language model, cooperating with a vision module and a voice module, and can not only achieve task decomposition and reasoning, multi-modal information perception and fusion, and voice interaction feedback, but also solve the problem that existing humanoid robots cannot perform real-time perception of the external environment and make timely adjustments due to the lack of a perception system; further improving the operation ability of the upper limbs of the humanoid robot to perform complex tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Some specific embodiments of the present invention will be described in detail hereinafter with reference to the accompanying drawings in an exemplary and non-limiting manner. The same reference numerals in the drawings denote the same or similar components or parts. Those skilled in the art should understand that these drawings are not necessarily drawn to scale. In the drawings:

[0020] Figure 1 is a schematic diagram of a collaborative motion control system for the upper limbs of a humanoid robot provided by an embodiment of the present invention;

[0021] Figure 2 is a schematic flowchart of an interaction method for upper limb collaborative control of a humanoid robot provided by an embodiment of the present invention;

[0022] Figure 3 is a schematic diagram of the upper limb grasping of a humanoid robot provided by an embodiment of the present invention;

[0023] Figure 4 Schematic diagram of the grasping gesture of the dexterous hand in an embodiment of the present invention;

[0024] Figure 5 Schematic diagram of the structure of a control device for the coordinated movement of the upper limbs of a humanoid robot provided in an embodiment of the present invention. Detailed implementation manners

[0025] To make the objectives, features, and advantages of the present invention more obvious and understandable, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0026] As Figure 1 shown, it is a schematic diagram of a coordinated movement control system for the upper limbs of a humanoid robot provided in an embodiment of the present invention.

[0027] A coordinated movement control system for the upper limbs of a humanoid robot, the system 100 includes: an embodied hierarchical decision-making interaction model, a robotic arm module, and a dexterous hand module; the embodied hierarchical decision-making interaction model includes: a large language model, a vision module, and a voice module. The voice module includes a voice recognition unit, a text conversion unit, and a voice synthesis unit. The vision module includes an object recognition unit, an object segmentation unit, and an object detection unit. The large language model includes an instruction disassembling unit, an instruction reasoning unit, and a task decision-making unit. The robotic arm module includes: a kinematic analysis unit and a motion trajectory planning unit. The dexterous hand module includes: an object model generation unit and a grasping gesture unit.

[0028] Both the first information output by the voice module and the second information output by the vision module are input into the large language model for processing, and a task sequence and the corresponding text instructions for each task are output. For any target task in the task sequence: the voice module processes the text instructions and broadcasts the current task status in voice form; when the system receives a feedback instruction from the user for the current task status to execute the task, it notifies the robotic arm module to control the robotic arm of the humanoid robot to move to a preset position according to the robotic arm motion trajectory; then it notifies the dexterous hand module to control the two hands of the humanoid robot to perform a grasping operation on the target object according to the grasping gesture.

[0029] The present invention proposes an interaction system for collaborative control of the upper limbs of a humanoid robot including "arm - hand - brain". Among them, after the embodied hierarchical decision - making interaction model decomposes the instructions, it realizes task planning, decision - making, and allocates them to the two arms and dexterous hands of the humanoid robot. The robotic arm conducts motion analysis and trajectory planning according to the corresponding content, and the dexterous hands generate grasping gestures according to the target model, and finally complete the operation task through collaborative motion.

[0030] As Figure 2 shown, it is a schematic flowchart of an interaction method for collaborative control of the upper limbs of a humanoid robot provided by an embodiment of the present invention; as Figure 4 shown, it is a schematic diagram of the grasping gesture of the dexterous hand in an embodiment of the present invention.

[0031] An interaction method for collaborative control of the upper limbs of a humanoid robot at least includes the following steps:

[0032] S101, using a large - language model to perform task analysis on the text information corresponding to the received voice request, and outputting a task analysis result;

[0033] S102, if the task analysis result indicates that there is a target object to be executed, then perform target object recognition processing according to the task analysis result, and output target recognition information and target pose information; based on the task analysis result, target recognition information, and target pose information, use the large - language model to perform task planning processing according to the principle of division of labor and cooperation of the two arms, generate a task sequence and a text instruction corresponding to each task;

[0034] S103, for any target task in the task sequence: broadcast the task status in voice form for the text instruction corresponding to the target task; when receiving the feedback voice from the user for the task status as "execute the task", then based on the target pose information and target recognition information corresponding to the target task, control the upper limbs of the humanoid robot to perform the operation action corresponding to the feedback voice for the target object;

[0035] S104, in the time sequence of the task sequence, successively control the upper limbs of the humanoid robot to perform the corresponding operation actions according to each target task, and generate an interaction result.

[0036] In S101, the specific process of performing task analysis on the large - language model in this embodiment is not limited in any way, as long as task planning can be achieved.

[0037] Exemplarily, the large language model is used to perform task analysis on the text information corresponding to the received voice request and output a task analysis result, including: converting the received voice request into text information; extracting features from the text information to generate key information; performing prompt word matching on the key information to generate a matching result; and performing analysis and reasoning based on the matching result to output a task analysis result. There are two situations for the task analysis result output after task analysis; the first situation is that there is a target object to be executed; the second situation is that there is no target object to be executed.

[0038] For example: the first situation: Please help me get an apple and put it in my hand.

[0039] The second situation includes: there is a target object, or there is a target object but no execution is required; for example: I'm hungry and I want to eat; or, there is an apple on the table.

[0040] In S102, the specific process and the applied algorithm for target object recognition and processing in this embodiment are not limited in any way, as long as the target recognition information and target pose information corresponding to the target object can be obtained.

[0041] Exemplarily, the target object is recognized and processed according to the task analysis result, and the target recognition information and target pose information are output, including: obtaining the key information of the target object from the task analysis result; obtaining the target recognition information corresponding to the key information; performing target segmentation based on the target recognition information to generate a target segmentation result; and performing pose detection on the target object based on the target segmentation result to generate target pose information, where the target pose information includes: target object position information and target object attitude information.

[0042] Specifically, a visual module is used to perform target object recognition and processing, and the recognition result is matched with the key information of the target object; if the match is successful, target recognition information is generated, where the target recognition information includes: target point cloud data, target image data, target type, size and other information; then, target segmentation processing is performed based on the target recognition information to generate a target segmentation result; pose detection is performed on the target segmentation result to obtain target pose information. Then, the target pose information and target recognition information are respectively input into the large language model; according to the principle of dual-arm division of labor and cooperation, the large language model disassembles the task instructions based on the task analysis result, target pose information and target recognition information, and outputs a task disassembly result; the task disassembly result is used for task instruction reasoning to output a task sequence and the text instruction corresponding to each task.

[0043] In S103, the speech synthesis unit converts the text instruction corresponding to the target task into voice information; and announces the task status to the user. Then, it determines whether to execute the target task through voice interaction; if the determination result indicates to execute the target task, based on the target pose information and target recognition information corresponding to the target task, it controls the upper limb of the humanoid robot to perform an operation action on the target object; if the determination result indicates not to execute the target task and to correct the target task, it takes the user's corrected voice regarding the task status as a new voice request; and continues to use the large language model to perform task analysis on the voice request, that is, returns to step S101.

[0044] Exemplarily, in response to the execution instruction corresponding to the feedback voice, based on the target pose information corresponding to the target task, a robotic arm motion trajectory is generated; and based on the target recognition information and target pose information corresponding to the target task, a grasping gesture is generated; it controls the robotic arm of the humanoid robot to move to a preset position according to the robotic arm motion trajectory, generating a two - hand execution instruction; in response to the two - hand execution instruction, it controls the two hands of the humanoid robot to perform a grasping operation on the target object according to the grasping gesture.

[0045] Specifically, the large language model sends the execution instruction corresponding to the feedback voice to the robotic arm module and the dexterous hand module; the robotic arm module, in response to the execution instruction, generates a robotic arm motion trajectory according to the target pose information corresponding to the target task; the dexterous hand module generates a grasping gesture based on the target recognition information and target pose information corresponding to the target task.

[0046] Here, there is no limitation on the generation process of the robotic arm motion trajectory and the grasping gesture, which can be implemented based on existing technologies or based on the following methods.

[0047] For example: after receiving the target pose information sent by the large language model, the robotic arm module converts the target pose information corresponding to the target task into the end - effector pose of the robotic arm, and then the robotic arm module performs motion analysis according to the end - effector pose of the robotic arm and the orocos KDL library; based on DMP, it performs motion trajectory planning on the motion analysis result to generate a robotic arm motion trajectory.

[0048] Just as Figure 4As shown, the dexterous hand module generates a 3D point cloud model of the target object and the contact points of the target object based on the target pose information and the target recognition information; the dexterous hand module inputs the 3D point cloud model of the target object and the contact points of the target object into the conditional variational autoencoder together for feature extraction processing to generate a contact map; the dexterous hand module matches the current finger positions of the two hands of the humanoid robot with the contact map to generate a grasping gesture. When the dexterous hands module controls the two hands of the humanoid robot to perform a grasping operation on the target object according to the grasping gesture, it also performs slip detection and impedance control on the grasping process so that the two hands receive real-time feedback from the dexterous hands module, thereby adjusting the grasping gesture in a timely manner to ensure that the two hands of the humanoid robot can successfully grasp the target object.

[0049] Thus, in this embodiment, the robotic arm module, the dexterous hand module, and the large language model are combined to achieve the humanoid motion of the upper limb of the humanoid robot.

[0050] In S104, according to the chronological order of the task sequence, the upper limb of the humanoid robot sequentially executes the target tasks until all the target tasks in the task sequence are executed; a completion instruction is generated; the completion instruction is converted into voice information, and the task status corresponding to the voice information is fed back to the user.

[0051] This embodiment realizes the perception of multi-modal information based on the vision module and the voice module, and can solve the problem that the existing humanoid robot cannot perform real-time perception of the external environment and make timely adjustments due to the lack of a perception system; the large language model in this embodiment analyzes complex natural language instructions through its powerful language understanding and generation capabilities, and performs task planning and decision-making on the voice request based on these natural language instructions; thus, not only the collaborative control of the robotic arm and the two hands of the humanoid robot is realized, but also the problem that the simple structure of the tool at the end of the robotic arm cannot achieve dexterous operation of the two hands can be solved.

[0052] In another preferred embodiment of this embodiment, the method further includes that if the task analysis result indicates that there is no target object to be executed, the large language model is used to perform task planning processing on the task analysis result to generate a text instruction corresponding to the task analysis result; the text instruction corresponding to the task analysis result is broadcast in the form of voice for the task status; when the feedback voice instruction received from the user for the task status is to correct the task, the feedback voice is used as a voice request, and the large language model is continued to be used to perform task analysis on the text information corresponding to the voice request.

[0053] Based on the large language model and combined with the vision module and the speech module, this embodiment realizes the disassembly and reasoning of voice request instructions, the perception and fusion of multi-modal information, and the voice interaction feedback. It can not only accurately control the upper limbs of the humanoid robot to execute target tasks, but also enable users to communicate with the humanoid robot more intuitively, improve the usability and user satisfaction of the humanoid robot, and enhance the emotional connection between humans and machines.

[0054] Another embodiment of the present invention provides an interaction method for collaborative control of the upper limbs of a humanoid robot, which at least includes the following steps:

[0055] S201, receive the voice request sent by the user and convert the voice request into text information; the large language model extracts features from the text information to generate key information; perform prompt word matching on the key information to generate a matching result; perform analysis and reasoning based on the matching result to output a task analysis result.

[0056] S202, determine whether there is a target object to be executed in the task analysis result; if not, execute step S203; if so, execute step S204.

[0057] S203, use the large language model to perform task planning processing on the task analysis result to generate a text instruction corresponding to the task analysis result; broadcast the task status in voice form for the text instruction corresponding to the task analysis result; when receiving the feedback voice instruction from the user for correcting the task status, use the feedback voice as a voice request and continue to execute S201.

[0058] S204, obtain the key information of the target object from the task analysis result; obtain the target recognition information corresponding to the key information; perform target segmentation based on the target recognition information to generate a target segmentation result; perform pose detection of the target object based on the target segmentation result to generate target pose information; wherein, the target pose information includes: target object position information and target object attitude information.

[0059] S205, based on the task analysis result, the target recognition information, and the target pose information, use the large language model to perform task planning processing according to the principle of collaborative work of both arms to generate a task sequence and a text instruction corresponding to each task.

[0060] S206. For any target task in the task sequence: announce the task status by converting the text instruction corresponding to the target task into speech; when the feedback speech received from the user regarding the task status is to execute the task, then in response to the execution instruction corresponding to the feedback speech, based on the target pose information corresponding to the target task, generate a robotic arm motion trajectory; and based on the target recognition information and target pose information corresponding to the target task, generate a grasping gesture; control the robotic arm of the humanoid robot to move to a preset position according to the robotic arm motion trajectory, and generate a two-handed execution instruction; in response to the two-handed execution instruction, control the two hands of the humanoid robot to perform a grasping operation on the target object according to the grasping gesture.

[0061] S207. According to the chronological order of the task sequence, sequentially control the upper limbs of the humanoid robot to perform corresponding operation actions according to each target task, and generate a completion instruction; and announce the task status corresponding to the completion instruction by voice.

[0062] As Figure 3 shown, it is a schematic diagram of the upper limb grasping of a humanoid robot in an embodiment of the present invention; the interactive method for the upper limb collaborative control of the humanoid robot provided in this embodiment will be described in detail below in combination with a specific application scenario.

[0063] For example: the voice request is "I want to eat something today". After the voice module receives the voice request, it is converted into text information; and the text information is sent to the large language model; the large language model extracts the key information as "I want to eat", and then through prompt word matching and analysis and reasoning by the large language model, a task analysis result is obtained.

[0064] The task analysis result indicates that there is no target object; the large language model generates a corresponding text instruction based on the task analysis result; the text instruction is "What do you want to eat"; the large language model sends this text instruction to the voice module, and the voice module announces "What do you want to eat" to the user in voice form; the user gives a voice request of "I want to eat apples"; then the voice module converts the voice request of "I want to eat apples" into text information again and sends it to the large language model.

[0065] The large language model extracts the key information "I want to eat" and "apples" from the user's two voice requests; then the task decision unit outputs a decision: I will first see where the apples are and then get them for you. Then the large language model mobilizes the visual module to query whether there is an "apple" in the data collected by the sensors and / or cameras; if so, the large language model obtains the apple pose information and apple recognition information from the visual module; and based on the apple pose information, apple recognition information, and task analysis result; according to the principle of two-arm division of labor and cooperation, perform task planning processing to generate a first task of "first move the robotic arm to a preset position and then grab the apple", and a second task of "pick up the apple and put the apple in the user's hand".

[0066] For the first task in the task sequence: The robotic arm module controls the robotic arm of the humanoid robot to move to a preset position according to the robotic arm movement trajectory, and generates a two-handed execution instruction; the dexterous hand module responds to the two-handed execution instruction, and controls the two hands of the humanoid robot to perform a grasping operation on the apple according to the grasping gesture.

[0067] After both the first task and the second task are executed, a completion instruction is generated; and the completion instruction is converted into voice information to be fed back to the user.

[0068] In this embodiment, through the embodied hierarchical decision-making interaction framework, integrating the trajectory planning of the dual robotic arms and the gesture generation of the dexterous hands, it not only enhances the ability of the humanoid robot's upper limbs to perform complex task operations, but also strengthens the anthropomorphic and versatility of the humanoid robot, providing technical support for the research and development and practical application of the humanoid robot.

[0069] As Figure 5 shown, it is a schematic structural diagram of a control device for the coordinated movement of the upper limbs of a humanoid robot provided by an embodiment of the present invention.

[0070] An interactive device for the coordinated control of the upper limbs of a humanoid robot, the device 500 includes: a first task analysis module 501, configured to perform task analysis on the text information corresponding to the received voice request by using the large language model, and output a task analysis result; a first instruction disassembling module 502, configured to, if the task analysis result indicates that there is a target object to be executed, perform target object recognition processing according to the task analysis result, and output target recognition information and target pose information; based on the task analysis result, the target recognition information, and the target pose information, use the large language model to perform task planning processing according to the principle of dual-arm division of labor and cooperation, generate a task sequence and a text instruction corresponding to each task; a control module 503, configured to, for any target task in the task sequence: broadcast the task state of the text instruction corresponding to the target task in a voice form; when receiving the feedback voice from the user for the task state as executing the task, based on the target pose information and the target recognition information corresponding to the target task, control the upper limbs of the humanoid robot to perform an operation action corresponding to the feedback voice on the target object; a first generation module 504, configured to sequentially control the upper limbs of the humanoid robot to perform corresponding operation actions according to each target task in the time sequence of the task sequence, and generate an interaction result.

[0071] In a preferred implementation manner of this embodiment, the device further includes: a second instruction disassembling module, configured to, if the task analysis result indicates that there is no target object to be executed, use the large language model to perform task planning processing on the task analysis result, generate a text instruction corresponding to the task analysis result; and broadcast the task status in a voice form for the text instruction corresponding to the task analysis result; a second task analysis module, configured to, when receiving a feedback voice instruction from the user for the task status as a task correction, use the feedback voice as a voice request, and continue to use the large language model to perform task analysis on the text information corresponding to the voice request.

[0072] In a preferred implementation manner of this embodiment, the control module includes: a first generating unit, configured to generate a manipulator motion trajectory based on the target pose information corresponding to the target task; a second generating unit, configured to generate a grasping gesture based on the target recognition information and the target pose information corresponding to the target task; a third generating unit, configured to, in response to the execution instruction corresponding to the feedback voice, control the manipulator of the humanoid robot to move to a preset position according to the manipulator motion trajectory, and generate a two-handed execution instruction; a control unit, configured to, in response to the two-handed execution instruction, control the two hands of the humanoid robot to perform a grasping operation on the target object according to the grasping gesture.

[0073] In a preferred implementation manner of this embodiment, the first generating unit includes: a conversion sub-unit, configured to convert the target pose information corresponding to the target task into the end pose of the manipulator; a trajectory planning sub-unit, configured to perform motion analysis on the end pose of the manipulator by using a manipulator motion model according to the orocos KDL library; and perform motion trajectory planning on the motion analysis result based on DMP to generate a manipulator motion trajectory.

[0074] In a preferred implementation manner of this embodiment, the second generating unit includes: a first generating sub-unit, configured to generate a 3D point cloud model of the target object and a target object contact point based on the target pose information and the target recognition information; a second generating sub-unit, configured to perform feature extraction processing on the 3D point cloud model of the target object and the target object contact point to generate a contact map; and a gesture generating sub-unit, configured to match the current finger positions of the two hands of the humanoid robot with the contact map to generate a grasping gesture.

[0075] In a preferred implementation manner of this embodiment, the first task analysis module includes: a conversion unit, configured to convert the received voice request into text information; a feature extraction unit, configured to perform feature extraction on the text information to generate key information; a matching unit, configured to perform prompt word matching on the key information to generate a matching result; and a task content analysis unit, configured to perform analysis and reasoning based on the matching result and output a task analysis result.

[0076] In a preferred embodiment of the present embodiment, the first instruction disassembling module further includes: a first obtaining unit, configured to obtain key information of the target object from the task analysis result; a second obtaining unit, configured to obtain target recognition information corresponding to the key information; a target segmentation unit, configured to perform target segmentation based on the target recognition information to generate a target segmentation result; a first generating unit, configured to perform pose detection of the target object based on the target segmentation result to generate target pose information; wherein, the target pose information includes: target object position information and target object attitude information.

[0077] In a preferred embodiment of the present embodiment, the first generating module includes: a generating unit, configured to sequentially control the upper limbs of the humanoid robot to execute corresponding operation actions according to each target task in the time sequence of the task sequence to generate a completion instruction; a feedback unit, configured to convert the completion instruction into voice information and feedback the task status corresponding to the voice information to the user.

[0078] The above device can execute an interaction method for collaborative control of the upper limbs of a humanoid robot provided by an embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the interaction method for collaborative control of the upper limbs of a humanoid robot. For technical details not described in detail in this embodiment, reference can be made to an interaction method for collaborative control of the upper limbs of a humanoid robot provided by an embodiment of the present invention.

[0079] The present invention also provides an electronic device, including: a processor; a memory for storing executable instructions of the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement an interaction method for collaborative control of the upper limbs of a humanoid robot according to the present invention.

[0080] In addition to the above methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the methods according to various embodiments of the present application described in the "Exemplary Method" section above of this specification.

[0081] The computer program product can be written in any combination of one or more programming languages for programming code to execute the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed completely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or completely on a remote computing device or server.

[0082] In addition, an embodiment of the present application may also be a computer-readable storage medium storing computer program instructions, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the methods according to the following embodiments of the present application described in the "Exemplary Methods" section above of this specification.

[0083] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0084] The basic principles of the present application have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present application are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present application. In addition, the above-disclosed specific details are only for illustrative purposes and for ease of understanding, rather than limitations, and the above details do not limit the present application to necessarily adopt the above specific details for implementation.

[0085] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present application are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended words, meaning "including but not limited to", and can be used interchangeably with each other. The words "or" and "and" used herein refer to the word "and / or", and can be used interchangeably with each other, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with each other.

[0086] It should also be noted that in the devices, equipment, and methods of the present application, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present application.

[0087] The foregoing description of the disclosed aspects is provided to enable any person skilled in the art to make or use this application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this application. Thus, this application is not intended to be limited to the aspects shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0088] The foregoing description has been presented for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of this application to the forms disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.

[0089] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0090] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of these features. In the description of the present invention, "a plurality of" means two or more unless otherwise specifically defined.

[0091] As described above, the foregoing are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily conceive of changes or substitutions within the technical scope disclosed by the present invention, and all of them should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. An interaction method for collaborative control of the upper limbs of a humanoid robot, characterized in that: Use a large language model to perform task analysis on the text information corresponding to the received voice request, and output the task analysis result; If the task analysis result indicates that there is a target object to be executed, perform target object recognition processing according to the task analysis result, and output target recognition information and target pose information; Based on the task analysis result, target recognition information, and the target pose information, use the large language model to perform task planning processing according to the principle of dual-arm division of labor and cooperation, generate a task sequence and a text instruction corresponding to each task; if the task analysis result indicates that there is no target object to be executed, use the large language model to perform task planning processing on the task analysis result, generate a text instruction corresponding to the task analysis result; broadcast the task status in the form of voice for the text instruction corresponding to the task analysis result; When the received feedback voice instruction from the user for the task status is to correct the task, use the feedback voice as a voice request, and continue to use the large language model to perform task analysis on the text information corresponding to the voice request; For any target task in the task sequence: broadcast the task status in the form of voice for the text instruction corresponding to the target task; when the received feedback voice from the user for the task status is to execute the task, based on the target pose information and target recognition information corresponding to the target task, control the upper limbs of the humanoid robot to perform an operation action corresponding to the feedback voice on the target object; According to the time sequence of the task sequence, sequentially control the upper limbs of the humanoid robot to perform corresponding operation actions according to each target task, and generate an interaction result.

2. The method according to claim 1, wherein The controlling the upper limbs of the humanoid robot to perform an operation action corresponding to the feedback voice on the target object based on the target pose information and target recognition information corresponding to the target task; includes: In response to the execution instruction corresponding to the feedback voice, generate a robotic arm motion trajectory based on the target pose information corresponding to the target task; and generate a grasping gesture based on the target recognition information and target pose information corresponding to the target task; Control the robotic arm of the humanoid robot to move to a preset position according to the robotic arm motion trajectory, and generate a two-handed execution instruction; In response to the two-handed execution instruction, control the two hands of the humanoid robot to perform a grasping operation on the target object according to the grasping gesture.

3. The method according to claim 2, wherein The generating a robotic arm motion trajectory based on the target pose information corresponding to the target task; includes: Convert the target pose information corresponding to the target task into the end pose of the robotic arm; According to the orocos KDL library, use the robotic arm motion model to perform motion analysis on the end pose of the robotic arm; perform motion trajectory planning on the motion analysis result based on DMP to generate a robotic arm motion trajectory.

4. The method according to claim 2, wherein The generating a grasping gesture based on the target recognition information and target pose information corresponding to the target task; includes: Generate a 3D point cloud model of the target object and a target object contact point based on the target pose information and the target recognition information; Extract features from the 3D point cloud model of the target object and the contact points of the target object to generate a contact map; Match the current finger positions of the two hands of the humanoid robot with the contact map to generate a grasping gesture.

5. The method according to claim 1, wherein Use the large language model to perform task analysis on the text information corresponding to the received voice request and output a task analysis result; including: Convert the received voice request into text information; Extract features from the text information to generate key information; Match the key information with prompt words to generate a matching result; Perform analysis and reasoning based on the matching result and output a task analysis result.

6. The method according to claim 1, characterized in that Perform target object recognition processing according to the task analysis result and output target recognition information and target pose information; including: Obtain the key information of the target object from the task analysis result; Obtain the target recognition information corresponding to the key information; Perform target segmentation based on the target recognition information to generate a target segmentation result; Perform pose detection of the target object based on the target segmentation result to generate target pose information; where the target pose information includes: target object position information and target object attitude information.

7. The method according to claim 1, characterized in that, According to the time sequence of the task sequence, sequentially control the upper limbs of the humanoid robot to perform corresponding operation actions according to each target task to generate an interaction result; including: According to the time sequence of the task sequence, sequentially control the upper limbs of the humanoid robot to perform corresponding operation actions according to each target task to generate a completion instruction; Convert the completion instruction into voice information and feedback the task status corresponding to the voice information to the user.

8. An interaction device for collaborative control of the upper limbs of a humanoid robot, characterized in that, Including: The first task analysis module is used to perform task analysis on the text information corresponding to the received voice request by using a large language model and output a task analysis result; The first instruction decomposition module is used to perform target object recognition processing according to the task analysis result if the task analysis result indicates that there is a target object to be executed, and output target recognition information and target pose information; Based on the task analysis result, target recognition information, and the target pose information, use the large language model to perform task planning processing according to the principle of dual-arm division of labor and cooperation to generate a task sequence and a text instruction corresponding to each task; The second instruction decomposition module is used to perform task planning processing on the task analysis result by using the large language model if the task analysis result indicates that there is no target object to be executed, and generate a text instruction corresponding to the task analysis result; broadcast the task status in the form of voice for the text instruction corresponding to the task analysis result; The second task analysis module is used to continue using the large language model to perform task analysis on the text information corresponding to the voice request when the received feedback voice instruction from the user for the task status is to correct the task; A control module, for any target task in the task sequence: broadcasting the task status by voice for the text instruction corresponding to the target task; when receiving the feedback voice from the user for the task status as executing the task, controlling the upper limb of the humanoid robot to perform an operation action corresponding to the feedback voice on the target object based on the target pose information and the target recognition information corresponding to the target task. A first generation module, for sequentially controlling the upper limb of the humanoid robot to perform corresponding operation actions according to each target task in the time sequence of the task sequence, and generating an interaction result.

9. A computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, the method as described in any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Double-arm operation task learning method and system based on large model

    CN117697763A

  • Flexible control method and device for double upper limb mechanical arms of humanoid robot

    CN118143954A

  • Universal system of intelligent robot with body, construction method and use method

    CN117549310A

  • Equipment, voice control method and device thereof and readable storage medium

    CN118942460A