Interaction method and device for upper limb cooperative control of humanoid robot

By using a large language model and visual and voice modules in humanoid robots, real-time perception and task planning of the external environment is achieved, and the problem that humanoid robots in the existing technology cannot be perceived and adjusted in real time is solved, and the operation ability of complex tasks is improved.

CN119927906AActive Publication Date: 2025-05-06江淮前沿技术协同创新中心

Patent Information

Application Number
CN202510103239.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-06
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

Due to the lack of a perception system, existing humanoid robots cannot perform real-time perception and timely adjustments to the external environment, resulting in insufficient operational capabilities when performing complex tasks.

Method used

The large language model is used to combine the interactive method of visual and voice modules, and the task analysis is performed through voice requests, the target position is identified, and task planning is carried out according to the principle of division of labor and cooperation between the two arms, and task sequences and text instructions are generated to realize the coordinated control of the upper limbs of the humanoid robot.

Benefits of technology

Real-time perception and interaction of the external environment by humanoid robots is realized, the operation ability of the upper limbs when performing complex tasks is improved, and the robot's adaptability and autonomy to the environment is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119927906A_ABST
    Figure CN119927906A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an interaction method for upper limb cooperative control of a humanoid robot. The interaction method comprises the following steps: performing task analysis on a received voice request; if the target object needing to be executed exists, target object recognition processing is carried out according to the task analysis result; based on target identification information and target pose information output by identification processing and a task analysis result, performing task planning processing by utilizing a large language model according to a double-arm division cooperation principle, and generating a task sequence and a character instruction corresponding to each task; sequentially executing the following operations on each target task in the task sequence: based on the target pose information and the target identification information, controlling upper limbs of the humanoid robot to execute an operation action corresponding to the target task on the target object; therefore, on the basis of the large language model, the visual module and the voice module are matched, so that a user can communicate with the robot more visually; and the operation capability of the upper limbs of the humanoid robot for executing complex tasks can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of humanoid robots, and in particular relates to an interactive method and device for coordinated control of upper limbs of a humanoid robot. Background Art

[0002] The robotic arm is an important component of the humanoid robot and is the direct execution structure for the robot to complete tasks. The dual-arm collaborative robot can have two arms like humans and coordinate work within a certain range. This design enables the robot to greatly enhance its adaptability to complex assembly tasks and improve the efficiency of workspace utilization. However, the upper limb collaboration of the humanoid robot is not a simple combination of two robotic arms + grippers, but a multi-structure combination of dual robotic arms + dexterous hands. In addition to the dual-arm control goals to be achieved, it is also necessary to meet the coordination control between the arms and between the arms and hands and the adaptability to the environment. This greatly increases the complexity of the upper limb operation of the humanoid robot, requiring more advanced integrated systems, efficient collaborative control algorithms, accurate environmental perception and positioning, and adjustable control methods. In addition, humanoid robots not only need basic limb movement capabilities, but also should have a "smart brain" that can command disassembly, task planning, and human-computer interaction tasks, in order to reflect the superiority of humanoid robots in intelligence and autonomy compared to other conventional industrial robots. Therefore, based on the language understanding ability of the large model and various perceptual information such as voice and vision, combined with the arm-hand joint system, the robot can realize real-time perception of the external environment, real-time interaction and real-time adjustment of motion planning.

[0003] At present, the coordinated control of dual-arm robots is usually based on kinematics and dynamics control. The kinematics-based coordinated control focuses on the study of the redundant characteristics of the robot arm, and the dynamics-based coordinated control focuses on the study of the end force control of the robot arm. However, the solution of the inverse kinematics is quite different from the kinematics of a single arm due to the large difference in the redundant degrees of freedom of the dual-arm system. The different dual-arm configurations also determine the differences in the inverse solution methods. The uncertainty of the inverse solution also affects the way the dynamic model is established, and the changes in the speed and acceleration of the dual-arm movement and the difference in the relative force at the end. In an unstructured environment, the dual-arm robot system cannot obtain all the kinematic and dynamic information. How the robot perceives the external environment, including the position, speed, contact force and other parameters of the manipulated object, and how the external environment information participates in the coordinated control of the dual arms have become new research hotspots.

[0004] Chinese patent specification CN 118143954 A discloses a compliant control method and device for the upper limb dual mechanical arms of a humanoid robot, which constructs a closed chain kinematic model based on the closed motion chain composed of the upper limb dual mechanical arms and the operation object, and constructs a joint dynamic model for the upper limb dual mechanical arms, and sets a dual-arm cooperative controller for the dual mechanical arms to achieve dual-arm cooperative motion control. However, this method is completely based on traditional strategies, cannot achieve accurate environmental perception and real-time self-adjustment, and will not be able to achieve complex task operations in practical applications.

[0005] Chinese patent specification CN 117697763 A discloses a dual-arm operation task learning method and system based on a large model, which receives the task environment map and task description information for dual-arm operation through the large model, determines the reward function, and obtains the dual-arm operation strategy through the imitation learning algorithm. This method uses the large model to reinforce the design of the learning reward function, but the setting and improvement of the reward function are difficult. Robot hardware can usually only demonstrate one skill at a time. At the same time, this method focuses on dual-arm operation without emphasizing multi-modal perception and interaction, which is essentially different from the upper limb operation of humanoid robots in practical applications. Summary of the invention

[0006] In response to the above-mentioned problems existing in the prior art, an embodiment of the present invention provides an interactive method and device for coordinated control of the upper limbs of a humanoid robot; this method can solve the problem that existing humanoid robots are unable to perceive and adjust the external environment in real time due to the lack of a perception system; and further improve the operational capability of the upper limbs of the humanoid robot to perform complex tasks.

[0007] According to a first aspect of an embodiment of the present invention, an interactive method for collaborative control of an upper limb of a humanoid robot is provided, the method comprising: performing task analysis on text information corresponding to a received voice request using a large language model, and outputting a task analysis result; if the task analysis result indicates that there is a target object that needs to be executed, performing target object recognition processing according to the task analysis result, and outputting target recognition information and target posture information; based on the task analysis result, the target recognition information, and the target posture information, performing task planning processing according to the principle of dual-arm division of labor and cooperation using the large language model, generating a task sequence and text instructions corresponding to each of the tasks; for any target task in the task sequence: broadcasting the task status of the text instructions corresponding to the target task in the form of voice; when receiving a feedback voice from the user for the task status that the task is to be executed, based on the target posture information and target recognition information corresponding to the target task, controlling the upper limb of the humanoid robot to perform an operation action corresponding to the feedback voice on the target object; and controlling the upper limb of the humanoid robot to perform corresponding operation actions according to each target task in turn according to the time sequence of the task sequence to generate an interactive result.

[0008] Optionally, the method also includes: if the task analysis result indicates that there is no target object that needs to be executed, using the large language model to perform task planning processing on the task analysis result to generate text instructions corresponding to the task analysis result; reporting the task status of the text instructions corresponding to the task analysis result in voice form; when the user's feedback voice instruction for the task status is received to correct the task, using the feedback voice as a voice request, and continuing to use the large language model to perform task analysis on the text information corresponding to the voice request.

[0009] Optionally, based on the target posture information and target recognition information corresponding to the target task, the upper limbs of the humanoid robot are controlled to perform an operation action corresponding to the feedback voice on the target object; including: generating a robotic arm motion trajectory based on the target posture information corresponding to the target task; generating a grasping gesture based on the target recognition information and target posture information corresponding to the target task; in response to the execution instruction corresponding to the feedback voice, the robotic arm of the humanoid robot is controlled to move to a preset position according to the robotic arm motion trajectory, and a two-hand execution instruction is generated; in response to the two-hand execution instruction, the two hands of the humanoid robot are controlled to perform a grasping operation on the target object according to the grasping gesture.

[0010] Optionally, the generating of the robot arm motion trajectory based on the target posture information corresponding to the target task includes: converting the target posture information corresponding to the target task into the posture of the end of the robot arm; performing motion analysis on the posture of the end of the robot arm using the robot arm motion model according to the orocos KDL library; performing motion trajectory planning on the motion analysis results based on DMP to generate the robot arm motion trajectory.

[0011] Optionally, the target recognition information and target posture information corresponding to the target task are used to generate a grasping gesture; including: generating a 3D point cloud model of the target object and contact points of the target object based on the target posture information and the target recognition information; performing feature extraction processing on the 3D point cloud model of the target object and the contact points of the target object to generate a contact map; matching the current finger positions of both hands of the humanoid robot with the contact map to generate a grasping gesture.

[0012] Optionally, the large language model is used to perform task analysis on the text information corresponding to the received voice request, and the task analysis result is output; the received voice request is converted into text information: feature extraction is performed on the text information to generate key information; prompt word matching is performed on the key information to generate matching results; analysis and reasoning are performed based on the matching results, and the task analysis result is output.

[0013] Optionally, the target object recognition processing is performed according to the task analysis result, and target recognition information and target posture information are output; including: obtaining key information of the target object from the task analysis result; obtaining target recognition information corresponding to the key information; performing target segmentation based on the target recognition information to generate a target segmentation result; performing posture detection of the target object based on the target segmentation result to generate target posture information; wherein the target posture information includes: target position information and target posture information.

[0014] Optionally, the method controls the upper limbs of the humanoid robot to perform corresponding operation actions according to each target task in turn according to the time sequence of the task sequence to generate an interaction result; including: controlling the upper limbs of the humanoid robot to perform corresponding operation actions according to each target task in turn according to the time sequence of the task sequence to generate a completion instruction; converting the completion instruction into voice information, and feeding back the task status corresponding to the voice information to the user.

[0015] According to a second aspect of an embodiment of the present invention, there is also provided an interactive device for collaborative control of the upper limbs of a humanoid robot, the device comprising: a first task analysis module, for performing task analysis on text information corresponding to a received voice request using the large language model, and outputting a task analysis result; a first instruction disassembly module, for performing target object recognition processing according to the task analysis result if the task analysis result indicates that there is a target object to be executed, and outputting target recognition information and target posture information; based on the task analysis result, the target recognition information, and the target posture information, performing task planning processing according to the principle of division of labor and cooperation of both arms using the large language model; The system comprises a control module, which is used to generate a task sequence and text instructions corresponding to each of the tasks; a control module, which is used to broadcast the task status of any target task in the task sequence by voice using the text instructions corresponding to the target task; when a user's feedback voice for the task status is received for executing the task, the upper limbs of the humanoid robot are controlled to execute the operation corresponding to the feedback voice on the target object based on the target posture information and target recognition information corresponding to the target task; a first generating module, which is used to control the upper limbs of the humanoid robot to execute the corresponding operation according to each target task in turn in the time sequence of the task sequence to generate an interaction result.

[0016] According to a third aspect of an embodiment of the present invention, there is further provided an electronic device, comprising: a processor; a memory for storing instructions executable by the processor; the processor for reading the executable instructions from the memory and executing the instructions to implement the method described in the first aspect.

[0017] According to a fourth aspect of an embodiment of the present invention, there is further provided a computer-readable medium on which a computer program is stored. When the program is executed by a processor, the method described in the first aspect is implemented.

[0018] The embodiment of the present invention provides an interactive method and device for collaborative control of the upper limbs of a humanoid robot. The method comprises: first, performing task analysis on text information corresponding to a received voice request using a large language model, and outputting a task analysis result; second, if the task analysis result indicates that there is a target object to be executed, performing target object recognition processing according to the task analysis result, and outputting target recognition information and target posture information; based on the task analysis result, the target recognition information, and the target posture information, performing task planning processing according to the principle of dual-arm division of labor and cooperation using the large language model, and generating a task sequence and text instructions corresponding to each of the tasks; then, for any target task in the task sequence: broadcasting the task status of the text instructions corresponding to the target task in the form of voice; when receiving a feedback voice from a user for the task status that the task is to be executed, based on the target posture information and target recognition information corresponding to the target task, controlling the upper limbs of the humanoid robot to perform an operation action corresponding to the feedback voice on the target object; finally, according to the time sequence of the task sequence, controlling the upper limbs of the humanoid robot to perform corresponding operation actions according to each target task in turn, and generating an interactive result. Therefore, the method of this embodiment is based on a large language model, and cooperates with a visual module and a voice module. It can not only realize task decomposition reasoning, multimodal information perception fusion, and voice interaction feedback, but also solve the problem that existing humanoid robots are unable to perceive and adjust the external environment in real time due to the lack of a perception system; it further improves the operational ability of the upper limbs of humanoid robots to perform complex tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Hereinafter, some specific embodiments of the present invention will be described in detail in an exemplary and non-limiting manner with reference to the accompanying drawings. The same reference numerals in the accompanying drawings indicate the same or similar components or parts. It should be understood by those skilled in the art that these drawings are not necessarily drawn to scale. In the accompanying drawings:

[0020] Figure 1 A schematic diagram of a coordinated motion control system for upper limbs of a humanoid robot provided by an embodiment of the present invention;

[0021] Figure 2 A schematic flow chart of an interactive method for coordinated control of upper limbs of a humanoid robot provided by an embodiment of the present invention;

[0022] Figure 3 A schematic diagram of grasping by the upper limbs of a humanoid robot in one embodiment of the present invention;

[0023] Figure 4 A schematic diagram of a grasping gesture of a dexterous hand in one embodiment of the present invention;

[0024] Figure 5 A schematic diagram of the structure of a control device for coordinated movement of the upper limbs of a humanoid robot provided in one embodiment of the present invention. DETAILED DESCRIPTION

[0025] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0026] like Figure 1 , which is a schematic diagram of a coordinated motion control system for the upper limbs of a humanoid robot provided by an embodiment of the present invention.

[0027] A collaborative motion control system for the upper limbs of a humanoid robot, the system 100 comprising: an embodied hierarchical decision-making interaction model, a robotic arm module, and a dexterous hand module; the embodied hierarchical decision-making interaction model comprises: a large language model, a visual module, and a speech module. The speech module comprises a speech recognition unit, a text conversion unit, and a speech synthesis unit. The visual module comprises a target recognition unit, a target segmentation unit, and a target detection unit. The large language model comprises an instruction disassembly unit, an instruction reasoning unit, and a task decision unit. The robotic arm module comprises: a kinematic analysis unit and a motion trajectory planning unit. The dexterous hand module comprises: a target model generation unit and a grasping gesture unit.

[0028] The first information output by the voice module and the second information output by the visual module are input into the large language model for processing, and the task sequence and the text instructions corresponding to each task are output. For any target task in the task sequence: the voice module processes the text instructions and broadcasts the current task status in the form of voice; when the system receives the user's feedback instruction for the current task status to execute the task, the mechanical arm module is notified to control the mechanical arm of the humanoid robot to move to the preset position according to the mechanical arm movement trajectory; then the dexterous hand module is notified to control the hands of the humanoid robot to perform the grasping operation on the target object according to the grasping gesture.

[0029] The present invention proposes an interactive system for coordinated control of the upper limbs of a humanoid robot, including "arm-hand-brain". The embodied hierarchical decision-making interaction model decomposes the instructions, implements task planning, decision-making, and distributes them to the arms and dexterous hands of the humanoid robot. The robotic arm performs motion analysis and trajectory planning according to the corresponding content, and the dexterous hands generate grasping gestures according to the target model, and finally completes the operation task through coordinated motion.

[0030] like Figure 2 As shown, it is a flow chart of an interactive method for collaborative control of upper limbs of a humanoid robot provided by an embodiment of the present invention; Figure 4 , which is a schematic diagram of a grasping gesture of a dexterous hand in one embodiment of the present invention.

[0031] An interactive method for coordinated control of upper limbs of a humanoid robot comprises at least the following steps:

[0032] S101, performing task analysis on text information corresponding to the received voice request using a large language model, and outputting a task analysis result;

[0033] S102, if the task analysis result indicates that there is a target object to be executed, target object recognition processing is performed according to the task analysis result, and target recognition information and target posture information are output; based on the task analysis result, target recognition information, and target posture information, task planning processing is performed using a large language model according to the principle of dual-arm division of labor and cooperation, and a task sequence and text instructions corresponding to each of the tasks are generated;

[0034] S103, for any target task in the task sequence: the text instruction corresponding to the target task is broadcasted in the form of voice to report the task status; when the user's feedback voice for the task status is received to execute the task, based on the target posture information and target recognition information corresponding to the target task, the upper limbs of the humanoid robot are controlled to execute the operation action corresponding to the feedback voice on the target object;

[0035] S104, in accordance with the time sequence of the task sequence, controlling the upper limbs of the humanoid robot to perform corresponding operation actions according to each target task in turn, and generating an interaction result.

[0036] In S101, this embodiment does not impose any limitation on the specific process of performing task analysis on the large language model, as long as task planning can be implemented.

[0037] Exemplarily, the large language model is used to perform task analysis on the text information corresponding to the received voice request, and the task analysis result is output; including: converting the received voice request into text information; extracting features from the text information to generate key information; matching the key information with prompt words to generate matching results; analyzing and reasoning based on the matching results to output the task analysis result. There are two situations for the task analysis result output after task analysis: the first situation is that there is a target object that needs to be executed; the second situation is that there is no target object that needs to be executed.

[0038] For example: Situation 1: Please help me get an apple and put it in my hand.

[0039] The second situation includes: the target object exists, or the target object exists but does not need to be executed; for example: I am hungry and I want to eat; or, there is an apple on the table.

[0040] In S102, this embodiment does not impose any limitation on the specific process of the target object recognition processing and the algorithm applied, as long as the target recognition information and target posture information corresponding to the target object can be obtained.

[0041] Exemplarily, the target object recognition processing is performed according to the task analysis result, and target recognition information and target posture information are output; including: obtaining key information of the target object from the task analysis result; obtaining target recognition information corresponding to the key information; performing target segmentation based on the target recognition information to generate a target segmentation result; performing posture detection of the target object based on the target segmentation result to generate target posture information; wherein the target posture information includes: target position information and target posture information.

[0042] Specifically, the visual module is used to perform target recognition processing, and the recognition result is matched with the key information of the target; if the match is successful, the target recognition information is generated; the target recognition information includes: target point cloud data, target image data, target type, size and other information; then, the target segmentation processing is performed based on the target recognition information to generate the target segmentation result; the target segmentation result is subjected to posture detection to obtain the target posture information. After that, the target posture information and target recognition information are respectively input into the large language model; according to the principle of two-arm division of labor and cooperation, the large language model performs task instruction disassembly based on the task analysis results, target posture information and target recognition information, and outputs the task disassembly result; the task disassembly result is subjected to task instruction reasoning, and the task sequence and the text instructions corresponding to each task are output.

[0043] In S103, the speech synthesis unit converts the text instructions corresponding to the target task into speech information, and reports the task status to the user. Then, it is determined whether the target task needs to be executed through speech interaction; if the judgment result indicates that the target task needs to be executed, the upper limbs of the humanoid robot are controlled to perform operation actions on the target object based on the target posture information and target recognition information corresponding to the target task; if the judgment result indicates that the target task is not executed and the target task is corrected, the user's corrected speech for the task status is used as a new speech request; the large language model is continued to be used to perform task analysis on the speech request, that is, return to step S101.

[0044] Exemplarily, in response to the execution instruction corresponding to the feedback voice, a robotic arm motion trajectory is generated based on the target posture information corresponding to the target task; and a grasping gesture is generated based on the target recognition information and target posture information corresponding to the target task; the robotic arm of the humanoid robot is controlled to move to a preset position according to the robotic arm motion trajectory, and a two-hand execution instruction is generated; in response to the two-hand execution instruction, the two hands of the humanoid robot are controlled to perform a grasping operation on the target object according to the grasping gesture.

[0045] Specifically, the large language model sends the execution instructions corresponding to the feedback voice to the robotic arm module and the dexterous hand module; the robotic arm module responds to the execution instructions and generates the robotic arm motion trajectory according to the target posture information corresponding to the target task; the dexterous hand module generates the grasping gesture based on the target recognition information and target posture information corresponding to the target task.

[0046] Here, there is no limitation on the generation process of the robot arm motion trajectory and the grasping gesture, which can be implemented based on the existing technology or based on the following method.

[0047] For example: After receiving the target posture information sent by the large language model, the robotic arm module converts the target posture information corresponding to the target task into the end posture of the robotic arm. Then the robotic arm module performs motion analysis based on the end posture of the robotic arm and the orocos KDL library; based on the DMP, the motion analysis results are used to plan the motion trajectory and generate the robotic arm motion trajectory.

[0048] As Figure 4As shown in the figure, the dexterous hand module generates a 3D point cloud model of the target object and the contact points of the target object based on the target posture information and target recognition information; the dexterous hand module inputs the 3D point cloud model of the target object and the contact points of the target object into the conditional variational autoencoder for feature extraction and processing to generate a contact map; the dexterous hand module matches the current finger positions of the humanoid robot's hands with the contact map to generate a grasping gesture. When the dexterous hand module controls the humanoid robot's hands to perform a grasping operation on the target object according to the grasping gesture, it also performs slip detection and impedance control on the grasping process, so that the hands receive real-time feedback from the dexterous hand module, thereby adjusting the grasping gesture in time to ensure that the humanoid robot's hands can successfully grasp the target object.

[0049] Therefore, this embodiment combines the robotic arm module, the dexterous hand module, and the large language model to achieve humanoid movement of the upper limbs of the humanoid robot.

[0050] In S104, the upper limbs of the humanoid robot perform the target tasks in sequence according to the time sequence of the task sequence until all the target tasks in the task sequence are completed; generate a completion instruction; convert the completion instruction into voice information, and feedback the task status corresponding to the voice information to the user.

[0051] This embodiment realizes the perception of multimodal information based on the visual module and the voice module, which can solve the problem that the existing humanoid robots are unable to perceive and adjust the external environment in real time due to the lack of a perception system; the large language model of this embodiment parses complex natural language instructions through its powerful language understanding and generation capabilities, and performs task planning and decision-making for voice requests based on these natural language instructions; thereby, not only the coordinated control of the humanoid robot's robotic arm and both hands is realized, but also the problem that the simple structure of the tool at the end of the robotic arm cannot achieve dexterous operation of both hands can be solved.

[0052] In another preferred implementation of the present embodiment, the method further includes, if the task analysis result indicates that there is no target object to be executed, performing task planning processing on the task analysis result using the large language model to generate text instructions corresponding to the task analysis result; reporting the task status of the text instructions corresponding to the task analysis result in voice form; and when a user's feedback voice instruction for the task status is received as a correction task, using the feedback voice as a voice request, and continuing to perform task analysis on the text information corresponding to the voice request using the large language model.

[0053] This embodiment is based on a large language model, and cooperates with a visual module and a voice module to realize the disassembly and reasoning of voice request instructions, multimodal information perception and fusion, and voice interaction feedback; it can not only accurately control the upper limbs of the humanoid robot to perform target tasks, but also allow users to communicate with the humanoid robot more intuitively, improve the usability and user satisfaction of the humanoid robot, and enhance the emotional connection between man and machine.

[0054] Another embodiment of the present invention provides an interactive method for coordinated control of upper limbs of a humanoid robot, comprising at least the following steps:

[0055] S201, receiving a voice request sent by a user and converting the voice request into text information; performing feature extraction on the text information by a large language model to generate key information; performing prompt word matching on the key information to generate matching results; performing analysis and reasoning based on the matching results and outputting task analysis results.

[0056] S202, determine whether there is a target object that needs to be executed in the task analysis result; if not, execute step S203; if so, execute step S204.

[0057] S203, use the large language model to perform task planning on the task analysis results, generate text instructions corresponding to the task analysis results; report the task status in voice form using the text instructions corresponding to the task analysis results; when the user's feedback voice instruction for the task status is received to correct the task, the feedback voice is used as a voice request and S201 is continued.

[0058] S204, obtaining key information of the target object from the task analysis results; obtaining target recognition information corresponding to the key information; performing target segmentation based on the target recognition information to generate a target segmentation result; performing pose detection of the target object based on the target segmentation result to generate target pose information; wherein the target pose information includes: target position information and target pose information.

[0059] S205, based on the task analysis results, target recognition information, and target posture information, a large language model is used to perform task planning and processing according to the principle of dual-arm division of labor and cooperation, and a task sequence and text instructions corresponding to each task are generated.

[0060] S206, for any target task in the task sequence: the text instructions corresponding to the target task are broadcast in the form of voice to report the task status; when the user's feedback voice for the task status is to execute the task, in response to the execution instruction corresponding to the feedback voice, based on the target posture information corresponding to the target task, a robot arm motion trajectory is generated; and based on the target recognition information and target posture information corresponding to the target task, a grasping gesture is generated; the robot arm of the humanoid robot is controlled to move to a preset position according to the robot arm motion trajectory, and a two-hand execution instruction is generated; in response to the two-hand execution instruction, the two hands of the humanoid robot are controlled to perform a grasping operation on the target object according to the grasping gesture.

[0061] S207, in accordance with the time sequence of the task sequence, control the upper limbs of the humanoid robot to perform corresponding operation actions according to each target task in turn, generate completion instructions; and broadcast the task status corresponding to the completion instructions by voice.

[0062] like Figure 3 The figure is a schematic diagram of the upper limb grasping of a humanoid robot in one embodiment of the present invention; the interactive method for coordinated control of the upper limbs of a humanoid robot provided in this embodiment will be described in detail below in conjunction with a specific application scenario.

[0063] For example: if the voice request is "I want to eat something today", the voice module converts the voice request into text information after receiving it; and sends the text information to the large language model; the large language model extracts the key information "I want to eat", and then the large language model performs prompt word matching, analysis and reasoning to obtain the task analysis result.

[0064] The task analysis result indicates that there is no target object; the large language model generates a corresponding text instruction based on the task analysis result; the text instruction is "what do you want to eat"; the large language model sends the text instruction to the voice module, and the voice module announces "what do you want to eat" to the user in the form of voice; the user responds with a voice request of "I want to eat apples"; the voice module then converts the voice request of "I want to eat apples" into text information and sends it to the large language model.

[0065] The big language model extracts the key information "I want to eat" and "apple" from the user's two voice requests; then the task decision unit outputs the decision: I will see where the apple is first, and then give it to you. After that, the big language model mobilizes the visual module to query whether there is "apple" in the data collected by the sensor and / or camera; if so, the big language model obtains the apple posture information and apple recognition information from the visual module; and based on the apple posture information, apple recognition information, and task analysis results; task planning is performed according to the principle of division of labor and cooperation between two arms, generating the first task of "moving the robotic arm to the preset position before grabbing the apple" and the second task of "picking up the apple and putting it in the user's hand".

[0066] For the first task in the task sequence: the robotic arm module controls the robotic arm of the humanoid robot to move to a preset position according to the robotic arm motion trajectory, and generates a two-hand execution instruction; the dexterous hand module responds to the two-hand execution instruction, and controls the two hands of the humanoid robot to perform a grabbing operation on the apple according to the grabbing gesture.

[0067] After both the first task and the second task are executed, a completion instruction is generated; and the completion instruction is converted into voice information to be fed back to the user.

[0068] This embodiment uses an embodied hierarchical decision-making interaction framework, combines the trajectory planning of dual robotic arms and the gesture generation of dexterous hands, which not only enhances the ability of the upper limbs of the humanoid robot to perform complex tasks, but also strengthens the anthropomorphism and versatility of the humanoid robot, providing technical support for the development and application of humanoid robots.

[0069] like Figure 5 , which is a schematic diagram of the structure of a control device for coordinated movement of the upper limbs of a humanoid robot provided by an embodiment of the present invention.

[0070] An interactive device for collaborative control of upper limbs of a humanoid robot, the device 500 comprising: a first task analysis module 501, for performing task analysis on text information corresponding to a received voice request using the large language model, and outputting a task analysis result; a first instruction disassembly module 502, for performing target object recognition processing according to the task analysis result if the task analysis result indicates that there is a target object to be executed, and outputting target recognition information and target posture information; based on the task analysis result, target recognition information, and target posture information, performing task planning processing according to the large language model according to the principle of dual-arm division of labor and cooperation, and generating a task sequence; A column and text instructions corresponding to each of the tasks; a control module 503, used for: for any target task in the task sequence: the text instructions corresponding to the target task are broadcast in the form of voice to report the task status; when the user's feedback voice for the task status is received for executing the task, based on the target posture information and target recognition information corresponding to the target task, the upper limbs of the humanoid robot are controlled to perform the operation action corresponding to the feedback voice on the target object; a first generation module 504, used to control the upper limbs of the humanoid robot to perform the corresponding operation action according to each target task in turn in the time order of the task sequence to generate an interaction result.

[0071] In a preferred implementation manner of this embodiment, the device also includes: a second instruction disassembly module, which is used to use the large language model to perform task planning processing on the task analysis result if the task analysis result indicates that there is no target object that needs to be executed, and generate text instructions corresponding to the task analysis result; the task status is broadcasted in the form of voice using the text instructions corresponding to the task analysis result; and a second task analysis module, which is used to use the feedback voice as a voice request when the user's feedback voice instruction for the task status is received to correct the task, and continue to use the large language model to perform task analysis on the text information corresponding to the voice request.

[0072] In a preferred implementation of this embodiment, the control module includes: a first generating unit, used to generate a robotic arm motion trajectory based on the target posture information corresponding to the target task; a second generating unit, used to generate a grasping gesture based on the target recognition information and target posture information corresponding to the target task; a third generating unit, used to respond to the execution instruction corresponding to the feedback voice, control the robotic arm of the humanoid robot to move to a preset position according to the robotic arm motion trajectory, and generate a two-hand execution instruction; a control unit, used to respond to the two-hand execution instruction, control the two hands of the humanoid robot to perform a grasping operation on the target object according to the grasping gesture.

[0073] In a preferred implementation of this embodiment, the first generation unit includes: a conversion subunit, used to convert the target posture information corresponding to the target task into the end posture of the robot arm; a trajectory planning subunit, used to perform motion analysis on the end posture of the robot arm using the robot arm motion model according to the orocos KDL library; based on the DMP, motion trajectory planning is performed on the motion analysis results to generate the robot arm motion trajectory.

[0074] In a preferred implementation manner of this embodiment, the second generation unit includes: a first generation subunit, used to generate a 3D point cloud model of the target object and the contact points of the target object based on the target posture information and the target identification information; a second generation subunit, used to perform feature extraction processing on the 3D point cloud model of the target object and the contact points of the target object to generate a contact map; a gesture generation subunit, used to match the current finger positions of both hands of the humanoid robot with the contact map to generate a grasping gesture.

[0075] In a preferred implementation manner of this embodiment, the first task analysis module includes: a conversion unit, used to convert the received voice request into text information; a feature extraction unit, used to extract features from the text information and generate key information; a matching unit, used to match the key information with prompt words and generate matching results; and a task content analysis unit, used to perform analysis and reasoning based on the matching results and output task analysis results.

[0076] In a preferred implementation of this embodiment, the first instruction disassembly module also includes: a first acquisition unit, used to obtain key information of the target object from the task analysis result; a second acquisition unit, used to obtain target identification information corresponding to the key information; a target segmentation unit, used to perform target segmentation based on the target identification information, and generate a target segmentation result; a first generation unit, used to perform pose detection of the target object based on the target segmentation result, and generate target pose information; wherein the target pose information includes: target position information and target pose information.

[0077] In a preferred implementation of this embodiment, the first generation module includes: a generation unit, which is used to control the upper limbs of the humanoid robot to perform corresponding operation actions according to each target task in turn in the time order of the task sequence, and generate completion instructions; a feedback unit, which is used to convert the completion instructions into voice information and feedback the task status corresponding to the voice information to the user.

[0078] The above-mentioned device can execute an interactive method for collaborative control of upper limbs of a humanoid robot provided in one embodiment of the present invention, and has functional modules and beneficial effects corresponding to the interactive method for collaborative control of upper limbs of a humanoid robot. For technical details not fully described in this embodiment, please refer to an interactive method for collaborative control of upper limbs of a humanoid robot provided in one embodiment of the present invention.

[0079] The present invention also provides an electronic device, comprising: a processor; a memory for storing executable instructions of the processor; the processor is used to read the executable instructions from the memory and execute the instructions to implement an interactive method for collaborative control of the upper limbs of a humanoid robot described in the present invention.

[0080] In addition to the above-mentioned methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the method according to various embodiments of the present application described in the above-mentioned "Exemplary Method" section of this specification.

[0081] The computer program product may be written in any combination of one or more programming languages ​​to write program codes for performing the operations of the embodiments of the present application, including object-oriented programming languages, such as Java, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0082] In addition, an embodiment of the present application may also be a computer-readable storage medium on which computer program instructions are stored. When the computer program instructions are executed by a processor, the processor executes the steps of the method according to the following embodiments of the present application described in the above "Exemplary Method" section of this specification.

[0083] The computer readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can include, for example, but is not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0084] The basic principles of the present application are described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, effects, etc. mentioned in the present application are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. are required by each embodiment of the present application. In addition, the specific details disclosed above are only for the purpose of illustration and ease of understanding, not for limitation, and the above details do not limit the present application to being implemented by adopting the above specific details.

[0085] The block diagrams of the devices, apparatuses, equipment, and systems involved in this application are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagram. As will be appreciated by those skilled in the art, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open words, referring to "including but not limited to", and can be used interchangeably with them. The words "or" and "and" used here refer to the words "and / or" and can be used interchangeably with them, unless the context clearly indicates otherwise. The words "such as" used here refer to the phrase "such as but not limited to", and can be used interchangeably with them.

[0086] It should also be noted that in the apparatus, device and method of the present application, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present application.

[0087] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present application. Therefore, the present application is not intended to be limited to the aspects shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

[0088] The above description has been given for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present application to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions and sub-combinations thereof.

[0089] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory.

[0090] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0091] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. An interactive method for coordinated control of upper limbs of a humanoid robot, characterized in that: Perform task analysis on the text information corresponding to the received voice request using the large language model, and output the task analysis result; If the task analysis result indicates that there is a target object that needs to be executed, target object recognition processing is performed according to the task analysis result, and target recognition information and target posture information are output; Based on the task analysis results, target recognition information, and target posture information, the large language model is used to perform task planning processing according to the principle of dual-arm division of labor and cooperation, and a task sequence and text instructions corresponding to each task are generated; For any target task in the task sequence: the text instruction corresponding to the target task is broadcasted in the form of voice to report the task status; when the user's feedback voice for the task status is received to execute the task, based on the target posture information and target recognition information corresponding to the target task, the upper limbs of the humanoid robot are controlled to execute the operation action corresponding to the feedback voice on the target object; According to the time sequence of the task sequence, the upper limbs of the humanoid robot are controlled to perform corresponding operation actions according to each target task in turn to generate interaction results.

2. The method according to claim 1, characterized in that Also includes: If the task analysis result indicates that there is no target object to be executed, the task analysis result is processed by the large language model to perform task planning, and a text instruction corresponding to the task analysis result is generated; the task status is broadcasted in the form of voice using the text instruction corresponding to the task analysis result; When the user's feedback voice instruction for the task status is received as a correction task, the feedback voice is used as a voice request, and the large language model is continued to be used to perform task analysis on the text information corresponding to the voice request.

3. The method according to claim 1, characterized in that The method of controlling the upper limbs of the humanoid robot to perform an operation action corresponding to the feedback voice on the target object based on the target posture information and the target recognition information corresponding to the target task comprises: In response to the execution instruction corresponding to the feedback voice, a robot arm motion trajectory is generated based on the target posture information corresponding to the target task; and a grasping gesture is generated based on the target recognition information and the target posture information corresponding to the target task; Controlling the mechanical arm of the humanoid robot to move to a preset position according to the movement trajectory of the mechanical arm, and generating a two-hand execution instruction; In response to the two-hand execution instruction, the two hands of the humanoid robot are controlled to perform a grasping operation on the target object according to the grasping gesture.

4. The method according to claim 3, characterized in that The step of generating a robot arm motion trajectory based on the target posture information corresponding to the target task comprises: Converting the target pose information corresponding to the target task into the end pose of the robot arm; According to the orocos KDL library, the robot arm motion model is used to perform motion analysis on the end position of the robot arm; based on the DMP, the motion trajectory of the motion analysis result is planned to generate the robot arm motion trajectory.

5. The method according to claim 3, characterized in that: The step of generating a grasping gesture based on the target recognition information and the target posture information corresponding to the target task comprises: Based on the target posture information and the target recognition information, a 3D point cloud model of the target object and contact points of the target object are generated; Performing feature extraction processing on the 3D point cloud model of the target object and the contact points of the target object to generate a contact map; The current finger positions of both hands of the humanoid robot are matched with the contact map to generate a grasping gesture.

6. The method according to claim 1, characterized in that Performing task analysis on text information corresponding to the received voice request using the large language model and outputting task analysis results; including: Converting the received voice request into a text message; Extracting features from the text information to generate key information; Perform prompt word matching on the key information to generate a matching result; Perform analysis and reasoning based on the matching results and output task analysis results.

7. The method according to claim 1, characterized in that The target object recognition processing is performed according to the task analysis result, and the target recognition information and target posture information are output; including: Obtaining key information of the target object from the task analysis results; Acquiring target identification information corresponding to the key information; Performing target segmentation based on the target recognition information to generate a target segmentation result; The target object's posture detection is performed based on the target segmentation result to generate target posture information; wherein the target posture information includes: target object position information and target object posture information.

8. The method according to claim 1, characterized in that According to the time sequence of the task sequence, the upper limbs of the humanoid robot are controlled to perform corresponding operation actions according to each target task in turn to generate an interaction result; including: According to the time sequence of the task sequence, the upper limbs of the humanoid robot are controlled to perform corresponding operation actions according to each target task in turn, and a completion instruction is generated; The completion instruction is converted into voice information, and the task status corresponding to the voice information is fed back to the user.

9. An interactive device for coordinated control of upper limbs of a humanoid robot, characterized in that: include: A first task analysis module, configured to perform task analysis on text information corresponding to the received voice request using the large language model, and output a task analysis result; A first instruction disassembly module is used for, if the task analysis result indicates that there is a target object that needs to be executed, performing target object recognition processing according to the task analysis result, and outputting target recognition information and target posture information; Based on the task analysis results, target recognition information, and target posture information, the large language model is used to perform task planning processing according to the principle of dual-arm division of labor and cooperation, and a task sequence and text instructions corresponding to each task are generated; The control module is used for: for any target task in the task sequence: reporting the task status by voice using the text instruction corresponding to the target task; when receiving the feedback voice of the user for the task status for executing the task, based on the target posture information and target recognition information corresponding to the target task, controlling the upper limbs of the humanoid robot to execute the operation action corresponding to the feedback voice on the target object; The first generating module is used to control the upper limbs of the humanoid robot to perform corresponding operation actions according to each target task in turn according to the time sequence of the task sequence to generate an interaction result.

10. A computer readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Double-arm operation task learning method and system based on large model

    CN117697763A

  • Flexible control method and device for double upper limb mechanical arms of humanoid robot

    CN118143954A

  • Universal system of intelligent robot with body, construction method and use method

    CN117549310A

  • Equipment, voice control method and device thereof and readable storage medium

    CN118942460A

  • Mechanical arm control method, device and equipment and storage medium

    CN119319568A

Cited By

  • Large model-based medical double-arm robot task and motion planning method and system

    CN120439319A

  • Two-hand teleoperation control method and device based on natural language

    CN121179419A

  • Question and answer mode interaction method and system for robot audio-visual fusion

    CN121315979A

  • A method and system for robot audio-visual fusion question and answer interaction

    CN121315979B