Humanoid robot control method and system and humanoid robot
Through the integration of the visual language perception model and a task planner based on a large language model into the humanoid robot control system, the problems of low integration, poor adaptability and poor interpretability of the humanoid robot control method in the prior art are solved, and end-to-end control with high integration and high adaptability are achieved.
Patent Information
- Application Number
- CN202510480187.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-17
AI Technical Summary
In the prior art, the humanoid robot control method has problems such as low integration, poor adaptability and poor interpretability.
The visual language perception model perceives the field of view image and task instructions, obtains the three-dimensional coordinates of the target object, and uses a task planner based on the prompt word and three-dimensional coordinates to determine the control information. The control information is controlled by the arm control module, the main control module and the hand control module respectively.
End-to-end control from perception to execution is realized, the integration and adaptability of the humanoid robot control process is improved, and the interpretability and transparency of the control process is enhanced.
Smart Images

Figure CN119974029A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of robot control technology, and in particular, to a humanoid robot control method, system and humanoid robot. Background Art
[0002] Over the years, robots have undergone significant development and research, and have achieved remarkable results in various forms. Among them, humanoid robots have received increasing attention in recent years due to their high similarity to humans. Since humanoid robots can operate directly in human living and working spaces, people expect them to be able to complete more complex tasks.
[0003] There are a lot of studies in the prior art showing that large language models (LLMs) and visual language models (VLMs) can endow robots with important semantic planning and logical reasoning capabilities, enabling them to directly understand and execute complex instructions through natural language. However, these studies are usually targeted at ordinary robots with relatively simple configurations and excellent controllers.
[0004] In contrast, there are very few large-scale models based on humanoid robots, although there are many models suitable for handling different tasks of humanoid robots, such as visual understanding, motion planning and control. 。 However, these models are often isolated and have not yet been successfully integrated and implemented on humanoid robots. Even if they can be integrated, it takes a lot of time. At the same time, it is difficult to ensure adaptability and interpretability in complex environments. Therefore, the humanoid robot control method in the prior art has the problems of low integration, poor adaptability and poor interpretability. Summary of the invention
[0005] The purpose of the present application is to provide a humanoid robot control method, system and humanoid robot to address the deficiencies in the above-mentioned prior art, so as to solve the problems of low integration, poor adaptability and poor interpretability of the humanoid robot control method in the prior art.
[0006] To achieve the above purpose, the technical solution adopted in the embodiment of the present application is as follows: In a first aspect, an embodiment of the present application provides a method for controlling a humanoid robot, the method comprising: Acquire a task instruction and a field of view image, wherein the task instruction is used to instruct the humanoid robot to perform an action on a target object; Inputting the visual field image and the task instruction into a pre-trained visual language perception model, and having the visual language perception model perceive and process the visual field image and the task instruction to obtain the three-dimensional coordinates of the target object in the visual field image; Obtain a prompt word, and input the prompt word and the three-dimensional coordinates into a pre-trained task planner based on a large language model, wherein the task planner performs task planning based on the prompt word and the three-dimensional coordinates, and determines control information corresponding to the prompt word, wherein the control information includes first position information, second position information, and hand state information, and the prompt word at least includes the task instruction; The arm movement of the humanoid robot is controlled by an arm control module in the motion control system according to the first position information, and the hand movement of the humanoid robot is controlled by a hand control module in the motion control system according to the hand state information, and the leg movement of the humanoid robot is controlled by a body control module in the motion control system according to a pre-deployed reinforcement learning strategy and the second position information, and the reinforcement learning strategy is obtained through pre-training.
[0007] In a possible implementation, the visual language perception model includes an encoding module and a decoding module connected in sequence; The visual language perception model performs perception processing on the field of view image and the task instruction to obtain the three-dimensional coordinates of the target object in the field of view image, including: Inputting the visual field image and the task instruction into the encoding module, and the encoding module encodes the visual field image and the task instruction to obtain encoded features; The encoded features are input into the decoding module, and the decoding module decodes the encoded features to obtain the three-dimensional coordinates.
[0008] In a possible implementation, the encoding module includes: an image encoding module, a text encoding module and an encoding fusion module; The inputting the visual field image and the task instruction into the encoding module, and the encoding module encoding the visual field image and the task instruction to obtain the encoded features, includes: Inputting the field of view image into the image encoding module, and the image encoding module performs image perception encoding processing on the field of view image to obtain image features; Inputting the task instruction into the text encoding module, and the text encoding module performs text-aware encoding processing on the task instruction to obtain text features; The image features and the text features are input into the encoding fusion module, and the encoding fusion module performs feature fusion on the image features and the text features to obtain the encoded features.
[0009] In a possible implementation, the prompt word further includes: pre-configured configuration information, the configuration information including: skill configuration information, operation space constraint configuration information, safety constraint configuration information and kinematic configuration information; The skill configuration information is used to indicate executable actions of the arms, hands and legs of the humanoid robot, the operation space constraint configuration information is used to indicate the operable space of the humanoid robot, and the kinematic configuration information is used to indicate the prior kinematic information of the humanoid robot.
[0010] In a possible implementation, the task planner performs task planning based on the prompt word and the three-dimensional coordinates, and determines the control information corresponding to the prompt word, including: Obtaining the current height of the humanoid robot; The task planner performs task planning based on the prompt word, the current height and the three-dimensional coordinates, and determines control information corresponding to the prompt word.
[0011] In a possible implementation, controlling the leg movement of the humanoid robot according to the reinforcement learning strategy and the second position information includes: Generate a target action sequence according to the reinforcement learning strategy and the second position information; The legs of the humanoid robot are controlled to move according to the target action sequence, and when the legs move to the target position indicated by the second position information, a first start instruction is sent to the arm control module.
[0012] In a possible implementation, controlling the movement of the arm of the humanoid robot according to the first position information includes: In response to the first start instruction, based on the first position information, a pre-built inverse kinematics model is used to predict the joint rotation angle of the arm of the humanoid robot, and the arm of the humanoid robot is controlled to rotate according to the joint rotation angle, and after the rotation is completed, a second start instruction is sent to the hand control module.
[0013] In a possible implementation manner, controlling the hand movement of the humanoid robot according to the hand state information includes: In response to the second start instruction, obtaining a current state of the hand of the humanoid robot; Determine whether the current state is consistent with the hand state information, and if not, adjust the state of the hand of the humanoid robot to be consistent with the hand state information.
[0014] In the second aspect, another embodiment of the present application provides a humanoid robot control system, which includes: a visual language perception model, a task planner based on a large language model, and a motion control system, and the motion control system includes: an arm control module, a hand control module, and a body control module; a reinforcement learning strategy is deployed in the body control module, and the reinforcement learning strategy is obtained through pre-training. The humanoid robot control system is used to execute the steps of the humanoid robot control method described in the first aspect.
[0015] In a third aspect, another embodiment of the present application provides a humanoid robot, comprising: the humanoid robot control system described in the second aspect.
[0016] The beneficial effects of the present application are: the visual language perception model is used to perceive and process the field of view image and the task instruction, and the three-dimensional coordinates of the target object in the field of view image are obtained; and the task planner based on the large language model is used to plan the task based on the prompt words and the three-dimensional coordinates, and the control information corresponding to the prompt words is determined, so that the arm movement, hand movement and leg movement of the humanoid robot can be controlled respectively through the arm control module, the main body control module and the hand control module in the motion control system, and multiple modules such as visual perception, language understanding, task planning, and motion control can be integrated into a unified framework, so as to realize end-to-end control from perception to execution, and improve the integration of the humanoid robot control process. At the same time, the large language model, the visual language perception model and the motion control system are integrated, and the visual language perception model can be used to adapt to the changes of the target object in the dynamic environment, and the task planner can be used to adapt to different task requirements, so as to enhance the adaptability of the robot in a complex environment. In addition, the visual language perception model provides intuitive perception results, and the task planner based on the large language model generates interpretable control information, which is convenient for human understanding and verification, thereby enhancing the interpretability of the humanoid robot control process and improving the transparency and credibility of the humanoid robot control process. In addition, by controlling the arm movement, hand movement and leg movement of the humanoid robot respectively through the arm control module, the body control module and the hand control module, the control accuracy and task execution capability of the robot can be improved, ensuring the flexibility of the humanoid robot control process. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1A schematic diagram of a flow chart of a humanoid robot control method provided in an embodiment of the present application; Figure 2 A schematic diagram of a structure of a visual language perception model provided in an embodiment of the present application; Figure 3 A schematic diagram of a flow chart for obtaining the three-dimensional coordinates of a target object in a field of view image in a humanoid robot control method provided in an embodiment of the present application; Figure 4 Another structural schematic diagram of the visual language perception model provided in the embodiment of the present application; Figure 5 A schematic diagram of a flow chart for obtaining encoded features in a humanoid robot control method provided in an embodiment of the present application; Figure 6 A schematic diagram of a flow chart for determining control information corresponding to a prompt word in a humanoid robot control method provided in an embodiment of the present application; Figure 7 A schematic diagram of a flow chart of controlling the leg movement of a humanoid robot in the humanoid robot control method provided in an embodiment of the present application; Figure 8 A schematic diagram of a flow chart of controlling the hand movement of a humanoid robot in the humanoid robot control method provided in an embodiment of the present application; Fig. 9 A schematic diagram of a humanoid robot control system provided in an embodiment of the present application. DETAILED DESCRIPTION
[0019] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of explanation and description and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn in real proportion. The flowchart used in this application shows the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowchart can be implemented out of sequence, and the steps without logical context can be reversed in order or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart under the guidance of the content of the present application, or remove one or more operations from the flowchart.
[0020] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.
[0021] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.
[0022] There is an endless stream of research and applications of robots based on large language models. A large number of studies in the prior art have shown that large language models (LLMs) and visual language models (VLMs) can give robots important semantic planning and logical reasoning capabilities, enabling them to directly understand and execute complex instructions through natural language. However, these studies are usually aimed at ordinary robots with relatively simple configurations and excellent controllers.
[0023] In contrast, there are very few large-scale models based on humanoid robots, although there are many models suitable for handling different tasks of humanoid robots, such as visual understanding, motion planning and control. 。 However, these models are often isolated and have not yet been successfully integrated and implemented on humanoid robots. Even if they can be integrated, it takes a lot of time. At the same time, it is difficult to ensure adaptability and interpretability in complex environments. Therefore, the humanoid robot control method in the prior art has the problems of low integration, poor adaptability and poor interpretability.
[0024] Based on the above problems, the embodiment of the present application proposes a humanoid robot control method, which senses and processes the field of view image and task instructions through a visual language perception model to obtain the three-dimensional coordinates of the target object in the field of view image, and performs task planning based on prompt words and three-dimensional coordinates through a task planner based on a large language model, and determines the control information corresponding to the prompt words, so that the arm movement, hand movement and leg movement of the humanoid robot can be controlled respectively through the arm control module, the main body control module and the hand control module in the motion control system, and multiple modules such as visual perception, language understanding, task planning, motion control, etc. can be integrated into a unified framework, realizing end-to-end control from perception to execution, improving the integration of the humanoid robot control process, and at the same time, enhancing the adaptability of the robot in a complex environment, and also enhancing the interpretability of the humanoid robot control process, and improving the transparency and credibility of the humanoid robot control process. In addition, it can also improve the control accuracy and task execution capability of the robot, and ensure the flexibility of the humanoid robot control process.
[0025] Figure 1 A schematic diagram of a flow chart of a humanoid robot control method provided in an embodiment of the present application, referring to Figure 1 As shown, the execution subject of the method can be any electronic device with processing capability, such as a main control device of a humanoid robot, and the method includes: S101, obtaining task instructions and visual field images.
[0026] The task instructions are used to instruct the humanoid robot to perform actions on the target object.
[0027] Optionally, the target object can be understood as an object that needs to be operated by the humanoid robot, such as a valve waiting to be opened or closed, goods waiting to be transported, etc. It can be obtained through voice commands, text commands or gesture commands, and the task information to be performed by the humanoid robot can be determined through the task command.
[0028] Optionally, the field of view image can be understood as an image within the field of view of the humanoid robot in the current environment, which can be obtained by using an RGB camera or an RGB-D camera (such as Intel RealSense, Azure Kinect) to capture the environment image, or by using a depth sensor (such as LiDAR, ToF camera) to obtain depth information of the environment.
[0029] S102, inputting the visual field image and the task instruction into a pre-trained visual language perception model, and the visual language perception model perceives and processes the visual field image and the task instruction to obtain the three-dimensional coordinates of the target object in the visual field image.
[0030] Optionally, the obtained task instructions and the visual field image are input into a pre-trained visual language perception model for perception processing to obtain the three-dimensional coordinates of the target object in the visual field image, wherein the three-dimensional coordinates may be the three-dimensional coordinates of the target object in the reference coordinate system of the humanoid robot.
[0031] Exemplarily, after the task instructions and the field of view image are input into the visual language perception model, the visual language perception model performs preprocessing on the field of view image such as pixel value normalization, image enhancement, and resolution adjustment, and performs preprocessing on the task instructions such as word segmentation, padding, or truncation to obtain the preprocessed image and text sequence, and the preprocessed image and text sequence are respectively input into the visual encoder and text encoder for encoding processing, and then feature fusion is performed, and target detection processing is performed on the fused features to obtain the bounding box and category of the target object, and then the target object is estimated and the coordinate system is converted to obtain the three-dimensional coordinates.
[0032] Exemplarily, the visual encoder can be implemented based on a convolutional neural network (CNN) or a visual Transformer, the language encoder can be implemented based on a Transformer or BERT, the object detection can be implemented based on an object detection head (such as FasterR-CNN, YOLO), and the coordinate estimation can be implemented based on depth information (such as the depth channel of an RGB-D image) or a monocular depth estimation model.
[0033] S103, obtaining a prompt word, and inputting the prompt word and the three-dimensional coordinates into a pre-trained task planner based on a large language model, and the task planner performs task planning based on the prompt word and the three-dimensional coordinates to determine control information corresponding to the prompt word.
[0034] The control information includes first position information, second position information and hand state information, and the prompt words at least include task instructions.
[0035] Optionally, a prompt word is obtained, where the prompt word is used to indicate relevant command and constraint information when the humanoid robot performs a task. The prompt word and three-dimensional coordinates are input into a pre-trained task planner based on a large prediction model, and task planning is performed through the task planner to obtain control information corresponding to the prompt word.
[0036] Exemplarily, the prompt words and three-dimensional coordinates are input into the task planner, which parses and standardizes the prompt words and three-dimensional coordinates. For example, natural language processing techniques (such as word segmentation, part-of-speech tagging, named entity recognition, etc.) can be used to pre-process the prompt words so that they can be better understood by the large language model. The input three-dimensional coordinates can be verified to ensure that they are valid and meet the workspace requirements of the humanoid robot, and the three-dimensional coordinates can be converted into a format suitable for processing by the large language model, such as encoding the coordinate values into a string or a sequence of numbers.
[0037] Exemplarily, the task planner combines the standardized prompt words and three-dimensional coordinates into an input sequence, and perceives the environment in which the robot is located in real time and models it. For example: through sensor data (such as cameras, lidar, etc.), objects, spatial relationships, and possible obstacles in the environment are identified, and a map or model of the robot's surrounding environment is constructed. Based on the input sequence and the results of environmental perception, task planning is performed, specifically, including determining the robot's motion path, the motion trajectory of the arm, and the state information of the hand. Among them, in the task planning process, the task planner can adopt a variety of algorithms and strategies, such as A* algorithm, Dijkstra algorithm, genetic algorithm, reinforcement learning, etc., so that the optimal solution or approximate optimal solution can be quickly found in a complex environment, and the corresponding control information can be generated.
[0038] Exemplarily, the task planner can plan the task of the humanoid robot into one or more subtasks, each subtask has corresponding control information. The task planner can exchange data with the motion control system of the humanoid robot, and send the control information corresponding to each subtask to the corresponding controller according to the execution order of each subtask and combined with the execution feedback of the previous task.
[0039] The control information of each subtask may include first position information for controlling the arm of the humanoid robot, hand state information for controlling the hand movement of the humanoid robot, and second position information for controlling the body of the humanoid robot. The first position information may include the position coordinates of the arm of the humanoid robot when performing the subtask and the speed or time information when the arm reaches the position coordinates, the second position information may include the position coordinates of the body of the humanoid robot when performing the subtask and the speed or time information when the body reaches the position coordinates, and the hand state information may include the opening and closing state of the hand and the size of the hand, etc. The position coordinates may be Cartesian space position coordinates.
[0040] S104. Control the arm movement of the humanoid robot according to the first position information through the arm control module in the motion control system, control the hand movement of the humanoid robot according to the hand state information through the hand control module in the motion control system, and control the leg movement of the humanoid robot according to the pre-deployed reinforcement learning strategy and the second position information through the body control module in the motion control system.
[0041] Among them, the reinforcement learning strategy is obtained through pre-training.
[0042] Optionally, the task planner may transmit the first position information in the control information to an arm control module in the motion control system, transmit the second position information to a body control module in the motion control system, and transmit the hand state information to a hand control module in the motion control system.
[0043] Optionally, data can also be exchanged between the arm control module, the body control module and the hand control module in the motion control system to ensure that the motion process of the humanoid robot is more accurate, vivid and anthropomorphic.
[0044] Optionally, a mapping relationship between pre-obtained position coordinates and joint angles can be deployed in the arm control module. After obtaining the first position information, the arm control module can calculate the joint angle of the arm of the humanoid robot based on the position coordinates and the mapping relationship in the first position information, thereby controlling the arm of the humanoid robot to move according to the joint angle.
[0045] Optionally, a pre-trained reinforcement learning strategy can be deployed in the main control module. After obtaining the second position information, the main control module can perform reasoning and control based on the reinforcement learning strategy and the position coordinates in the second position information, so that the humanoid robot can move to the target position. Exemplarily, the main control module can obtain the current state of the humanoid robot in real time through encoders, IMUs, plantar force sensors, etc., and encode the current state and the position coordinates in the second position information as the input vector of the reinforcement learning strategy, thereby generating control actions through the reinforcement learning strategy, and decoding the control actions into specific leg joint control instructions (such as leg joint angles or torques), and then the leg joint controller (such as a motor driver) in the main control module adjusts the joint angles, speeds or torques according to the leg joint control instructions, generates a suitable gait (such as walking, running), adjusts the leg motion trajectory, and uses zero moment point (ZMP) control or model predictive control (MPC) to ensure that the robot maintains balance during movement. And measure the new state of the humanoid robot in real time, and dynamically adjust the control action according to the new state and the position coordinates in the second position information, and repeat the above process until the humanoid robot reaches the position coordinates in the second position information.
[0046] Optionally, the hand control module may control hand movement according to hand state information.
[0047] In this embodiment, the visual language perception model is used to perceive and process the field of view image and the task instruction to obtain the three-dimensional coordinates of the target object in the field of view image, and the task planner based on the large language model is used to perform task planning based on the prompt words and the three-dimensional coordinates to determine the control information corresponding to the prompt words, so that the arm movement, hand movement and leg movement of the humanoid robot can be controlled respectively through the arm control module, the main body control module and the hand control module in the motion control system, and multiple modules such as visual perception, language understanding, task planning, and motion control can be integrated into a unified framework, so as to realize end-to-end control from perception to execution, and improve the integration of the humanoid robot control process. At the same time, the large language model, the visual language perception model and the motion control system are integrated, and the visual language perception model can be used to adapt to the changes of the target object in the dynamic environment, and the task planner can be used to adapt to different task requirements, so as to enhance the adaptability of the robot in a complex environment. In addition, the visual language perception model provides intuitive perception results, and the task planner based on the large language model generates interpretable control information, which is convenient for human understanding and verification, thereby enhancing the interpretability of the humanoid robot control process and improving the transparency and credibility of the humanoid robot control process. In addition, by controlling the arm movement, hand movement and leg movement of the humanoid robot respectively through the arm control module, the body control module and the hand control module, the control accuracy and task execution capability of the robot can be improved, ensuring the flexibility of the humanoid robot control process.
[0048] In one possible implementation, Figure 2 A schematic diagram of a structure of a visual language perception model provided in an embodiment of the present application, Figure 3 A schematic diagram of a flow chart of obtaining the three-dimensional coordinates of a target object in a field of view image in a humanoid robot control method provided in an embodiment of the present application, referring to Figure 2 as well as Figure 3 As shown, the visual language perception model includes an encoding module and a decoding module connected in sequence; in the above S102, the visual language perception model perceives and processes the visual field image and the task instruction to obtain the three-dimensional coordinates of the target object in the visual field image, including: S301, inputting the visual field image and the task instruction into the encoding module, and the encoding module encodes the visual field image and the task instruction to obtain encoded features.
[0049] Optionally, the field of view image and the task instruction may be pre-processed and then input into a coding module, which encodes the field of view image and the task instruction to capture key information in the field of view image and the task instruction to obtain encoded features.
[0050] Exemplarily, the encoding module can be implemented based on a deep learning framework (such as TensorFlow, PyTorch, etc.), for example, it can be implemented based on a convolutional neural network (CNN) and a recurrent neural network (RNN) or a Transformer.
[0051] S302: Input the encoded features into a decoding module, and the decoding module decodes the encoded features to obtain three-dimensional coordinates.
[0052] Optionally, the encoded features are input into a decoding module, and the decoding module decodes the encoded features to obtain three-dimensional coordinates.
[0053] Exemplarily, the decoding module can be implemented based on dense layers, upsampling layers or transposed convolutional layers.
[0054] By processing field of view images and task instructions in an encoding and decoding manner, multimodal information fusion can be achieved, the comprehensiveness of perception can be improved, and features can be efficiently extracted and compressed, reducing computational complexity. At the same time, this modular design can enhance system scalability, suppress noise and improve robustness. In addition, it also supports end-to-end learning and improves adaptability.
[0055] In one possible implementation, Figure 4 Another structural diagram of the visual language perception model provided in the embodiment of the present application is shown in FIG. Figure 5 A schematic diagram of a flow chart of obtaining encoded features in the humanoid robot control method provided in the embodiment of the present application, referring to Figure 4 as well as Figure 5 As shown, the encoding module includes: an image encoding module, a text encoding module and an encoding fusion module; the above S301 inputs the field of view image and the task instruction into the encoding module, and the encoding module encodes the field of view image and the task instruction to obtain the encoded features, including: S501: Input a field of view image into an image encoding module, and the image encoding module performs image perception encoding processing on the field of view image to obtain image features.
[0056] Optionally, the image encoding module may preprocess the input field of view image first, including denoising, enhancement, cropping and other operations, to improve the image quality and reduce the complexity of subsequent encoding, and then extract key features in the image, such as edges, textures, colors, shapes, etc., and then encode the extracted image features to obtain image features. Among them, discrete encoding methods such as One-hot encoding, hash encoding, etc. or continuous encoding methods such as floating point representation may be adopted.
[0057] S502: Input the task instruction into a text encoding module, and the text encoding module performs text-aware encoding processing on the task instruction to obtain text features.
[0058] Optionally, the text encoding module can first perform natural language processing on the input task instructions, such as word segmentation, part-of-speech tagging, named entity recognition, etc., to understand the semantics and structure of the task instructions, and then parse the processed task instructions into machine-understandable formats such as semantic trees and abstract syntax trees, and then encode the parsed instructions into a feature matrix of a preset type or format to obtain text features.
[0059] S503: Input the image features and the text features into a coding fusion module, and the coding fusion module performs feature fusion on the image features and the text features to obtain encoded features.
[0060] Optionally, the obtained image features and text features are input into a coding fusion module, and the coding fusion module performs multimodal feature fusion on the image features and text features to obtain encoded features.
[0061] Among them, multimodal feature fusion can be implemented based on deep learning models, such as attention mechanism or neural network.
[0062] By encoding text and images separately and then performing multimodal fusion, we can fully utilize the characteristics of each modality, improve perception accuracy, suppress unimodal noise, improve robustness, and enhance information complementarity, which helps to fully understand the task and environment. In addition, this modular design is easy to optimize and expand, supports complex task scenarios, and helps to improve the intelligence level of humanoid robots.
[0063] In a possible implementation, the prompt word also includes: pre-configured configuration information, the configuration information includes: skill configuration information, operation space constraint configuration information, safety constraint configuration information and kinematic configuration information.
[0064] The skill configuration information is used to indicate the executable actions of the arms, hands and legs of the humanoid robot, the operation space constraint configuration information is used to indicate the operable space of the humanoid robot, and the kinematic configuration information is used to indicate the prior kinematic information of the humanoid robot.
[0065] Optionally, the prompt words include task instructions, skill configuration information, operation space constraint configuration information, safety constraint configuration information and kinematic configuration information.
[0066] Among them, the skill configuration information may include the arm skills, hand skills and leg skills of the humanoid robot. For example, the arm skills may include left arm movement, right arm movement, arm swinging, folding and two-handed coordinated movements, etc., the hand skills include grasping, grasping, changing hands, releasing, etc., and the leg skills include standing upright, bending, center of gravity transfer, kicking, leg lifting, etc. Through the skill configuration information, the humanoid robot can adapt to the requirements of various tasks and improve the efficiency and success rate of task execution.
[0067] The operation space constraint configuration information is used to define the spatial range of the robot's operation, such as the working area, obstacle location, etc. The operation space constraint configuration information can avoid collisions between the humanoid robot and the environment to ensure safe operation. It can also optimize path planning and improve the movement efficiency of the humanoid robot.
[0068] Among them, the safety constraint configuration information is used to set safety restrictions, such as: the minimum safe distance between the humanoid robot and humans to prevent the robot from colliding with humans when moving or performing tasks. A list of fragile items, valuable items, etc. that the humanoid robot must not touch or damage. Limit the force used by the humanoid robot when performing tasks to prevent damage to humans or objects due to excessive force. The maximum movement speed of the humanoid robot in different scenarios to ensure that it can remain safe when interacting with humans or objects, and the emergency stop mechanism of the humanoid robot to prevent the humanoid robot from overloading or performing dangerous actions. The reliability and safety of the system can be enhanced through the safety constraint configuration information.
[0069] The kinematic configuration information is used to indicate the kinematic characteristics of the humanoid robot, including motion parameters such as joint angles, angular velocities, accelerations, and the position and posture information of the humanoid robot, so as to ensure the feasibility and accuracy of motion planning. It can also ensure the execution of complex actions and improve the quality of task completion.
[0070] In one possible implementation, Figure 6 A schematic diagram of a flow chart for determining control information corresponding to a prompt word in a humanoid robot control method provided in an embodiment of the present application, referring to Figure 6 As shown, in the above S103, the task planner performs task planning based on the prompt word and the three-dimensional coordinates, and determines the control information corresponding to the prompt word, including: S601. Obtain the current height of the humanoid robot.
[0071] Optionally, the current height of the humanoid robot may be obtained by performing calculations based on the environmental depth information obtained by a depth sensor (such as LiDAR or ToF camera).
[0072] Exemplarily, the environmental depth information can be filtered and segmented through the Open3D point cloud processing library, and the ground plane can be fitted through the RANSAC algorithm to obtain the plane equation. The point cloud data belonging to the ground is extracted as a reference for height calculation, and the bottom point of the humanoid robot is detected. Then, according to the ground plane equation, the vertical distance from the bottom point of the humanoid robot to the ground is calculated to obtain the current height of the humanoid robot.
[0073] S602: The task planner performs task planning based on the prompt word, the current height, and the three-dimensional coordinates, and determines control information corresponding to the prompt word.
[0074] Exemplarily, the task planner can encode the prompt word, current height and three-dimensional coordinates respectively to obtain the prompt word coding features, height coding features and coordinate coding features, and perform feature fusion on the prompt word coding features, height coding features and coordinate coding features to obtain fused features, and then perform task decomposition and task planning to obtain control information corresponding to the prompt word.
[0075] By obtaining the current height of the humanoid robot and performing task planning based on the prompt word, current height and three-dimensional coordinates, the control information corresponding to the prompt word can be obtained. The height and coordinate information can be combined to achieve accurate motion planning and execution. It can also avoid collisions and dangerous actions through height and space constraints, thereby optimizing path and action planning and improving task execution efficiency.
[0076] In one possible implementation, Figure 7 A schematic diagram of a flow chart of controlling the leg movement of a humanoid robot in a humanoid robot control method provided in an embodiment of the present application, referring to Figure 7 As shown, S104 controls the leg movement of the humanoid robot according to the reinforcement learning strategy and the second position information, including: S701: Generate a target action sequence according to a reinforcement learning strategy and second position information.
[0077] Optionally, the main control module can obtain the current state of the humanoid robot in real time through an encoder, IMU, plantar force sensor, etc., and encode the current state and the position coordinates in the second position information as an input vector of the reinforcement learning strategy, thereby generating a target action sequence through the reinforcement learning strategy.
[0078] The target action sequence includes that when the humanoid robot moves to the position corresponding to the second position information, the subject needs to perform at least one action. For example, it can be the angle, speed or torque of the leg joint, which is used to realize the humanoid robot's leg standing, bending, center of gravity transfer, kicking, leg lifting and other actions.
[0079] S702: Control the legs of the humanoid robot to move according to the target action sequence, and send a first start instruction to the arm control module when the legs move to the target position indicated by the second position information.
[0080] Optionally, after obtaining the target action sequence, the subject control module may traverse the target action sequence to implement execution of the target action sequence.
[0081] Optionally, during the traversal process, for the current action in the traversed target action sequence, the main control module obtains the state of the humanoid robot after the previous action is executed in real time through encoders, IMUs, plantar force sensors, etc., and combines the reinforcement learning strategy, the current action and the state after the previous action is executed to determine whether to adjust the current action. If so, the current action is adjusted through the reinforcement learning strategy and the state after the previous action is executed, and the current action is executed. If not, the current action is executed directly. Specifically, the execution of the current action can be achieved by controlling the leg joint controller of the leg until the traversal is completed.
[0082] Optionally, after the traversal is completed, the main control module can determine the current position of the humanoid robot again, and when it is determined that the humanoid robot has moved to the target position indicated by the second position information, send a first start instruction to the arm control module to make the arm control module start moving.
[0083] Exemplarily, the body control module may also send a first start instruction to the arm control module when executing the last action in the target action sequence, so as to make the movement process of the humanoid robot more rich and anthropomorphic.
[0084] In a possible implementation, the above S104 controls the arm movement of the humanoid robot according to the first position information, including: In response to the first start instruction, based on the first position information, a pre-built inverse kinematics model is used to predict the joint rotation angle of the arm of the humanoid robot, and the arm of the humanoid robot is controlled to rotate according to the joint rotation angle, and after the rotation is completed, a second start instruction is sent to the hand control module.
[0085] Optionally, the arm control module responds to the first start instruction sent by the main control module, calculates and predicts the joint rotation angle of the arm of the humanoid robot based on the first position information and a pre-built inverse kinematics model, and sends the joint rotation angle to the joint controller corresponding to the arm, so that the arm of the humanoid robot rotates according to the joint rotation angle, and sends a second start instruction to the controlled module after the rotation is completed.
[0086] In one possible implementation, Figure 8 A schematic diagram of a flow chart of controlling the hand movement of a humanoid robot in a humanoid robot control method provided in an embodiment of the present application, referring to Figure 8 As shown, the above S104 controls the hand movement of the humanoid robot according to the hand state information, including: S801. Respond to the second start instruction and obtain the current state of the hand of the humanoid robot.
[0087] Optionally, the hand control module may respond to the second start instruction, obtain the current state of the humanoid robot in real time through an encoder, IMU, etc., and determine the current state of the hand of the humanoid robot.
[0088] S802: Determine whether the current state is consistent with the hand state information; if not, adjust the state of the hand of the humanoid robot to be consistent with the hand state information.
[0089] Optionally, after determining the current state of the hand of the humanoid robot, the hand control module may determine whether the current state of the hand is consistent with the hand state information, and if so, maintain the current state of the hand.
[0090] Optionally, if not, the state of the hand of the humanoid robot can be adjusted to be consistent with the hand state information by controlling the joint driver corresponding to the hand.
[0091] By controlling the main body to move first, and then controlling the arms after the legs have finished moving, and finally controlling the hands, this control method from lower limbs to upper limbs to the end can ensure that the humanoid robot maintains balance and stability during movement, provide a good platform for subsequent arm and hand operations, and enable the humanoid robot to gradually enter the task state. It can also decompose complex control tasks into simpler subtasks, thereby reducing the overall control difficulty. In addition, it can ensure that the robot will not suddenly make unpredictable movements when performing operations, thereby avoiding potential collisions or injuries. It can also optimize energy consumption.
[0092] Based on the same inventive concept, a humanoid robot control device corresponding to the humanoid robot control method is also provided in the embodiment of the present application. Since the principle of solving the problem by the device in the embodiment of the present application is similar to the above-mentioned humanoid robot control method in the embodiment of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0093] Fig. 9 A schematic diagram of a humanoid robot control system provided in an embodiment of the present application, referring to Fig. 9 As shown, the humanoid robot control system includes: a visual language perception model, a task planner based on a large language model and a motion control system, and the motion control system includes: an arm control module, a hand control module and a main body control module; a reinforcement learning strategy is deployed in the main body control module, and the reinforcement learning strategy is obtained through pre-training. The humanoid robot control system is used to execute the steps of the above-mentioned humanoid robot control method.
[0094] For the description of the visual language perception model, the task planner based on the large language model, and the processing flow of the motion control system in the humanoid robot control system, and the interaction flow between the visual language perception model, the task planner based on the large language model, and the motion control system, please refer to the relevant instructions in the above method embodiments and will not be described in detail here.
[0095] The present application also provides a humanoid robot, which includes the above-mentioned humanoid robot control system, and the humanoid robot can execute user instructions through interaction with the user.
[0096] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, the specific working process of the system and device described above can refer to the corresponding process in the method embodiment, and will not be repeated in this application. In the several embodiments provided in this application, it should be understood that the disclosed system, device and method can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0097] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or part of the technical solution that contributes to the prior art or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, RandomAccess Memory), disk or optical disk and other media that can store program code.
[0098] The above are only specific implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be covered by the protection scope of the present application.
Claims
1. A method for controlling a humanoid robot, characterized in that: include: Acquire a task instruction and a field of view image, wherein the task instruction is used to instruct the humanoid robot to perform an action on a target object; Inputting the visual field image and the task instruction into a pre-trained visual language perception model, and having the visual language perception model perceive and process the visual field image and the task instruction to obtain the three-dimensional coordinates of the target object in the visual field image; Obtain a prompt word, and input the prompt word and the three-dimensional coordinates into a pre-trained task planner based on a large language model, wherein the task planner performs task planning based on the prompt word and the three-dimensional coordinates, and determines control information corresponding to the prompt word, wherein the control information includes first position information, second position information, and hand state information, and the prompt word at least includes the task instruction; The arm movement of the humanoid robot is controlled by an arm control module in the motion control system according to the first position information, and the hand movement of the humanoid robot is controlled by a hand control module in the motion control system according to the hand state information, and the leg movement of the humanoid robot is controlled by a body control module in the motion control system according to a pre-deployed reinforcement learning strategy and the second position information, and the reinforcement learning strategy is obtained through pre-training.
2. The humanoid robot control method according to claim 1, characterized in that: The visual language perception model includes an encoding module and a decoding module connected in sequence; The visual language perception model performs perception processing on the field of view image and the task instruction to obtain the three-dimensional coordinates of the target object in the field of view image, including: Inputting the visual field image and the task instruction into the encoding module, and the encoding module encodes the visual field image and the task instruction to obtain encoded features; The encoded features are input into the decoding module, and the decoding module decodes the encoded features to obtain the three-dimensional coordinates.
3. The humanoid robot control method according to claim 2, characterized in that: The encoding module includes: an image encoding module, a text encoding module and an encoding fusion module; The inputting the field of view image and the task instruction into the encoding module, and the encoding module encoding the field of view image and the task instruction to obtain the encoded features, includes: Inputting the field of view image into the image encoding module, and the image encoding module performs image perception encoding processing on the field of view image to obtain image features; Inputting the task instruction into the text encoding module, and the text encoding module performs text-aware encoding processing on the task instruction to obtain text features; The image features and the text features are input into the encoding fusion module, and the encoding fusion module performs feature fusion on the image features and the text features to obtain the encoded features.
4. The humanoid robot control method according to claim 1, characterized in that: The prompt word also includes: pre-configured configuration information, the configuration information includes: skill configuration information, operation space constraint configuration information, safety constraint configuration information and kinematic configuration information; The skill configuration information is used to indicate executable actions of the arms, hands and legs of the humanoid robot, the operation space constraint configuration information is used to indicate the operable space of the humanoid robot, and the kinematic configuration information is used to indicate the prior kinematic information of the humanoid robot.
5. The humanoid robot control method according to claim 1, characterized in that: The task planner performs task planning based on the prompt word and the three-dimensional coordinates, and determines the control information corresponding to the prompt word, including: Obtaining the current height of the humanoid robot; The task planner performs task planning based on the prompt word, the current height and the three-dimensional coordinates, and determines control information corresponding to the prompt word.
6. The humanoid robot control method according to claim 1, characterized in that: The controlling the leg movement of the humanoid robot according to the reinforcement learning strategy and the second position information comprises: Generate a target action sequence according to the reinforcement learning strategy and the second position information; The legs of the humanoid robot are controlled to move according to the target action sequence, and when the legs move to the target position indicated by the second position information, a first start instruction is sent to the arm control module.
7. The humanoid robot control method according to claim 6, characterized in that: The controlling the movement of the arm of the humanoid robot according to the first position information comprises: In response to the first start instruction, based on the first position information, a pre-built inverse kinematics model is used to predict the joint rotation angle of the arm of the humanoid robot, and the arm of the humanoid robot is controlled to rotate according to the joint rotation angle, and after the rotation is completed, a second start instruction is sent to the hand control module.
8. The humanoid robot control method according to claim 7, characterized in that: The controlling the hand movement of the humanoid robot according to the hand state information comprises: In response to the second start instruction, obtaining a current state of the hand of the humanoid robot; Determine whether the current state is consistent with the hand state information, and if not, adjust the state of the hand of the humanoid robot to be consistent with the hand state information.
9. A humanoid robot control system, characterized in that: include: A visual language perception model, a task planner based on a large language model, and a motion control system, wherein the motion control system includes: an arm control module, a hand control module, and a body control module; A reinforcement learning strategy is deployed in the main control module, and the reinforcement learning strategy is obtained through pre-training. The humanoid robot control system is used to execute the humanoid robot control method described in any one of claims 1-8.
10. A humanoid robot, characterized in that: include: The humanoid robot control system as claimed in claim 9.
Citation Information
Patent Citations
Robot instruction operation method and system based on natural language and medium
CN116690616A
Robot manipulation method based on visual language large model
CN118559711A
Multi-modal sensing humanoid robot action self-adaptive control method and multi-modal sensing humanoid robot action self-adaptive control system
CN119610112A
Desktop-level multi-task mechanical arm control method and related device
CN119704175A
Multi-robot collaborative navigation method and system based on visual language large model
CN119756375A
Cited By
Robot control model training method, robot control model control method, robot control model training device, robot control model control device and electronic equipment
CN120620237A
Robot control model training method, control method, device and electronic equipment
CN120620237B
Flat part sorting method and device, electronic equipment and storage medium
CN121573430A