Humanoid robot double-arm cooperation method, device, electronic device and storage medium

By obtaining the surrounding environment and double arms status information of the humanoid robot, building a neural network model, generating control instructions and predicting trajectory points, the efficiency and accuracy of the coordinated operation of the humanoid robot with both arms is solved, and stable operation and efficient coordination are achieved in complex environments.

CN119610085BActive Publication Date: 2025-07-01广州里工实业有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411598808.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-07-01
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

The prior art is difficult to achieve efficient and accurate coordinated operation of humanoid robot arms in complex environments, and lacks an end-to-end solution from information acquisition to control instruction generation.

Method used

By obtaining the feature vectors of the surrounding environment information of the humanoid robot and the joint state information of the two arms, a neural network model is built, including the motion branch model and the trajectory branch model, a control command and predict the trajectory point are generated, and the coordinated control of the two arms is realized.

Benefits of technology

The efficiency and accuracy of the humanoid robot's two-arm operation are improved, allowing it to operate stably in complex environments and task scenarios, and have better generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119610085B_ABST
    Figure CN119610085B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, electronic device and storage medium for the collaborative operation of the two arms of a humanoid robot, relating to the technical field of robots. The method includes: obtaining a target feature vector, where the target feature vector includes surrounding environment information and the joint state information of the two arms of the humanoid robot; constructing a neural network model and then respectively constructing a motion branch model and a trajectory branch model based on the neural network model; inputting the target feature vector into the motion branch model and the trajectory branch model respectively to obtain the control instructions for the two arms of the humanoid robot output by the motion branch model and the predicted trajectory points for the two arms of the humanoid robot output by the trajectory branch model; generating a control strategy according to the control instructions and the predicted trajectory points; and controlling the two arms of the humanoid robot to move according to the control strategy. By fusing the surrounding environment information and the joint state information of the two arms, the present application can more comprehensively and accurately perceive the environment and the state of the two arms, enabling the humanoid robot to better adapt to complex environments and tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robot technology, and particularly to a method, device, electronic device and storage medium for collaborative operation of the arms of a humanoid robot. Background Art

[0002] With the development of humanoid robot technology, collaborative operation of the arms has become an important research direction. At present, there are some problems in the operation of the arms of humanoid robots. On the one hand, in terms of information acquisition, it is difficult to obtain comprehensive state information of humanoid robots. On the other hand, in terms of control strategies, traditional solutions often have difficulty achieving efficient collaboration of the arms and lack an effective end-to-end solution from information acquisition to control instruction generation. Summary of the Invention

[0003] The main purpose of the embodiments of this application is to propose a method, device, electronic device and storage medium for collaborative operation of the arms of a humanoid robot, so as to achieve efficient and accurate collaborative operation of the arms of a humanoid robot for different tasks in a complex environment, and improve the intelligence and adaptability of the humanoid robot.

[0004] To achieve the above purpose, on the one hand, the embodiments of this application propose a method for collaborative operation of the arms of a humanoid robot, and the method includes the following steps:

[0005] Obtain a target feature vector; wherein, the target feature vector includes a feature vector corresponding to the surrounding environment information of the humanoid robot and a feature vector corresponding to the joint state information of the arms of the humanoid robot;

[0006] Construct a neural network model and then respectively construct a motion branch model and a trajectory branch model based on the neural network model;

[0007] Input the target feature vector into the motion branch model and the trajectory branch model respectively, and obtain the control instructions of the arms of the humanoid robot output by the motion branch model and the predicted trajectory points of the arms of the humanoid robot output by the trajectory branch model;

[0008] Generate a control strategy according to the control instructions and the predicted trajectory points;

[0009] Control the arms of the humanoid robot to move according to the control strategy.

[0010] In some embodiments, the obtaining of the target feature vector includes the following steps:

[0011] Obtain a time-series image of the surrounding environment information of the humanoid robot through a sensor arranged on the head of the humanoid robot; convert the time-series image into a first branch feature vector through a depth convolutional neural network encoder based on visual perception;

[0012] Obtain the joint state information of the left arm and the right arm of the humanoid robot through sensors arranged on the two arms of the humanoid robot; convert the joint state information into a second branch feature vector corresponding to the left arm and a second branch feature vector corresponding to the right arm through a multi-layer perceptron network encoder based on joint state perception;

[0013] Fuse the first branch feature vector and the two second branch feature vectors to obtain the target feature vector.

[0014] In some embodiments, constructing the neural network model includes the following steps:

[0015] Construct multiple network layers and multiple different types of neurons; wherein, the types of the neurons include sensory neurons, internal neurons, command neurons, and motor neurons;

[0016] Configure each of the neurons into each of the network layers;

[0017] In any two consecutive network layers, insert synapses for any source neuron and connect the synapses to target neurons to obtain the neural network model; wherein, the source neuron is the neuron that is the starting point of information transmission, and the target neuron is the neuron that is the ending point of information transmission.

[0018] In some embodiments, the step of inserting synapses for any source neuron and connecting the synapses to target neurons in any two consecutive network layers to obtain the neural network model includes the following steps:

[0019] In any two consecutive network layers, insert synapses for any source neuron whose number and polarity satisfy the Bernoulli distribution, and randomly connect the synapses to the target neurons according to the binomial distribution to obtain the neural network model.

[0020] In some embodiments, the step of inserting synapses for any source neuron and connecting the synapses to target neurons in any two consecutive network layers to obtain the neural network model includes the following steps:

[0021] In any two consecutive network layers, set the command neurons to be circularly connected and insert a preset number of synapses between them to obtain the neural network model.

[0022] In some embodiments, constructing the motion branch model and the trajectory branch model based on the neural network model includes the following steps:

[0023] Determine the motion branch loss function based on the difference between the predicted control distribution and the expert distribution, and the feature loss between the predicted control distribution and the supervision signal and the prediction signal at the current moment of the expert data; construct the motion branch model according to the neural network model and the motion branch loss function;

[0024] Determine the trajectory branch loss function based on the difference between the predicted trajectory points and the true trajectory points, and the feature loss between the predicted trajectory points and the supervision signal and the prediction signal at the current moment of the expert data; construct the trajectory branch model according to the neural network model and the trajectory branch loss function.

[0025] In some embodiments, the generating the control strategy according to the control instruction and the predicted trajectory points includes the following steps:

[0026] Perform weighted summation on the control instruction and the predicted trajectory points to generate the control strategy.

[0027] To achieve the above object, on the other hand, an embodiment of the present application proposes a humanoid robot dual-arm collaboration device, and the device includes:

[0028] A feature vector acquisition unit for acquiring a target feature vector; wherein, the target feature vector includes a feature vector corresponding to the surrounding environment information of the humanoid robot and a feature vector corresponding to the joint state information of the two arms of the humanoid robot;

[0029] A model construction unit for constructing a neural network model and then constructing a motion branch model and a trajectory branch model based on the neural network model respectively;

[0030] An output prediction unit for inputting the target feature vector into the motion branch model and the trajectory branch model respectively, and obtaining the control instruction of the two arms of the humanoid robot output by the motion branch model and the predicted trajectory points of the two arms of the humanoid robot output by the trajectory branch model;

[0031] A control strategy generation unit for generating a control strategy according to the control instruction and the predicted trajectory points;

[0032] A dual-arm driving unit for controlling the two arms of the humanoid robot to move according to the control strategy.

[0033] To achieve the above object, on the other hand, an embodiment of the present application proposes an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the above method is implemented.

[0034] To achieve the above object, another aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the above method.

[0035] The embodiments of the present application at least include the following beneficial effects:

[0036] The present application can obtain a target feature vector, where the target feature vector includes a feature vector corresponding to the surrounding environment information of the humanoid robot and a feature vector corresponding to the joint state information of the two arms of the humanoid robot; construct a neural network model and then respectively construct a motion branch model and a trajectory branch model based on the neural network model; input the target feature vector into the motion branch model and the trajectory branch model respectively to obtain the control instructions for the two arms of the humanoid robot output by the motion branch model and the predicted trajectory points of the two arms of the humanoid robot output by the trajectory branch model; generate a control strategy according to the control instructions and the predicted trajectory points; control the two arms of the humanoid robot to move according to the control strategy. By fusing the surrounding environment information and the joint state information of the two arms, the present application overcomes the limitations of a single sensor, more comprehensively and accurately senses the environment and the state of the two arms, and provides a reliable basis for subsequent operations; the motion branch model and the trajectory branch model constructed based on the neural network model can effectively realize the coordinated operation of the two arms, improve the efficiency and accuracy of the two-arm operation of the humanoid robot, and enable the humanoid robot to better adapt to complex environments and tasks; through the end-to-end solution from the target feature vector to the control strategy, combined with the characteristics of the humanoid robot itself for optimization, the humanoid robot has better generalization ability and can operate stably in different environment and task scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.

[0038] Figure 1 It is a schematic flowchart of the method for the coordinated operation of the two arms of the humanoid robot provided by the embodiments of the present application;

[0039] Figure 2 It is an example flowchart of the method for the coordinated operation of the two arms of the humanoid robot provided by the embodiments of the present application;

[0040] Figure 3 It is an example flowchart of a more specific method for the coordinated operation of the two arms of the humanoid robot provided by the embodiments of the present application;

[0041] Figure 4It is a schematic flowchart of an example architecture of the humanoid robot's dual-arm collaboration method provided by an embodiment of the present application;

[0042] Figure 5 It is a schematic flowchart of an example of information fusion provided by an embodiment of the present application;

[0043] Figure 6 It is a schematic diagram of the connection of each neuron in the neural network model provided by an embodiment of the present application;

[0044] Figure 7 It is a flowchart for calculating the motion branch loss function provided by an embodiment of the present application;

[0045] Figure 8 It is a flowchart for calculating the trajectory branch loss function provided by an embodiment of the present application;

[0046] Figure 9 It is a schematic flowchart of an example of the humanoid robot's dual-arm collaborative operation provided by an embodiment of the present application;

[0047] Figure 10 It is a schematic structural diagram of the humanoid robot's dual-arm collaboration device provided by an embodiment of the present application;

[0048] Figure 11 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0049] In order to make the purpose, technical solutions and advantages of the present application clearer, the following further details the present application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. When the following description involves the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of the present application. They are only examples of devices and methods that are consistent with some aspects of the embodiments of the present application as detailed in the appended claims.

[0050] It can be understood that the terms "first", "second", etc. used in the present application can be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information can also be called the second information, and similarly, the second information can also be called the first information. Depending on the context, the words "if", "when" as used herein can be interpreted as "when...", "while...", or "in response to determining".

[0051] The terms "at least one", "a plurality", "each", "any one", etc. used in this application, at least one includes one, two or more than two, a plurality includes two or more than two, each refers to each of the corresponding plurality, and any one refers to any one of the plurality.

[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0053] Referring to Figure 1 , embodiments of this application provide a method for collaborative operation of the arms of a humanoid robot. This method may include but is not limited to S100 to S140, specifically as follows:

[0054] S100: Obtain a target feature vector; wherein, the target feature vector includes a feature vector corresponding to the surrounding environment information of the humanoid robot and a feature vector corresponding to the joint state information of the arms of the humanoid robot.

[0055] Further, S100 may include S101 to S103:

[0056] S101: Obtain a time-series image of the surrounding environment information of the humanoid robot through a sensor disposed on the head of the humanoid robot; convert the time-series image into a first branch feature vector through a depth convolutional neural network encoder based on visual perception;

[0057] S102: Obtain the joint state information of the left arm and the right arm of the humanoid robot through sensors disposed on the arms of the humanoid robot; convert the joint state information into a second branch feature vector corresponding to the left arm and a second branch feature vector corresponding to the right arm through a multi-layer perceptron network encoder based on joint state perception;

[0058] S103: Fuse the first branch feature vector and the two second branch feature vectors to obtain the target feature vector.

[0059] Specifically, taking the time-series image obtained by the visual sensor on the head of the humanoid robot as input, a first branch feature vector is obtained through a depth convolutional neural network encoder based on visual perception (VPCNN-Encoder); the time-series image obtained by the visual sensor includes information such as the position, shape, and color of objects in the environment.

[0060] Taking the sensor measurement information (including joint angles, angular velocities, torques, etc.) of each joint of the left and right arms of the humanoid robot as inputs, the second branch feature vectors (one for each of the left and right arms) are obtained through a multi-layer perceptron network encoder based on joint state perception (JSP-MLP-Encoder);

[0061] The first branch feature vector is fused with the second branch feature vectors of the left and right arms to obtain the effective information of the environment and the states of the two arms, that is, the target feature vector. When fusing each feature vector, it is necessary to consider how to reasonably integrate visual information and the state information of the two arms so as to provide comprehensive and accurate inputs for subsequent operations.

[0062] More specifically, the humanoid robot obtains the sequential images of the environment through the head visual sensor, and the first branch feature vector is obtained after being processed by a deep convolutional neural network encoder for visual perception (VPCNN-Encoder). This vector contains information such as the positions and shapes of objects in the environment.

[0063] Meanwhile, the sensor measurement information (including joint angles, angular velocities, torques, etc.) of each joint of the left and right arms is respectively obtained, and the second branch feature vectors of the left and right arms are obtained through a multi-layer perceptron network encoder based on joint state perception (JSP-MLP-Encoder).

[0064] The first branch feature vector is fused with the second branch feature vectors of the left and right arms to obtain the effective information of the environment and the states of the two arms. In the fusion process, the importance of different information needs to be considered, and methods such as weighted summation can be used to ensure the comprehensiveness and accuracy of the information.

[0065] S110: Construct a neural network model and then construct a motion branch model and a trajectory branch model based on the neural network model respectively.

[0066] Furthermore, the step of constructing the neural network model in S110 may include S111~S113:

[0067] S111: Construct multiple network layers and multiple different types of neurons; among them, the types of the neurons include sensory neurons, internal neurons, command neurons, and motor neurons;

[0068] S112: Configure each of the neurons into each of the network layers;

[0069] S113: In any two consecutive network layers, insert synapses for any source neuron and connect the synapses to the target neuron to obtain the neural network model; where the source neuron is the neuron that is the starting point of information sending, and the target neuron is the neuron that is the ending point of information sending.

[0070] Specifically, in this embodiment, a neural network model suitable for the cooperative operation of the two arms of a humanoid robot can be constructed, including multiple layers of neurons. Among them, sensory neurons are used to sense environmental and two-arm state information, internal neurons are used for information processing and transmission, command neurons are used to generate control commands for the cooperative operation of the two arms, and motor neurons are used to convert control commands into specific joint movements; between any two consecutive layers, for any source neuron, synapses are inserted according to certain rules, and the number and polarity of the synapses satisfy a specific distribution, and the target neurons are randomly selected through the corresponding distribution.

[0071] Furthermore, S113 may include: in any two consecutive network layers, insert the synapses with the number and polarity satisfying the Bernoulli distribution for any source neuron, and randomly connect the synapses to the target neurons according to the binomial distribution to obtain the neural network model.

[0072] As another implementation manner, the step of constructing the neural network model in S110 may include S114:

[0073] S114: In any two consecutive network layers, set the command neurons to be circularly connected and insert a preset number of the synapses to obtain the neural network model.

[0074] It can be understood that both the source neurons and the target neurons in this embodiment are command neurons. Circular connections can be set between the command neurons and an appropriate number of synapses are inserted, and their parameters satisfy the corresponding rules.

[0075] Further, constructing the motion branch model and the trajectory branch model based on the neural network model in S110 includes S115 - S116:

[0076] S115: Determine the motion branch loss function according to the difference between the predicted control distribution and the expert distribution and the feature loss between the predicted control distribution and the supervision signal and the prediction signal of the expert data at the current moment; construct the motion branch model according to the neural network model and the motion branch loss function;

[0077] S116: Determine the trajectory branch loss function according to the difference between the predicted trajectory points and the real trajectory points and the feature loss between the predicted trajectory points and the supervision signal and the prediction signal of the expert data at the current moment; construct the trajectory branch model according to the neural network model and the trajectory branch loss function.

[0078] Specifically, the motion branch loss function and the trajectory branch loss function are respectively defined, and the neural network model of this embodiment is used for modeling. The motion branch loss function considers the difference between the predicted control distribution and the expert distribution, as well as the feature loss between the supervision signal and the prediction signal at the current moment of the expert data. The trajectory branch loss function considers the difference between the predicted trajectory points and the true trajectory points, as well as the feature loss between the supervision signal and the prediction signal at the current moment of the expert data.

[0079] More specifically, the motion branch loss function considers the difference (such as KL divergence) between the predicted control distribution and the expert distribution (which can be obtained from human dual-arm collaborative operation data), as well as the feature loss (which can be measured by calculating the L2 measure between feature vectors) between the supervision signal and the prediction signal at the current moment of the expert data. The trajectory branch loss function considers the difference (such as spatial position difference, etc.) between the predicted trajectory points (the predicted trajectory points of the end effectors or joints of the two arms) and the true trajectory points, as well as the feature loss (measured by calculating the L2 measure between feature vectors) between the supervision signal and the prediction signal at the current moment of the expert data.

[0080] S120: Input the target feature vector into the motion branch model and the trajectory branch model respectively, and obtain the control instructions of the two arms of the humanoid robot output by the motion branch model and the predicted trajectory points of the two arms of the humanoid robot output by the trajectory branch model.

[0081] Specifically, input the effective information integrating the environment and the states of the two arms into the motion branch model and the trajectory branch model at the same time; the motion branch model outputs the control instructions of the two arms of the humanoid robot, including the motion angles, speeds, torques, etc. of each joint of the left and right arms; the trajectory branch model outputs the predicted trajectory points of the two arms of the humanoid robot at different future moments, such as the predicted trajectory points of the end effectors of the two arms at moments of 1 second, 2 seconds, 3 seconds, etc. in the future, providing a reference for the motion planning of the two arms.

[0082] S130: Generate a control strategy according to the control instructions and the predicted trajectory points.

[0083] Specifically, based on the control instructions of the two arms of the humanoid robot and the predicted trajectory points at different future moments, a final control strategy of the two arms of the humanoid robot is obtained through a fusion mechanism to achieve the collaborative and accurate operation of the two arms.

[0084] Further, S130 may include S131:

[0085] S131: Perform weighted summation on the control instructions and the predicted trajectory points to generate the control strategy.

[0086] Exemplarily, different weights can be assigned according to the different motion requirements and importance of the two arms, and then the control instructions and predicted trajectory points can be fused by means of weighted summation, etc., to ensure the accurate cooperation of the two arms and enable the robot to better complete various tasks, such as grasping, assembling, etc.

[0087] S140: Control the two arms of the humanoid robot to move according to the control strategy.

[0088] It can be understood that in this embodiment, the two arms of the humanoid robot can be controlled to move according to the control strategy through the control of the humanoid robot.

[0089] Next, specific application examples will be combined to introduce and illustrate the solutions of the embodiments of the present application in detail:

[0090] Specifically, referring to Figure 2 、 Figure 3 and Figure 4 , this embodiment may include the following solutions:

[0091] (1) Information extraction.

[0092] 1.1 Visual information processing.

[0093] The visual sensor on the head of the humanoid robot continuously acquires the sequential images of the environment, and the sequential images contain information such as the positions and shapes of the objects in the environment. These images are input into the deep convolutional neural network encoder based on visual perception (VPCNN-Encoder), and the first branch feature vector is obtained through encoding processing.

[0094] 1.2 Acquisition of the state information of the two arms.

[0095] The sensors of each joint of the left and right arms of the humanoid robot continuously measure information such as the angles, angular velocities, and torques of the joints. The sensor measurement information of the left and right arms is respectively input into the multi-layer perceptron network encoder based on joint state perception (JSP-MLP-Encoder), and the second branch feature vectors of the left and right arms are obtained through encoding processing.

[0096] 1.3 Information fusion.

[0097] The first branch feature vector is fused with the second branch feature vectors of the left and right arms. In the fusion process, it is necessary to consider how to reasonably integrate the visual information and the state information of the two arms. For example, different weights can be assigned according to the importance of different information, and then fusion can be performed by means of weighted summation, etc., to obtain the effective information of the environment and the state of the two arms. Among them, Figure 5 is an example flowchart of information fusion.

[0098] (2) Neural network model construction.

[0099] 2.1 Neuron Definition and Connection.

[0100] The constructed neural network model includes sensory neurons, internal neurons, instruction neurons, and motor neurons. Sensory neurons are used to sense the effective information of the environment and the state of the two arms. Internal neurons are used for information processing and transmission. Instruction neurons are used to generate control instructions for the coordinated operation of the two arms. Motor neurons are used to convert the control instructions into specific executable joint movements. Exemplarily, Figure 6 FIG. is a connection example diagram of each neuron in the neural network model provided in this embodiment.

[0101] Next, each neuron will be specifically described.

[0102] Sensory neurons:

[0103] 1) Environmental information perception:

[0104] Receive the environmental-related part of the effective information of the environment and the state of the two arms from the fusion unit. For visual information, it parses the feature vector processed by the depth convolutional neural network encoder based on visual perception (VPCNN-Encoder), and extracts information such as the position, shape, color, texture, etc. of the objects in the environment, as well as the spatial relationship information between the objects. For example, it can identify the position and shape of obstacles in the environment and determine whether they are within the working area of the robot's two arms.

[0105] The perception of the two-arm state information is mainly auxiliary. It obtains the encoded features of the approximate position and posture information of the two arms, and these information together with the environmental information constitute a preliminary perception of the overall scene.

[0106] 2) Two-arm state information perception:

[0107] From the two-arm state information obtained from the fusion unit, the sensory neurons can identify the approximate features of the initial position and posture of the two arms. For example, it can judge whether the two arms are extended or bent, and the approximate angle range, etc., providing a basis for subsequent precise processing.

[0108] Internal neurons:

[0109] 1) Information integration and preprocessing:

[0110] Receive the environmental and two-arm state information from the sensory neurons. For environmental information, it integrates the feature information of different objects. For example, it combines the position and shape information of multiple objects to construct a more complete environmental scene model. At the same time, it further refines the two-arm state information, and combines the environmental information to analyze whether the relative position and posture of the two arms in the current environment are reasonable.

[0111] Internal neurons also preprocess the information transmitted by sensory neurons, such as performing operations like data normalization and noise removal, to improve the quality and accuracy of the information and provide better input for the processing of command neurons.

[0112] 2) Feature extraction and abstraction:

[0113] Extract higher-level features from the integrated information. For environmental information, it may extract some representative environmental pattern features, such as spatial layout patterns, object combination patterns, etc. For the information of the two-arm state, it can extract the trend features of the two-arm movement, such as whether the two arms tend to approach or move away from each other, and the speed trend of the movement. These features will help the command neurons generate more reasonable control commands.

[0114] Command neurons:

[0115] 1) Basis for generating control commands:

[0116] Receive the information processed by internal neurons. It first judges the goals and requirements of the current task based on the environmental information and the information of the two-arm state. For example, if the task is to grasp an object, it needs to determine information such as the position, shape, size of the object, and the distance and angle relationship between the two arms and the object.

[0117] At the same time, the command neurons will consider the current state of the two arms, such as joint angles, angular velocities, torques, etc., as well as the movement capabilities and limitations of the two arms to ensure that the generated control commands are feasible.

[0118] 2) Generation of two-arm coordinated control commands:

[0119] Based on the above information, the command neurons generate control commands for the two-arm coordinated operation. For the left arm, it will determine the parameters such as the angles, speeds, and torques that each joint needs to adjust to achieve the correct movement of the left arm. For the right arm, corresponding control commands will also be generated. These control commands will consider the coordination relationship between the two arms, such as the movement sequence of the two arms and the requirements for synchronism. For example, if it is to grasp a large object, it may be necessary for the left arm to move to a specific position first, and then the right arm to cooperate. The command neurons will accurately generate such a sequence of control commands.

[0120] Motor neurons:

[0121] Convert commands into joint movements: Receive the control commands from the command neurons. For the motor neurons of the left arm, it converts the parameters such as the angles, speeds, and torques of each joint of the left arm in the command into specific electrical signals or mechanical movement commands, directly driving the actuators of the left arm joints to make the left arm move according to the command requirements.

[0122] Similarly, for the motor neurons of the right arm, they convert the control instructions of the right arm into actual joint movements, ensuring that the right arm can accurately execute the instructions and achieve coordinated operation with the left arm. For example, if the instruction requires the elbow joint of the right arm to bend at a certain angle, the motor neurons will convert this angle instruction into a signal that the elbow joint actuator can recognize, causing the elbow joint to bend accurately to the specified angle.

[0123] Between any two consecutive layers, for any source neuron, synapses are inserted according to certain rules. For example, the number and polarity of the synapses satisfy the Bernoulli distribution, and the target neurons are randomly selected through the binomial distribution. Recurrent connections can be set between the instruction neurons, and an appropriate number of synapses are inserted.

[0124] Next, an explanation of the source neuron is given:

[0125] The source neuron is a concept involved when describing the connection relationship between any two consecutive layers in a neural network. It refers to the neuron that serves as the starting point for information transmission in a certain layer.

[0126] Differences between the source neuron and other neurons:

[0127] 1) With sensory neurons:

[0128] The main function of sensory neurons is to sense and preliminarily process environmental and bilateral arm state information, which is the first processing link for information to enter the neural network. The source neuron is a more general concept used to describe the information sending end when connecting layers. It can appear between any two consecutive layers of the neural network, not limited to the layer where the sensory neurons are located. For example, when transmitting information from the sensory neuron layer to the internal neuron layer, the sensory neurons can act as source neurons, but when the internal neuron layer transmits information to other layers, the internal neurons can also become source neurons.

[0129] 2) With internal neurons:

[0130] Internal neurons focus on information integration, preprocessing, feature extraction, and abstraction processes. The source neuron emphasizes the starting point of information transmission. When an internal neuron serves as the starting point for information transmission and sends information to other layers, it becomes a source neuron, which is defined from the perspective of information flow direction and is different from the focus of the functional processing of internal neurons themselves.

[0131] 3) With instruction neurons:

[0132] The command neurons are mainly responsible for generating control commands for the coordinated operation of the two arms. The source neuron is a concept in describing the connection structure of the neural network. When a command neuron transmits information to a motor neuron or other relevant layers, the command neuron can act as a source neuron, which is a different dimension from the control command generation function of the command neuron itself. One is about functional processing, and the other is about structural connection.

[0133] 4) With motor neurons:

[0134] The role of motor neurons is to convert control commands into specific executable joint movements. The source neuron is used to illustrate the starting situation of information transmission in the neural network. Motor neurons can act as source neurons in some cases, such as when the output of motor neurons is transmitted as feedback information to other layers for further processing, but this is different from its main function of driving joint movements.

[0135] Between any two consecutive layers, for any source neuron, synapses are inserted according to certain rules. For example, the number and polarity of synapses satisfy the Bernoulli distribution, and target neurons are randomly selected through the binomial distribution. Recurrent connections can be set between command neurons, and appropriate specific synaptic parameters (such as the number satisfying certain conditions and the polarity satisfying the Bernoulli distribution) are inserted to achieve information transmission and interaction between command neurons.

[0136] Next, the role of inserting synapses is described.

[0137] 1) Information transmission:

[0138] Synapses are the key structures for information transmission between neurons. In the neural network, inserting synapses enables the source neuron to transmit the electrical or chemical signals it generates (in biological neural networks, they are electrical impulses and neurotransmitters, analogized here as processed information) to the target neuron. For example, sensory neurons transmit the environmental and two-arm state information they perceive to internal neurons through synapses for further processing and integration.

[0139] 2) Constructing the connection structure of the neural network:

[0140] The insertion method of synapses determines the connection topology of the neural network. Different neurons are connected by inserting synapses according to specific rules, thus forming a complex network structure. This structure has a crucial impact on the way information flows and is processed in the network. For example, reasonable synaptic connections can enable the neural network to effectively learn and recognize different input patterns, just like in the coordinated operation of the two arms of a humanoid robot, it can accurately generate appropriate control commands and predict trajectory points according to environmental and two-arm state information.

[0141] Next, the role of the number and polarity of synapses satisfying the Bernoulli distribution will be described.

[0142] 1) Controlling the complexity and diversity of information transmission:

[0143] The Bernoulli distribution is a discrete probability distribution. The fact that the number and polarity of synapses satisfy the Bernoulli distribution means that the number and polarity of synapses are randomly determined but follow certain probability rules. This randomness increases the complexity and diversity of the connections in the neural network. Different combinations of the number and polarity of synapses will result in different information transmission paths and intensities, enabling the neural network to learn and process a wider range of input patterns. For example, in the cooperative operation of the two arms of a humanoid robot, in the face of different environments and task scenarios, this complex and diverse connection method can enable the neural network to better adapt to and process various situations, rather than being limited to a fixed connection pattern.

[0144] 2) Simulating the characteristics of biological neural networks:

[0145] In biological neural networks, the connections between neurons are also complex and have a certain degree of randomness. The fact that the number and polarity of synapses satisfy the Bernoulli distribution can, to a certain extent, simulate this characteristic of biological neural networks, thereby possibly endowing artificial neural networks with some advantages similar to biological neural networks, such as better adaptability and learning ability. This helps the neural network for the cooperative operation of the two arms of a humanoid robot to better learn the patterns and skills of human two-arm cooperative operation, and improve the accuracy and flexibility of its operation.

[0146] Next, the meaning of the target neuron will be described.

[0147] Definition: The target neuron is a concept corresponding to the source neuron in the neural network. When the source neuron transmits information through synapses, the receiving neuron to which the information is directed is the target neuron.

[0148] 1) Position in the neural network structure:

[0149] The target neuron is located in the next layer of the layer where the source neuron is located (in the context of describing the transmission of information from one layer to the adjacent next layer). For example, when transmitting information from the sensory neuron layer to the internal neuron layer, if the sensory neuron is the source neuron, then some neurons in the internal neuron layer will become target neurons and receive information from the sensory neurons.

[0150] 2) The role of randomly selecting target neurons through the binomial distribution:

[0151] Increase the randomness and diversity of network connections: The binomial distribution is a discrete probability distribution. Randomly selecting target neurons through the binomial distribution makes the connections between source neurons and target neurons in the neural network random. This randomness increases the diversity of network connections and avoids fixed and single connection patterns. Different random connection combinations enable the neural network to explore more different information transmission paths and processing methods during the learning process. For example, in the neural network of a humanoid robot's dual-arm collaborative operation, this randomness can enable the network to better adapt to various complex environments and task situations and improve its generalization ability.

[0152] 3) Optimize information processing and learning ability:

[0153] Randomly selecting target neurons can prompt the neural network to discover more effective information processing methods during the learning process. Since the connections are random, the network will continuously adjust its own weights and connection relationships during training to adapt to different input and output requirements. This helps the network better learn the relationship between input information (such as environmental and dual-arm state information) and output information (such as control instructions and predicted trajectory points), thereby improving the performance and accuracy of the network. For example, when learning how to accurately control a humanoid robot's dual arms to perform a grasping task, random target neuron connections can enable the network to find the optimal control strategy faster.

[0154] 2.2 Define the loss function.

[0155] Motion branch loss function: Considering the difference between the predicted control distribution and the expert distribution (such as the expert distribution obtained from human dual-arm collaborative operation data), this difference can be measured by the KL divergence; at the same time, it includes the feature loss between the supervision signal and the prediction signal of the expert data at the current moment, and this loss can be measured by calculating the L2 measure between the feature vectors. Exemplarily, Figure 7 is the calculation flowchart of the motion branch loss function.

[0156] Trajectory branch loss function: Considering the difference between the predicted trajectory points (predicted trajectory points of the end effectors or joints of the dual arms) and the real trajectory points, such as spatial position differences, etc.; at the same time, it includes the feature loss between the supervision signal and the prediction signal of the expert data at the current moment, and this loss can be measured by calculating the L2 measure between the feature vectors. Exemplarily, Figure 8 is the calculation flowchart of the trajectory branch loss function.

[0157] (III) Information input and output.

[0158] 3.1 Input:

[0159] Simultaneously input the fused effective information of the environment and the dual-arm state into the motion branch model and the trajectory branch model.

[0160] 3.2 Output:

[0161] The motion branch model outputs control commands for the two arms of the humanoid robot, including motion angle, speed, torque and other commands for each joint of the left and right arms. For example, for the shoulder joint of the left arm, an angle value and an angular velocity value may be output; for the elbow joint of the right arm, a different angle value and angular velocity value may be output, etc.

[0162] The trajectory branch model outputs predicted trajectory points of the two arms of the humanoid robot at different future times, such as predicted trajectory points of the end effectors of the two arms at 1 second, 2 seconds, 3 seconds, etc. in the future. These predicted trajectory points can provide a reference for the motion planning of the two arms and help the robot better plan the motion path of the two arms.

[0163] (4) Generation of the final control strategy.

[0164] Based on the control commands of the two arms of the humanoid robot and the predicted trajectory points at different future times, a final control strategy is generated through a fusion mechanism. For example, a weighted average method can be adopted. Different weights are assigned according to the different motion requirements and importance of the two arms, and then the control commands and predicted trajectory points are fused by weighted summation and other methods to obtain the final control strategy. This fusion mechanism can ensure the coordinated and accurate operation of the two arms, enabling the robot to better complete various tasks, such as grasping, assembly and other tasks. Exemplarily, Figure 9 is an example flowchart of the coordinated operation of the two arms of the humanoid robot.

[0165] Referring to Figure 10 , the embodiments of the present application also provide a device for the coordinated operation of the two arms of the humanoid robot, which can implement the above-mentioned method for the coordinated operation of the two arms of the humanoid robot. The device includes:

[0166] A feature vector acquisition unit for acquiring a target feature vector; wherein, the target feature vector includes a feature vector corresponding to the surrounding environment information of the humanoid robot and a feature vector corresponding to the joint state information of the two arms of the humanoid robot;

[0167] A model construction unit for constructing a neural network model and then constructing a motion branch model and a trajectory branch model based on the neural network model;

[0168] An output prediction unit for respectively inputting the target feature vector into the motion branch model and the trajectory branch model to obtain the control commands of the two arms of the humanoid robot output by the motion branch model and the predicted trajectory points of the two arms of the humanoid robot output by the trajectory branch model;

[0169] A control strategy generation unit for generating a control strategy according to the control commands and the predicted trajectory points;

[0170] A two - arm drive unit for controlling the movement of the two arms of the humanoid robot according to the control strategy.

[0171] It can be understood that the content in the above - mentioned method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented by the device embodiments of the present application are the same as those of the above - mentioned method embodiments, and the beneficial effects achieved are also the same as those of the above - mentioned method embodiments.

[0172] An embodiment of the present application also provides an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above - mentioned knowledge extraction method is implemented. The electronic device can be any intelligent terminal including a tablet computer, an in - vehicle computer, etc.

[0173] It can be understood that the content in the above - mentioned method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented by the device embodiments of the present application are the same as those of the above - mentioned method embodiments, and the beneficial effects achieved are also the same as those of the above - mentioned method embodiments.

[0174] Please refer to Figure 11 , Figure 11 which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0175] A processor 1101, which can be implemented in a general - purpose CPU (Central Processing Unit), a micro - processor, an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0176] A memory 1102, which can be implemented in the form of a Read Only Memory (ROM), a static storage device, a dynamic storage device, or a Random Access Memory (RAM), etc. The memory 1102 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1102 and are called by the processor 1101 to execute the knowledge extraction method of the embodiments of the present application;

[0177] An input / output interface 1103 for implementing information input and output;

[0178] A communication interface 1104, which is used to implement the communication interaction between this device and other devices. It can achieve communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0179] A bus 1105, which transmits information between various components of the device (such as a processor 1101, a memory 1102, an input / output interface 1103, and a communication interface 1104);

[0180] Among them, the processor 1101, the memory 1102, the input / output interface 1103, and the communication interface 1104 achieve communication connections with each other inside the device through the bus 1105.

[0181] The embodiment of the present application also provides a computer-readable storage medium. This computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned knowledge extraction method is implemented.

[0182] It can be understood that the content in the above method embodiments is applicable to the embodiments of this storage medium. The functions specifically implemented by the embodiments of this storage medium are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0183] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and their combinations.

[0184] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0185] Those skilled in the art can understand that the technical solutions shown in the figure do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figure, or combine certain steps, or different steps.

[0186] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0187] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.

[0188] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0189] It should be understood that in this application, "at least one (item)" means one or more, and "multiple" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one)" or its similar expression below refers to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0190] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0191] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0192] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0193] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. And the aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks or optical discs and other various media that can store programs.

[0194] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, and thus do not limit the scope of rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of rights of the embodiments of the present application.

Claims

1. A humanoid robot dual-arm collaboration method, characterized in that: The method comprises the following steps: Acquire a target feature vector; wherein the target feature vector includes a feature vector corresponding to the surrounding environment information of the humanoid robot and a feature vector corresponding to the joint state information of both arms of the humanoid robot; Constructing a neural network model and then constructing a motion branch model and a trajectory branch model based on the neural network model; Inputting the target feature vector into the motion branch model and the trajectory branch model respectively, obtaining control instructions of the two arms of the humanoid robot output by the motion branch model and predicted trajectory points of the two arms of the humanoid robot output by the trajectory branch model; Generate a control strategy according to the control instruction and the predicted trajectory point; The arms of the humanoid robot are controlled to move according to the control strategy.

2. The humanoid robot dual-arm collaboration method according to claim 1, characterized in that: The step of obtaining the target feature vector comprises the following steps: Acquire a time-series image of the environment information around the humanoid robot through a sensor arranged on the head of the humanoid robot; convert the time-series image into a first branch feature vector through a deep convolutional neural network encoder based on visual perception; Acquire joint state information of the left arm and the right arm of the humanoid robot through sensors arranged on the arms of the humanoid robot; convert the joint state information into a second branch feature vector corresponding to the left arm and a second branch feature vector corresponding to the right arm through a multi-layer perceptron network encoder based on joint state perception; The first branch feature vector and the two second branch feature vectors are merged to obtain the target feature vector.

3. The humanoid robot dual-arm collaboration method according to claim 1, characterized in that: The construction of the neural network model comprises the following steps: Constructing multiple network layers and multiple different types of neurons; wherein the types of neurons include perception neurons, internal neurons, command neurons and motor neurons; Allocating each of the neurons to each of the network layers; In any two consecutive network layers, synapses are inserted into any source neurons, and the synapses are connected to target neurons to obtain the neural network model; wherein the source neuron is the neuron at the starting point of information transmission, and the target neuron is the neuron at the end point of information transmission.

4. The humanoid robot dual-arm coordination method according to claim 3, characterized in that: The method of inserting synapses into any source neurons in any two consecutive network layers and connecting the synapses to target neurons to obtain the neural network model comprises the following steps: In any two consecutive network layers, the synapses whose quantity and polarity satisfy the Bernoulli distribution are inserted into any source neurons, and the synapses are randomly connected to the target neurons according to the binomial distribution to obtain the neural network model.

5. The humanoid robot dual-arm coordination method according to claim 3, characterized in that: The method of inserting synapses into any source neurons in any two consecutive network layers and connecting the synapses to target neurons to obtain the neural network model comprises the following steps: In any two consecutive network layers, each of the instruction neurons is set to be cyclically connected and a preset number of synapses are inserted to obtain the neural network model.

6. The humanoid robot dual-arm coordination method according to claim 1, characterized in that: The step of constructing a motion branch model and a trajectory branch model based on the neural network model comprises the following steps: Determine the motion branch loss function according to the difference between the predicted control distribution and the expert distribution and the feature loss between the supervisory signal and the prediction signal of the predicted control distribution and the expert data at the current moment; construct the motion branch model according to the neural network model and the motion branch loss function; The trajectory branch loss function is determined according to the difference between the predicted trajectory point and the true trajectory point and the feature loss between the predicted trajectory point and the supervisory signal and the predicted signal of the expert data at the current moment; and the trajectory branch model is constructed according to the neural network model and the trajectory branch loss function.

7. The humanoid robot dual-arm collaboration method according to any one of claims 1 to 6, characterized in that: The generating of the control strategy according to the control instruction and the predicted trajectory point comprises the following steps: The control command and the predicted trajectory point are weightedly summed to generate the control strategy.

8. A humanoid robot dual-arm collaborative device, characterized in that: The device comprises: A feature vector acquisition unit, used to acquire a target feature vector; wherein the target feature vector includes a feature vector corresponding to the surrounding environment information of the humanoid robot and a feature vector corresponding to the joint state information of both arms of the humanoid robot; A model building unit, used to build a neural network model and then build a motion branch model and a trajectory branch model based on the neural network model; an output prediction unit, used to input the target feature vector into the motion branch model and the trajectory branch model respectively, to obtain control instructions of the two arms of the humanoid robot output by the motion branch model and predicted trajectory points of the two arms of the humanoid robot output by the trajectory branch model; A control strategy generating unit, used for generating a control strategy according to the control instruction and the predicted trajectory point; The dual-arm driving unit is used to control the dual arms of the humanoid robot to move according to the control strategy.

9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • End-to-end automatic driving method and system based on brain-like neural circuit trajectory guidance

    CN117163067A

  • Double-arm robot control method and device, electronic equipment and storage medium

    CN118514079A