A control method and device for the collaborative operation of the arms of a humanoid robot

By using Transformer encoder and decoder to process action image sequences and multi-source input data, a model for co-motion control of two-arms of humanoid robots is generated, which solves the problems of complexity and uncertainty of co-motion control of two-arms in the prior art, and improves the operation accuracy and success rate.

CN119748429BActive Publication Date: 2025-06-27江淮前沿技术协同创新中心
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411791525.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-06-27
Estimated Expiration
2044-12-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the complexity and uncertainty in the coordinated motion control of humanoid robots with both arms, resulting in poor control effects and long development cycles and high costs.

Method used

By obtaining the sequence of action images generated during the humanoid robot's two arms performing target tasks, the multi-source input data is encoded and decoded using the Transformer encoder and decoder to generate a model for the co-motion control of the two arms.

Benefits of technology

The operation accuracy and success rate of humanoid robots' joint tasks with both arms are improved, and the ability to coordinate and complete tasks independently is enhanced, so that humanoid robots' arms can complete more complex long-action sequence tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119748429B_ABST
    Figure CN119748429B_ABST
Patent Text Reader

Abstract

The present invention discloses a control method and device for the collaborative operation of the two arms of a humanoid robot. The method includes: acquiring an action image sequence generated during the process of the two arms of the humanoid robot executing a target task; determining a first feature vector sequence corresponding to the action image sequence, a mechanical arm joint position vector sequence, and a second feature vector sequence corresponding to a target mask image sequence; performing encoding processing on the first feature vector sequence, the second feature vector sequence, the mechanical arm joint position vector sequence, and a target style variable corresponding to the action image sequence, and outputting a sequence of feature key-value pairs; using the sequence of feature key-value pairs and the humanoid robot body position vector as training samples together; performing supervised imitation learning on a number of training samples to generate a motion control model for the collaborative operation of the two arms. Thus, based on the motion control model, the two arms of the humanoid robot can complete more complex long-action sequence tasks, effectively improving the success rate of target task execution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a control method and device for the cooperative movement of the two arms of a humanoid robot. Background Art

[0002] With the rapid development of artificial intelligence technology, robot technology has made remarkable progress and has been widely applied in various fields. Among them, humanoid robots, as robots that can simulate human behaviors and actions, have extremely high research and application value. However, the motion control of the two arms of humanoid robots has always been a challenging problem. Traditional two-arm motion control methods mainly rely on accurate modeling and complex control algorithms. However, due to the complexity and uncertainty of the two arms of humanoid robots, the control effects of these methods are often not satisfactory. In addition, these methods usually require a large amount of debugging and optimization, resulting in a long development cycle and high costs.

[0003] In recent years, imitation learning, as an emerging machine learning technology, has provided new ideas for the motion control of robot arms. Imitation learning learns from the action data of human experts and uses a sequence of continuous action pictures as the original input of the model, enabling the robot to directly imitate the human motion trajectory and behavior, and thus execute a series of complex continuous actions. This method has the characteristics of simplicity, intuitiveness, and easy implementation, and has achieved good results in some specific scenario tasks. However, pure imitation learning also has some limitations. First, imitation learning requires a large amount of high-quality data to train the model, and such data is often difficult to obtain. Second, imitation learning may be affected by data bias and noise, resulting in limited generalization ability of the learned model. Therefore, using a single sequence of continuous action pictures as the only input for imitation learning makes the content learned by the model more divergent, and the action learning for the target object to be operated is not focused enough, resulting in a reduction in the generalization ability and robustness of the results of imitation learning.

[0004] In existing end-to-end motion control algorithms for robotic arms based on imitation learning, only a single camera image sequence is usually used as the single image data source for training the imitation learning model, and the algorithms are generally applied to traditional robotic arms to perform simple tasks. However, with the development of humanoid robots and the diversification of the application scenarios of robotic arms, using only image training of the imitation learning model for end-to-end motion control of robotic arms can no longer meet the requirements of humanoid robots for performing long-time two-arm cooperative tasks in complex scenarios. Summary of the Invention

[0005] In view of the above problems existing in the prior art, an embodiment of the present invention provides a control method and device for the cooperative movement of the two arms of a humanoid robot. This method can improve the operation accuracy when the two arms of a humanoid robot cooperate to execute a target task, and further improve the success rate of the humanoid robot in executing the target task.

[0006] According to the first aspect of the embodiments of the present invention, a control method for the collaborative operation of the two arms of a humanoid robot is provided, including: obtaining an action image sequence generated during the process of the two arms of the humanoid robot executing a target task; wherein, the action image sequence is used to indicate a set formed by arranging a plurality of original images in chronological order; determining a first feature vector sequence corresponding to the action image sequence, a second feature vector sequence of the mask image sequence corresponding to the action image sequence, and a robotic arm joint position vector sequence corresponding to the action image sequence; wherein, the mask image sequence includes a plurality of target mask images arranged in chronological order, and the target mask image is used to indicate an image obtained by performing mask processing on a target object in the original image; using a Transformer encoder to encode the first feature vector sequence, the second feature vector sequence, the robotic arm joint position vector sequence, and a target style variable corresponding to the action image sequence, and outputting a feature key-value pair sequence; using the feature key-value pair sequence and the humanoid robot body position vector as training samples together; and based on a Transformer decoder, performing supervised imitation learning on a plurality of the training samples to generate a motion control model for the collaborative operation of the two arms.

[0007] Optionally, the determining the first feature vector sequence corresponding to the action image sequence, and the second feature vector sequence of the mask image sequence corresponding to the action image sequence; includes: for any original image in the action image sequence: respectively performing feature extraction processing on the original image and the target mask image corresponding to the original image, and outputting an original image feature and a target mask image feature; performing feature mapping processing on the original image feature and the target mask image feature to generate a fused image and a mapped feature; determining position information of the mapped feature in the fused image; using a linear layer to process the original image feature and the target mask image feature respectively, and outputting an original image feature vector and a target mask image feature vector; adding the position information to the original image feature vector to generate a first feature vector corresponding to the original image; adding the position information to the target mask image feature vector to generate a second feature vector corresponding to the target mask image; generating a first feature vector sequence based on the first feature vector corresponding to each original image in the action image sequence; and generating a second feature vector sequence based on the second feature vector corresponding to the target mask image of each original image in the action image sequence.

[0008] Optionally, determining the sequence of robotic arm joint position vectors corresponding to the action image sequence includes: collecting robotic arm joint position information at different moments during the process of the humanoid robot's two arms performing a target task to obtain a robotic arm joint position sequence; aligning the robotic arm joint position sequence with the action image sequence in terms of time to obtain the sequence of robotic arm joint positions corresponding to the action image sequence; and encoding the robotic arm joint position sequence to output a sequence of robotic arm joint position vectors.

[0009] Optionally, the method further includes: obtaining the original image sequence corresponding to the process of the humanoid robot's two arms performing a target task; for any original image in the original image sequence: obtaining the prior knowledge of the target object and the robotic arm joint position information corresponding to the original image; and forming action image knowledge data from the original image, the prior knowledge of the target object, and the robotic arm joint position information; wherein the robotic arm joint position information corresponding to the original image is used to indicate the robotic arm joint position coordinates collected at the acquisition moment of the original image; generating an action image knowledge sequence based on the action image knowledge data corresponding to each original image in the original image sequence; and performing action chunking on the action image knowledge sequence to generate an action image sequence.

[0010] Optionally, the method further includes: performing target object detection on the original image to generate a target detection box and target object category knowledge; using the target detection box to perform masking processing on the original image to generate a masked image; and using the masked image and the target object category knowledge together as the prior knowledge of the target object.

[0011] Optionally, performing action chunking on the action image knowledge sequence to generate an action image sequence includes: at each moment corresponding to the action image knowledge sequence, performing chunking on the action image knowledge sequence according to a preset step size to output a number of action image chunks; wherein each action image chunk includes a number of action image knowledge data; performing exponential weighting processing on the number of action image chunks to output the chunk weight corresponding to each action image chunk; for any one of the action image chunks: based on the chunk weight, performing weighted sampling on the action image knowledge data in the action image chunk to output a sampling result; and generating an action image sequence based on the sampling results corresponding to each action image chunk.

[0012] Optionally, the method further includes: generating a target object category sequence based on the target object category knowledge corresponding to each of the original images in the action image sequence; respectively encoding the action image sequence and the target object category knowledge sequence, and outputting a target object category knowledge vector sequence and an action vector sequence; using a Transformer encoder to respectively encode the robotic arm joint position vector sequence, the target object category knowledge vector sequence, and the action vector sequence to generate a target style variable.

[0013] Optionally, the method further includes: obtaining a sequence of to-be-tested images generated during the process of a humanoid robot's two arms executing a target task; using the motion control model to perform prediction processing on the sequence of to-be-tested images, and outputting a predicted action of the humanoid robot.

[0014] According to a second aspect of the embodiments of the present invention, there is also provided a control device for the coordinated operation of a humanoid robot's two arms. The device includes: a first acquisition module, configured to acquire an action image sequence generated during the process of a humanoid robot's two arms executing a target task; wherein, the action image sequence is used to indicate a set formed by arranging a plurality of original images in chronological order; a first determination module, configured to determine a first feature vector sequence corresponding to the action image sequence, a second feature vector sequence of a mask image sequence corresponding to the action image sequence, and a robotic arm joint position vector sequence corresponding to the action image sequence; wherein, the mask image sequence includes a plurality of target mask images arranged in chronological order, and the target mask image is used to indicate an image obtained by performing a masking process on a target object in the original image; a first encoding module, configured to use a Transformer encoder to encode the first feature vector sequence, the second feature vector sequence, the robotic arm joint position vector sequence, and a target style variable corresponding to the action image sequence, and output a feature key-value pair sequence; a first generation module, configured to use the feature key-value pair sequence and the humanoid robot body position vector as training samples together; and based on a Transformer decoder, perform supervised imitation learning on a plurality of the training samples to generate a motion control model for the coordinated operation of the two arms.

[0015] According to a third aspect of the embodiments of the present invention, there is also provided an electronic device, which includes: a processor; a memory for storing executable instructions that can be executed by the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the method as described in the first aspect.

[0016] According to a fourth aspect of the embodiments of the present invention, there is also provided a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method as described in the first aspect.

[0017] An embodiment of the present invention provides a control method for the collaborative operation of the two arms of a humanoid robot. The method includes: First, obtain an action image sequence generated during the process of the two arms of the humanoid robot executing a target task; wherein, the action image sequence is used to indicate a set formed by arranging a plurality of original images in chronological order; Second, determine a first feature vector sequence corresponding to the action image sequence, a second feature vector sequence of a mask image sequence corresponding to the action image sequence, and a robotic arm joint position vector sequence corresponding to the action image sequence; wherein, the mask image sequence includes a plurality of target mask images arranged in chronological order, and the target mask image is used to indicate an image obtained by performing mask processing on a target object in the original image; After that, use a Transformer encoder to encode the first feature vector sequence, the second feature vector sequence, the robotic arm joint position vector sequence, and a target style variable corresponding to the action image sequence, and output a feature key-value pair sequence; Finally, use the feature key-value pair sequence and the humanoid robot body position vector as training samples together; based on a Transformer decoder, perform supervised imitation learning on a plurality of the training samples to generate a motion control model for the collaborative operation of the two arms. In this embodiment, the action image, the target mask image, and the robotic arm joint position information corresponding to the same moment during the process of the two arms of the humanoid robot executing a target task are used as multi-source inputs for imitation learning; thus, the recognition and positioning capabilities of the two arms of the humanoid robot for the target object can be strengthened, thereby improving the control accuracy of the collaborative operation of the two arms of the humanoid robot. In this embodiment, the Transformer architecture is applied to the imitation learning of the collaborative operation of the two arms of the humanoid robot, effectively enhancing the collaborative operation ability of the two arms of the humanoid robot and the task autonomous completion ability, so that the two arms of the humanoid robot can complete more complex long action sequence tasks, effectively improving the success rate of the humanoid robot in executing tasks. Description of the Drawings

[0018] Some specific embodiments of the present invention will be described in detail below with reference to the drawings in an exemplary and non-limiting manner. The same reference numerals in the drawings denote the same or similar components or parts. Those skilled in the art should understand that these drawings are not necessarily drawn to scale. In the drawings:

[0019] Figure 1 is a flowchart showing a control method for the collaborative operation of the two arms of a humanoid robot provided by an embodiment of the present invention;

[0020] Figure 2 is a flowchart showing the process of generating an action image sequence in an embodiment of the present invention;

[0021] Figure 3 is a flowchart showing the process of generating a target style variable in an embodiment of the present invention;

[0022] Figure 4 Schematic flow diagram of generating the first feature vector and the second feature vector in an embodiment of the present invention;

[0023] Figure 5 Schematic structural diagram of a control device for the collaborative operation of the two arms of a humanoid robot provided in an embodiment of the present invention. Specific embodiments

[0024] To make the objectives, features, and advantages of the present invention more obvious and understandable, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present invention.

[0025] In the imitation learning method, by directly mapping the RGB images captured by the visible light camera to the actions of the two arms of the humanoid robot, this pixel-to-action mapping is very suitable for the humanoid robot to perform fine operations. However, fine operation tasks usually involve the complex physical properties of the target object to be operated, such as category, shape, size, etc. Therefore, the present invention uses the prior knowledge of the target object as supplementary data for the end-to-end motion control algorithm of the two-arm collaboration.

[0026] In recent years, the Transformer architecture has been widely used in various fields of deep learning and achieved good results. The self-attention mechanism is introduced into the Transformer deep learning model architecture, enabling the entire model to simultaneously consider all position information of the input sequence and assign different attention weights to different parts of the sequence, thereby better capturing the semantic information in the sequence data. Therefore, the Transformer architecture is very suitable for processing sequence-to-sequence tasks and has the ability to synthesize information across sequences and generate new sequences.

[0027] For the complex tasks that may be encountered during the cooperative control of the arms of a humanoid robot, the control method of the present invention first detects and identifies the target object to be operated during the execution of the target task by the arms of the humanoid robot based on the target detection algorithm, and obtains the prior knowledge of the target object (i.e., the target detection box and the target object category knowledge); then uses the target detection box to perform mask cropping on the original image to obtain a masked image; secondly, constructs an imitation learning algorithm based on the Transformer architecture and the action chunking technique, and learns and generates a motion control model for arm cooperation on the action sequence. By performing target detection and masking operations on the original images in each action sequence, and combining the manipulator joint position information in each moment's action, an action image knowledge sequence is formed; finally, the action image knowledge sequence is used as the multi-source training data of the end-to-end motion control algorithm.

[0028] The control method of the present invention is based on the Transformer architecture and the action chunking technique as the main method of imitation learning, and uses the action image knowledge sequence to train the imitation learning algorithm model. Finally, a motion control model for arm cooperation is obtained, providing a new control method for the arms of the humanoid robot to cooperate in executing target tasks, effectively enhancing the arm cooperation ability and the task autonomous completion ability of the humanoid robot, enabling the arms of the humanoid robot to complete more complex long action sequence tasks, and effectively improving the task execution success rate.

[0029] As Figure 1 shown, it is a schematic flowchart of the control method for the arms of a humanoid robot provided by an embodiment of the present invention.

[0030] A control method for the arms of a humanoid robot at least includes the following steps:

[0031] S101, obtain an action image sequence generated during the execution of the target task by the arms of the humanoid robot; wherein, the action image sequence is used to indicate a set formed by arranging a plurality of original images in chronological order;

[0032] S102, determine a first feature vector sequence corresponding to the action image sequence, a second feature vector sequence of the masked image sequence corresponding to the action image sequence, and a manipulator joint position vector sequence corresponding to the action image sequence; wherein, the masked image sequence includes a plurality of target masked images arranged in chronological order, and the target masked image is used to indicate an image obtained by masking the target object in the original image;

[0033] S103, use the Transformer encoder to perform encoding processing on the first feature vector sequence, the second feature vector sequence, the manipulator joint position vector sequence, and the target style variable corresponding to the action image sequence, and output a feature key-value pair sequence;

[0034] S104. Use the feature key-value pair sequence and the humanoid robot body position vector as training samples together; perform supervised imitation learning on a number of training samples based on a Transformer decoder to generate a motion control model for dual-arm collaboration.

[0035] In S101, there are many application scenarios for the humanoid robot to perform the target task. For example, a humanoid robot used in daily life grabs an apple based on dual-arm collaboration, or a humanoid robot used on a production line processes a certain device based on dual-arm collaboration.

[0036] The action image sequence is used to indicate a set formed by arranging in chronological order the original image-related data groups of the key actions of the humanoid robot's dual arms collected by a visible light camera during the process of the humanoid robot's dual arms performing the target task.

[0037] The original image-related data group includes: the original image, and / or the prior knowledge of the target object corresponding to the original image, and / or the manipulator joint position information corresponding to the original image.

[0038] In S102, based on the first feature vector corresponding to each original image in the action image sequence, determine the first feature vector sequence corresponding to the action image sequence; based on the second feature vector of the target mask image corresponding to each original image in the action image sequence, determine the second feature vector sequence corresponding to the action image sequence; obtain the manipulator joint position information corresponding to the acquisition time of each original image in the action image sequence, and determine the manipulator joint position vector corresponding to the original image as the manipulator joint position vector corresponding to the original image; based on the manipulator joint position vector corresponding to each original image in the action image sequence, determine the manipulator joint position vector sequence corresponding to the action image sequence.

[0039] The target object is used to indicate the object to be manipulated by the dual arms during the process of the humanoid robot's dual arms performing the target task.

[0040] In S103, use the first feature vector sequence, the second feature vector sequence, the manipulator joint position vector sequence, and the target style variable corresponding to the action image sequence as the input of the Transformer encoder, and perform processing based on 4 self-attention modules to generate a feature key-value pair sequence.

[0041] The target style variable is used to indicate the unique features of the target object to be manipulated. The use of the target style variable helps the end-to-end motion control model of the humanoid robot's dual arms based on imitation learning to learn more information related to the target object and the manipulator joint actions, thereby improving the success rate of the humanoid robot in performing the target task.

[0042] In S104, the humanoid robot body position vector is generated based on the humanoid robot body position information. The humanoid robot body position information is used to indicate the position information of the humanoid robot body relative to the robot arm base.

[0043] The feature key-value pair sequence and the humanoid robot body position vector are used as training samples, and the training samples are used as the input of the Transformer decoder. They are processed using 7 cross-attention modules, and downsampled and projected to k*N (where N is the sum of the number of degrees of freedom of the humanoid robot's arms) by a multi-layer perceptron, corresponding to the target joint position within the next preset time adjacent to the current moment; the target joint position is compared with the target joint position label to generate an L1 loss function. The above operations are performed based on several training samples; when the L1 loss function tends to be minimized, model parameters are generated; the model is adjusted based on the model parameters to generate a motion control model for dual-arm coordination.

[0044] This embodiment uses the action image, target mask image, and mechanical arm joint position information corresponding to the same moment in the process of the humanoid robot's dual arms performing the target task as multi-source input for imitation learning; thereby, it can enhance the recognition and positioning capabilities of the humanoid robot's dual arms for the target object, thereby improving the control accuracy of the humanoid robot's dual arms in coordination. This embodiment applies the Transformer architecture to the imitation learning of the humanoid robot's dual arms in coordination, effectively enhancing the humanoid robot's dual arms' coordination capabilities and autonomous task completion capabilities, so that the humanoid robot's dual arms can complete more complex long action sequence tasks, effectively improving the success rate of the humanoid robot in performing tasks.

[0045] In addition, this embodiment uses an imitation learning strategy based on the Transformer encoding and decoding structure to achieve end-to-end motion control of the humanoid robot's dual arms. The self-attention structure in the encoder can perform high-dimensional mapping of the feature information of multi-source data, and the cross-attention module and L1 loss function in the decoder can accurately model the predicted action sequence of the humanoid robot's dual arms, which can improve the operation accuracy when the dual arms perform tasks in collaboration, thereby improving the overall task execution success rate.

[0046] In a preferred implementation of this embodiment, the method also includes: a prediction stage; obtaining a test image sequence generated when the humanoid robot's arms perform a target task; using the motion control model to perform prediction processing on the test image sequence, and outputting the predicted action of the humanoid robot.

[0047] Here, the image sequence to be tested is used to indicate a set of images to be tested corresponding to each moment within a preset step time including the current moment and before the current moment, which are arranged in time sequence.

[0048] The predicted action of the humanoid robot is used to indicate the target joint position of the humanoid robot corresponding to the next preset step time adjacent to the current moment.

[0049] Therefore, this embodiment accurately predicts the humanoid robot's movements within the next preset step time based on the motion control model, thereby effectively enhancing the coordination ability and autonomous task completion ability of the humanoid robot's arms in the process of performing the target task, allowing the humanoid robot's arms to complete more complex long motion sequence tasks and effectively improve the success rate of task execution.

[0050] In a preferred implementation manner of this embodiment, the determination of the robotic arm joint position vector sequence corresponding to the action image sequence includes: collecting the robotic arm joint position information at different moments during the process of the humanoid robot's two arms performing the target task to obtain a robotic arm joint position sequence; time-aligning the robotic arm joint position sequence with the action image sequence to obtain a robotic arm joint position sequence corresponding to the action image sequence; encoding the robotic arm joint position sequence to output a robotic arm joint position vector sequence.

[0051] Specifically, the acquisition time of each original image in the action image sequence is obtained to generate an acquisition time sequence; for any acquisition time in the acquisition time sequence: the manipulator joint position information corresponding to the acquisition time is selected from the manipulator joint position sequence; the manipulator joint position information is projected into a manipulator joint position vector based on a linear encoder; based on the manipulator joint position vector corresponding to each acquisition time in the acquisition time sequence, a manipulator joint position vector sequence corresponding to the action image sequence is obtained.

[0052] This embodiment uses the robot arm joint position vector sequence as the input of the humanoid robot dual-arm imitation learning, which can improve the accuracy of the mapping learning between the original image and the robot arm movement, thereby improving the success rate of the humanoid robot dual-arm in performing the target task.

[0053] like Figure 2 FIG. 1 is a flow chart of generating an action image sequence in one embodiment of the present invention.

[0054] Generating an action image sequence includes at least the following steps:

[0055] S201, obtaining a sequence of original images corresponding to the process in which the humanoid robot's arms perform a target task;

[0056] S202. For any original image in the original image sequence: Obtain the prior knowledge of the target object corresponding to the original image and the robotic arm joint position information; Combine the original image, the prior knowledge of the target object, and the robotic arm joint position information into action image knowledge data; where the robotic arm joint position information corresponding to the original image is used to indicate the robotic arm joint position coordinates collected at the acquisition moment of the original image.

[0057] S203. Based on the action image knowledge data corresponding to each original image in the original image sequence, generate an action image knowledge sequence.

[0058] S204. Perform action chunking on the action image knowledge sequence to generate an action image sequence.

[0059] In S201, during the process of the humanoid robot's two arms performing the target task, for end-to-end motion control of the two arms based on the imitation learning method, it is necessary to obtain the robotic arm joint position information and the original image of the camera at each moment during the execution of the target task. Use the joint module and camera in the humanoid robot system, and publish the Topic topics containing the robotic arm joint position information and the original image through the ROS system. Subscribe to the topics through the Python programming language to obtain the original image sequence and the robotic arm joint position sequence.

[0060] In S202 and S203, according to the original image, generate the prior knowledge of the target object corresponding to the original image based on a model or preset rules. Exemplarily, perform target object detection on the original image to generate a target detection box and target object category knowledge; use the target detection box to perform masking processing on the original image to generate a masked image; jointly use the masked image and the target object category knowledge as the prior knowledge of the target object. Here, the target category knowledge is used to indicate the category attributes of the target object to be operated.

[0061] In S204, this embodiment does not make any limitations on the action chunking process. For example: Perform action chunking on the action image knowledge sequence based on a model or preset rules to generate an action image sequence.

[0062] Exemplarily, at each moment corresponding to the action image knowledge sequence, perform chunking on the action image knowledge sequence according to a preset step size to output a number of action image chunks; where the action image chunks include a number of action image knowledge data; perform exponential weighting processing on the number of action image chunks to output the chunk weight corresponding to each action image chunk; for any one of the action image chunks: based on the chunk weight, perform weighted sampling on the action image knowledge data in the action image chunk to output a sampling result; based on the sampling results corresponding to each action image chunk, generate an action image sequence.

[0063] Here, the preset step size is determined according to the actions of the two arms of the humanoid robot during the execution of the target task.

[0064] The algorithm of the present invention performs action chunking on the actions of the two arms of the humanoid robot when executing the target task according to a specific step size k, that is, observing the action every k steps, generating the next k-step action, and executing it in sequence. At the same time, to avoid the instability of the robot's movement caused by the discreteness of the actions after action chunking, a time set method is adopted to perform chunking processing at each moment, and an exponential weighting scheme is used to perform weighted averaging on the prediction of the next moment. In this way, overlapping action image chunks are obtained, and there will be multiple predicted actions within a given time step.

[0065] Assume that the time series corresponding to the action image knowledge sequence is [1min, 2min, 3min, 4min, 5min, 6min, 7min, 8min, 9min, 10min]. An action image sequence is formed based on the original images corresponding to each moment in the time series. If the action image sequence is chunked according to 3 time steps at each moment in this time series, 8 action image chunks are output; among them, each action image chunk includes three original images. The chunk weights of each action image chunk are determined based on the exponential weighting method. Since each original image in the action image chunk has a corresponding image weight, the chunk weight and the image weight are multiplied to obtain the acquisition weight corresponding to each original image. Then, the original images with acquisition weights greater than the preset threshold are selected from each action image chunk as the acquisition results; based on the original images corresponding to the acquisition results of each action image chunk in several action image chunks, an action image sequence is generated.

[0066] This embodiment uses the method of action chunking processing and time set to perform continuous smoothing chunking and exponential weighting prediction on the actions of the two arms of the humanoid robot, enhancing the smoothness of the predicted actions of the two arms of the robot, and avoiding the discrete switching of the actions of the two arms caused by discontinuous chunking, thereby making the movement of the two arms of the robot more stable.

[0067] As Figure 3 shown, it is a schematic flowchart of generating the target style variable in an embodiment of the present invention.

[0068] Generating the target style variable includes at least the following steps:

[0069] S301, generating a target object category sequence based on the target object category knowledge corresponding to each of the original images in the action image sequence;

[0070] S302, respectively performing encoding processing on the action image sequence and the target object category knowledge sequence, and outputting a target object category knowledge vector sequence and an action vector sequence;

[0071] S303. Use the Transformer encoder to encode the robotic arm joint position vector sequence, the target object category knowledge vector sequence, and the action vector sequence respectively to generate the target style variable.

[0072] Specifically, take the robotic arm joint position vector sequence, the target object category knowledge vector sequence, and the action vector sequence as the input of the Transformer encoder, and use 4 self-attention modules and linear layers to obtain the target style variable Z(mean, std), where mean is the mean and std is the variance.

[0073] It should be noted that the target object category knowledge vector sequence, the action vector sequence, and the robotic arm joint position vector sequence have a corresponding relationship in time.

[0074] Figure 4 It is a schematic flow chart for generating the first feature vector and the second feature vector in an embodiment of the present invention.

[0075] S401. Perform feature extraction processing on the original image and the target mask image corresponding to the original image respectively, and output the original image feature and the target mask image feature;

[0076] S402. Perform feature mapping processing on the original image feature and the target mask image feature to generate a fused image and a mapped feature;

[0077] S403. Determine the position information of the mapped feature in the fused image;

[0078] S404. Use the linear layer to process the original image feature and the target mask image feature respectively, and output the original image feature vector and the target mask image feature vector;

[0079] S405. Add the position information to the original image feature vector to generate the first feature vector corresponding to the original image;

[0080] S406. Add the position information to the target mask image feature vector to generate the second feature vector corresponding to the target mask image.

[0081] Specifically, the original image and the target mask image are respectively normalized using a convolutional neural network, and the normalized original image and the normalized target mask image are output; deep feature extraction is respectively performed on the normalized original image and the normalized target mask image to output the original image features and the target mask image features; among them, the original image features contain the global information in the target task scenario, and the target mask image features contain more in-depth features of the target object to be operated. Feature mapping is performed on these two types of features, and the features are projected into embedding vectors through a linear layer, and sinusoidal position information is added to the feature mapping embedding vectors to preserve the feature space position information. Finally, the first feature vector corresponding to the original image and the second feature vector corresponding to the target mask image are obtained, improving the accuracy of the motion control model for predicting the actions of the two arms.

[0082] Next, a control method for the coordinated operation of the two arms of a humanoid robot provided in this embodiment will be described in detail in conjunction with a specific application scenario.

[0083] A control method for the coordinated operation of the two arms of a humanoid robot at least includes the following steps:

[0084] S1. Obtain the original image sequence corresponding to the process of the two arms of the humanoid robot executing the target task.

[0085] S2. For any original image in the original image sequence: perform target object detection on the original image to generate a target detection box and target object category knowledge; use the target detection box to perform masking processing on the original image to generate a mask image; use the mask image and the target object category knowledge together as the prior knowledge of the target object corresponding to the original image; obtain the mechanical arm joint position information corresponding to the original image; form the action image knowledge data from the original image, the prior knowledge of the target object, and the mechanical arm joint position information; where the mechanical arm joint position information corresponding to the original image is used to indicate the mechanical arm joint position coordinates collected at the acquisition moment of the original image.

[0086] S3. Generate an action image knowledge sequence based on the action image knowledge data corresponding to each original image in the original image sequence.

[0087] S4. At each moment corresponding to the action image knowledge sequence, perform block processing on the action image knowledge sequence according to a preset step size to output a number of action image blocks; where the action image blocks include a number of action image knowledge data; perform exponential weighting processing on the number of action image blocks to output the block weight corresponding to each action image block; for any action image block: based on the block weight, perform weighted sampling on the action image knowledge data in the action image block to output a sampling result; generate an action image sequence based on the sampling results corresponding to each action image block.

[0088] S5. Collect the mechanical arm joint position information at different moments during the process of the humanoid robot's two arms performing the target task to obtain the mechanical arm joint position sequence; perform time alignment on the mechanical arm joint position sequence and the action image sequence to obtain the mechanical arm joint position sequence corresponding to the action image sequence; perform encoding processing on the mechanical arm joint position sequence and output the mechanical arm joint position vector sequence.

[0089] S6. Generate the target object category knowledge sequence based on the target object category knowledge corresponding to each original image in the action image sequence; perform encoding processing on the action image sequence and the target object category knowledge sequence respectively, and output the target object category knowledge vector sequence and the action vector sequence; use the Transformer encoder to perform encoding processing on the mechanical arm joint position vector sequence, the target object category knowledge vector sequence, and the action vector sequence respectively to generate the target style variable.

[0090] S7. For any original image in the action image sequence: perform feature extraction processing on the original image and the target mask image corresponding to the original image respectively, and output the original image feature and the target mask image feature; perform feature mapping processing on the original image feature and the target mask image feature to generate the fused image and the mapped feature; determine the position information of the mapped feature in the fused image; use the linear layer to process the original image feature and the target mask image feature respectively, and output the original image feature vector and the target mask image feature vector; add the position information to the original image feature vector to generate the first feature vector corresponding to the original image; add the position information to the target mask image feature vector to generate the second feature vector corresponding to the target mask image. Generate the first feature vector sequence based on the first feature vector corresponding to each original image in the action image sequence; generate the second feature vector sequence based on the second feature vector corresponding to the target mask image of each original image in the action image sequence.

[0091] S8. Use the Transformer encoder to perform encoding processing on the first feature vector sequence, the second feature vector sequence, the mechanical arm joint position vector sequence, and the target style variable corresponding to the action image sequence, and output the feature key-value pair sequence.

[0092] S9. Use the feature key-value pair sequence and the humanoid robot body position vector as training samples together; perform supervised imitation learning on several training samples based on the Transformer decoder to generate the motion control model for two-arm cooperation.

[0093] S10. Obtain the to-be-tested image sequence generated during the process of the humanoid robot's two arms performing the target task; use the motion control model for two-arm cooperation to perform prediction processing on the to-be-tested image sequence and output the predicted action of the humanoid robot.

[0094] In this embodiment, for the end-to-end motion control algorithm of the humanoid robot's dual arms, the original image, the prior knowledge of the target object, and the robotic arm joint position information are used as the main training data for the imitation learning algorithm based on the Transformer architecture. On the basis of the original image data features, the recognition and positioning capabilities of the robotic arm's imitation learning are enhanced by using the masked image, the target object category knowledge, and the robotic arm joint position information, thereby improving the robotic arm motion control accuracy. The target object category knowledge is beneficial for the robotic arm motion control algorithm to autonomously predict operations such as target grasping in unseen target states. The target masked image can enhance the robotic arm's perception ability of the object to be operated, so as to achieve more accurate recognition and positioning. Therefore, the control method for the dual-arm coordination of the humanoid robot in this embodiment can effectively enhance the dual-arm coordination ability and the task autonomous completion ability of the humanoid robot, so that the dual arms of the humanoid robot can complete more complex long action sequence tasks, effectively improving the success rate of target task execution.

[0095] As Figure 5 shown, it is a schematic structural diagram of a control device for the dual-arm coordination of a humanoid robot provided by an embodiment of the present invention.

[0096] A control device for the dual-arm coordination of a humanoid robot, the device 500 includes: a first acquisition module 501, configured to acquire an action image sequence generated during the execution of a target task by the dual arms of the humanoid robot; wherein, the action image sequence is used to indicate a set formed by arranging a plurality of original images in chronological order; a first determination module 502, configured to determine a first feature vector sequence corresponding to the action image sequence, a second feature vector sequence of a masked image sequence corresponding to the action image sequence, and a robotic arm joint position vector sequence corresponding to the action image sequence; wherein, the masked image sequence includes a plurality of target masked images arranged in chronological order, and the target masked image is used to indicate an image obtained by masking the target object in the original image; a first encoding module 503, configured to use a Transformer encoder to encode the first feature vector sequence, the second feature vector sequence, the robotic arm joint position vector sequence, and a target style variable corresponding to the action image sequence, and output a feature key-value pair sequence; a first generation module 504, configured to use the feature key-value pair sequence and the humanoid robot body position vector as training samples together; and perform supervised imitation learning on a plurality of the training samples based on a Transformer decoder to generate a motion control model for dual-arm coordination.

[0097] In a preferred implementation manner of this embodiment, the first determination module includes: a first determination unit, configured to, for any original image in the action image sequence: perform normalization processing on the original image and the target mask image corresponding to the original image respectively, and output the normalized original image and the normalized target mask image; perform depth feature extraction on the normalized original image and the normalized target mask image respectively, to generate an original image feature and a target mask image feature; perform feature mapping processing on the original image feature and the target mask image feature, to generate a fused image and a mapped feature; determine the position information of the mapped feature in the fused image; use a linear layer to process the original image feature and the target mask image feature respectively, and output an original image feature vector and a target mask image feature vector; add the position information to the original image feature vector to generate a first feature vector corresponding to the original image; add the position information to the target mask image feature vector to generate a second feature vector corresponding to the target mask image; a second determination unit, configured to generate a first feature vector sequence based on the first feature vectors corresponding to each original image in the action image sequence; a third determination unit, configured to generate a second feature vector sequence based on the second feature vectors corresponding to the target mask images of each original image in the action image sequence.

[0098] In a preferred implementation manner of this embodiment, the first determination module further includes: a collection unit, configured to collect the mechanical arm joint position information at different moments during the process of the humanoid robot's two arms executing a target task, to obtain a mechanical arm joint position sequence; a time alignment unit, configured to perform time alignment on the mechanical arm joint position sequence and the action image sequence, to obtain a mechanical arm joint position sequence corresponding to the action image sequence; an encoding processing unit, configured to perform encoding processing on the mechanical arm joint position sequence, and output a mechanical arm joint position vector sequence.

[0099] In a preferred implementation manner of this embodiment, the device further includes: a second acquisition module, configured to acquire an original image sequence corresponding to the process of the humanoid robot's two arms executing a target task; a second determination module, configured to, for any original image in the original image sequence: acquire the prior knowledge of the target object and the mechanical arm joint position information corresponding to the original image; form action image knowledge data from the original image, the prior knowledge of the target object, and the mechanical arm joint position information; wherein, the mechanical arm joint position information corresponding to the original image is used to indicate the mechanical arm joint position coordinates collected at the acquisition moment of the original image; a second generation module, further configured to generate an action image knowledge sequence based on the action image knowledge data corresponding to each original image in the original image sequence; a block processing module, configured to perform action block processing on the action image knowledge sequence, to generate an action image sequence.

[0100] In a preferred embodiment of this embodiment, the device further includes: a target detection module, configured to detect a target object in the original image, and generate a target detection box and target object category knowledge; a masking processing module, configured to perform masking processing on the original image by using the target detection box to generate a masked image; a third determination module, configured to use the masked image and the target object category knowledge together as target object prior knowledge.

[0101] In a preferred embodiment of this embodiment, the block processing module includes: a block processing unit, configured to perform block processing on the action image knowledge sequence at each moment corresponding to the action image knowledge sequence according to a preset step length, and output a plurality of action image blocks; wherein, each action image block includes a plurality of action image knowledge data; an exponential weighting processing unit, configured to perform exponential weighting processing on the plurality of action image blocks, and output a block weight corresponding to each action image block; a weighted sampling unit, configured to, for any one of the action image blocks: based on the block weight, perform weighted sampling on the action image knowledge data in the action image block, and output a sampling result; a generation unit, configured to generate an action image sequence based on the sampling result corresponding to each action image block.

[0102] In a preferred embodiment of this embodiment, the device further includes: a third generation module, configured to generate a target object category sequence based on the target object category knowledge corresponding to each original image in the action image sequence; a second encoding processing module, configured to perform encoding processing on the action image sequence and the target object category knowledge sequence respectively, and output a target object category knowledge vector sequence and an action vector sequence; a third encoding processing module, configured to use a transformer encoder to perform encoding processing on the robotic arm joint position vector sequence, the target object category knowledge vector sequence, and the action vector sequence respectively, and generate a target style variable.

[0103] In a preferred embodiment of this embodiment, the device further includes: a third acquisition module, configured to acquire a sequence of to-be-tested images generated during the process of a humanoid robot's two arms executing a target task; a prediction processing module, configured to perform prediction processing on the sequence of to-be-tested images by using the motion control model, and output a predicted action of the humanoid robot.

[0104] The above device can execute a control method for the cooperation of the two arms of a humanoid robot provided in an embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing a control method for the cooperation of the two arms of a humanoid robot. For technical details not described in detail in this embodiment, reference may be made to a control method for the cooperation of the two arms of a humanoid robot provided in an embodiment of the present invention.

[0105] The present invention also provides an electronic device, comprising: a processor; a memory for storing executable instructions executable by the processor; and the processor for reading the executable instructions from the memory and executing the instructions to implement a control method for the coordinated operation of the two arms of a humanoid robot according to the present invention.

[0106] In addition to the above methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions that, when run by a processor, cause the processor to execute the steps in the methods according to various embodiments of the present application described in the "Exemplary Method" section above of this specification.

[0107] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0108] Furthermore, an embodiment of the present application may also be a computer-readable storage medium, on which computer program instructions are stored, and the computer program instructions, when run by a processor, cause the processor to execute the steps in the methods according to the following embodiments of the present application described in the "Exemplary Method" section above of this specification.

[0109] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0110] The basic principles of the present application have been described in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present application are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present application. Additionally, the specific details disclosed above are only for illustrative and facilitating understanding purposes, rather than limitations. These details do not limit the present application to necessarily implementing with the above specific details.

[0111] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present application are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The word "or" and "and" used herein refer to the word "and / or", and can be used interchangeably with it, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with it.

[0112] It should also be noted that in the devices, equipment, and methods of the present application, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present application.

[0113] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present application. Various modifications to these aspects will be very apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the present application. Therefore, the present application is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

[0114] The above description has been given for purposes of illustration and description. In addition, this description does not intend to limit the embodiments of the present application to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof.

[0115] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0116] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality of" means two or more, unless otherwise specifically defined.

[0117] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claimed rights.

Claims

1. A control method for the coordination of two arms of a humanoid robot, characterized in that: include: Acquire a sequence of action images generated when the humanoid robot's arms perform a target task; wherein the sequence of action images is used to indicate a set of several original images arranged in time sequence; Determine a first feature vector sequence corresponding to the action image sequence, a second feature vector sequence of a mask image sequence corresponding to the action image sequence, and a robot arm joint position vector sequence corresponding to the action image sequence; wherein the mask image sequence includes a plurality of target mask images arranged in chronological order, and the target mask image is used to indicate an image after masking the target object in the original image; Using a Transformer encoder to encode the first feature vector sequence, the second feature vector sequence, the robot arm joint position vector sequence, and the target style variables corresponding to the action image sequence, and output a feature key-value pair sequence; The feature key-value pair sequence and the humanoid robot body position vector are used together as training samples; supervised imitation learning is performed on several of the training samples based on the Transformer decoder to generate a motion control model for dual-arm collaboration.

2. The method according to claim 1, characterized in that The step of determining a first feature vector sequence corresponding to the action image sequence and a second feature vector sequence of a mask image sequence corresponding to the action image sequence comprises: For any original image in the action image sequence: perform feature extraction processing on the original image and the target mask image corresponding to the original image respectively, and output the original image features and the target mask image features; perform feature mapping processing on the original image features and the target mask image features to generate a fused image and mapping features; determine the position information of the mapping features in the fused image; use linear layers to process the original image features and the target mask image features respectively, and output the original image feature vector and the target mask image feature vector; add the position information to the original image feature vector to generate a first feature vector corresponding to the original image; add the position information to the target mask image feature vector to generate a second feature vector corresponding to the target mask image; Generate a first feature vector sequence based on a first feature vector corresponding to each of the original images in the action image sequence; A second feature vector sequence is generated based on the second feature vector corresponding to the target mask image of each original image in the action image sequence.

3. The method according to claim 1, characterized in that The step of determining the robot arm joint position vector sequence corresponding to the action image sequence comprises: Collecting the mechanical arm joint position information at different moments during the process of the two arms of the humanoid robot performing the target task, and obtaining the mechanical arm joint position sequence; Temporally aligning the robot arm joint position sequence with the action image sequence to obtain a robot arm joint position sequence corresponding to the action image sequence; The robot arm joint position sequence is encoded and a robot arm joint position vector sequence is output.

4. The method according to claim 1, characterized in that Also includes: Acquire a sequence of original images corresponding to the process in which the humanoid robot's two arms perform a target task; For any original image in the original image sequence: obtaining target object prior knowledge and robot arm joint position information corresponding to the original image; combining the original image, target object prior knowledge, and robot arm joint position information into action image knowledge data; wherein the robot arm joint position information corresponding to the original image is used to indicate the robot arm joint position coordinates collected at the time of collecting the original image; generating an action image knowledge sequence based on the action image knowledge data corresponding to each of the original images in the original image sequence; The action image knowledge sequence is subjected to action block processing to generate an action image sequence.

5. The method according to claim 4, characterized in that Also includes: Performing target object detection on the original image to generate target detection frames and target object category knowledge; Using the target detection frame to perform mask processing on the original image to generate a mask image; The mask image and target object category knowledge are used together as target object prior knowledge.

6. The method according to claim 4, characterized in that The step of performing action block processing on the action image knowledge sequence to generate an action image sequence comprises: At each moment corresponding to the action image knowledge sequence, the action image knowledge sequence is processed into blocks according to a preset step length, and a plurality of action image blocks are output; wherein the action image blocks include a plurality of action image knowledge data; Performing exponential weighting processing on the plurality of action image blocks, and outputting a block weight corresponding to each action image block; For any of the action image blocks: based on the block weights, weighted sampling is performed on the action image knowledge data in the action image block, and a sampling result is output; Based on the sampling results corresponding to each of the action image blocks, an action image sequence is generated.

7. The method according to claim 1, characterized in that Also includes: generating a target object category sequence based on target object category knowledge corresponding to each of the original images in the action image sequence; Encoding the action image sequence and the target object category knowledge sequence respectively, and outputting a target object category knowledge vector sequence and an action vector sequence; The robot arm joint position vector sequence, the target object category knowledge vector sequence and the action vector sequence are respectively encoded using a transformer encoder to generate a target style variable.

8. The method according to claim 1, characterized in that Also includes: Acquire a sequence of images to be tested generated when the humanoid robot's arms perform a target task; The motion control model is used to perform prediction processing on the image sequence to be tested, and the predicted action of the humanoid robot is output.

9. A control device for the coordination of two arms of a humanoid robot, characterized in that: include: A first acquisition module is used to acquire a sequence of action images generated when the humanoid robot's arms perform a target task; wherein the sequence of action images is used to indicate a set of several original images arranged in time sequence; A first determination module is used to determine a first feature vector sequence corresponding to the action image sequence, a second feature vector sequence of a mask image sequence corresponding to the action image sequence, and a robot arm joint position vector sequence corresponding to the action image sequence; wherein the mask image sequence includes a plurality of target mask images arranged in chronological order, and the target mask image is used to indicate an image after mask processing is performed on a target object in the original image; A first encoding module is used to encode the first feature vector sequence, the second feature vector sequence, the robot arm joint position vector sequence, and the target style variable corresponding to the action image sequence using a Transformer encoder, and output a feature key-value pair sequence; The first generation module is used to use the feature key-value pair sequence and the humanoid robot body position vector as training samples; based on the Transformer decoder, supervised imitation learning is performed on several of the training samples to generate a motion control model for dual-arm collaboration.

10. A computer readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Simulation learning mechanical arm grabbing method and device based on multi-scale sequence model

    CN116901071A

  • Simulation learning fruit picking method and device based on multi-modal information fusion

    CN116985132A