Mechanical arm control method and system based on imitation learning

Through the ACT algorithm of VR remote operation and multi-modal image fusion, the problem of robotic arm movement deviation in complex lighting environments is solved, and low-cost and efficient robotic arm control is achieved.

CN120244985APending Publication Date: 2025-07-04HARBIN ENG UNIV

Patent Information

Application Number
CN202510605975.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, teaching data collection takes time and is costly. The performance of mainstream imitation learning algorithms is limited in complex lighting environments, and the movement of the robotic arm is prone to deviation.

Method used

The VR remote operation data collection method combined with the ACT algorithm of multi-modal image fusion is used to optimize the feature processing link through infrared and visible light features, and improve the operation accuracy and robustness of the robotic arm in complex lighting environments.

Benefits of technology

It realizes low-cost and efficient teaching data collection, and the robotic arm operates accurately and smoothly in complex lighting environments, and can complete tasks independently.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120244985A_ABST
    Figure CN120244985A_ABST
Patent Text Reader

Abstract

The invention discloses a mechanical arm control method and system based on imitation learning, and belongs to the field of intelligent robots and automatic control. The data collection layer designs a data collection method based on a VR platform, and is responsible for collecting teaching data to make a complex scene data set, performing coordinate conversion to solve angle information of each joint of the mechanical arm, realizing VR teleoperation teaching and teaching data collection, and transmitting the teaching data to the ACT algorithm layer for multi-modal image fusion; the multi-modal image fusion ACT algorithm layer carries out feature extraction and fusion denoising on complex illumination scene data, then infrared and visible light information fusion complementary features are combined with an ACT algorithm to obtain predicted action parameters, the predicted action parameters are converted into corresponding action plans through a motion control strategy, and a mechanical arm is driven to autonomously complete tasks. The method solves the problems that the performance of a mainstream imitation learning algorithm is limited in a complex illumination environment, the motion of the mechanical arm is prone to deflection and the like, and the perceptual performance and the capacity of adapting to complex scenes of the mechanical arm under the complex illumination condition are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent robots and automatic control, and particularly relates to a manipulator control method and system based on imitation learning. Background Art

[0002] Imitation learning technology has been widely applied in the field of robots. Thanks to the recent explosion of this technology, robots have shown greater potential in complex environments. The collection of demonstration data is the core problem faced by imitation learning first, which directly affects the performance of imitation learning. However, at present, the collection of demonstration device data takes a long time, and the cost and operation threshold are high, so it cannot be popularized on a large scale. At the same time, traditional imitation learning algorithms have poor stability and low efficiency, and the accuracy and robustness of the optimal strategy of the advanced ACT algorithm will be greatly reduced in complex lighting environments, and the robot actions are prone to deviation.

[0003] Facing the above problems, we designed a VR teleoperation data collection method, developed in combination with VR devices and the unity3D platform, and directly taught through VR devices; and proposed an ACT algorithm for multi-modal image fusion, introducing a gradient fusion attention mechanism and an image weight generation network to perform feature fusion on infrared and visible light modalities, optimizing the ACT image feature processing link, and constructing a manipulator control method. The VR teleoperation data collection method has high teaching efficiency and simple operation. The algorithm can provide high-quality action strategies in complex environments. The manipulator's autonomous task actions are smooth, coherent, efficient and safe.

[0004] In the process of VR teleoperation data collection, coordinate transformation is one of the core steps. It obtains the pose information of the VR eyes to determine the pose of the binocular camera mapped to the VR coordinate system, constructs a pose transformation matrix, and realizes the coordinate transformation from the binocular camera coordinate system to the VR global coordinate system. In addition, in the ACT algorithm for multi-modal image fusion, the feature extraction backbone module, the gradient fusion attention module, and the bimodal feature fusion module jointly complete the feature extraction, fusion and denoising of complex image data. The ACT algorithm is an algorithm for imitation learning, which can generate a predicted action sequence by inputting data such as joint positions and image features, and use a motion control strategy to drive the manipulator to complete tasks autonomously.

[0005] The present invention provides a manipulator control method based on imitation learning. By integrating a virtual reality (VR) teleoperation data collection method and an ACT algorithm for multi-modal image fusion, it uses VR devices to collect demonstration data efficiently, and combines the feature fusion technology of infrared and visible light images to perform data association and processing, so as to improve the action accuracy and robustness of the manipulator in complex lighting environments, and designs three autonomous tasks to verify the performance of this method. Summary of the Invention

[0006] The object of the present invention is to provide a manipulator control method and system based on imitation learning, aiming at the problems of low efficiency and high cost in teaching data collection in the prior art, the limited performance of mainstream imitation learning algorithms in complex lighting environments, and the easy deflection of manipulator movements.

[0007] The object of the present invention is achieved by the following technical solutions:

[0008] A manipulator control method based on imitation learning, the specific steps are as follows:

[0009] Step 1: Configure the robot and run it, and collect the operator's teaching data through the VR interaction module, including the pose information of using the handle to control the manipulator to complete tasks;

[0010] Step 2: Establish a D-H coordinate model of the manipulator using the actual model of the robot, and combine the D-H parameters of each link and the expression of the transformation relationship between adjacent link coordinate systems to obtain the forward and inverse kinematic relationships of the manipulator;

[0011] Step 3: Map the VR handle pose information to the manipulator base coordinate system, solve the manipulator joint angles in combination with the cyclic coordinate descent method, and drive the manipulator to move through the manipulator Python SDK;

[0012] Step 4: Use a binocular camera to collect environmental image data, calculate the transformation matrix from the camera coordinate system to the VR global coordinate system through hand-eye calibration and the Rodrigues formula, and synchronously construct a complex scene dataset containing multi-modal images and teaching actions;

[0013] Step 5: Perform shallow feature extraction on the infrared image and visible light images of the two modalities respectively to generate an infrared feature map and a visible light feature map;

[0014] Step 6: Extract the gradient texture features of the infrared and visible light images through the gradient fusion attention module, and enhance the complementary information of the infrared and visible light modalities in combination with the coordinate attention mechanism;

[0015] Step 7: Perform cross-modal fusion on the infrared and visible light features through the dual-modal feature fusion module to generate a fused feature map;

[0016] Step 8: Input the fused feature map into the ACT algorithm module, and use the conditional variational autoencoder to generate the manipulator prediction action sequence;

[0017] Step 9: Transmit the predicted action sequence to the motion control strategy part to enable the manipulator to autonomously complete tasks.

[0018] Further, the step 2 is implemented by a robotic arm kinematic analysis module. A D-H coordinate model of the robotic arm is established according to the actual model of the robot and specific DH parameters. Combining the D-H parameters of each connecting rod and the expression of the transformation relationship between adjacent connecting rod coordinate systems The kinematic relationship between adjacent connecting rods is obtained. The formula is as follows:

[0019]

[0020] where i is the serial number of the robotic arm joint, α i-1 is the connecting rod twist angle, d is the connecting rod offset, a i-1 is the connecting rod length, θ is the joint angle, is the transformation relationship matrix between adjacent joint coordinate systems, is the transformation relationship matrix of the robotic arm end relative to the base.

[0021] Further, the step 4 obtains the transformation matrix from the camera coordinate system to the robot base coordinate system through hand-eye calibration The transformation matrix from the camera coordinate system to the VR global coordinate system is calculated through the Rodrigues formula The formula is as follows:

[0022]

[0023] where is the transformation matrix from the camera coordinate system to the VR glasses coordinate system, is the matrix for transforming the VR glasses coordinate system to the VR global coordinate system;

[0024] where

[0025] where R i (θ) is the augmented matrix composed of the rotation matrix obtained by rotating θ angles around the i-axis and the position coordinates;

[0026] Calculated through the Rodrigues formula

[0027]

[0028] where g c 、g g are the gravity vectors of the camera and the VR glasses, n i is n x 、n y 、n z , which is the common rotation axis when the two coordinate systems of g c 、g g coincide, and θ is the included angle between the two vectors;

[0029]

[0030] Among them, the identity matrix A provides the initial state of rotation, and B actually encodes the direction information of the rotation axis. and Given this, the transformation matrix from the camera coordinate system to the VR global coordinate system can be solved.

[0031] Furthermore, the specific steps of step 6 are as follows:

[0032] (1) Use a gradient fusion interaction-based residual-like structure. The mainstream part uses ordinary CBL convolution combinations, including a 3×3 convolutional layer, a regularization BN layer, and a LeakyRelu activation function. The residual part uses two-way gradient operator convolutions and a pointwise convolution. In the gradient operator convolution, a gradient operator is introduced to convolve the input features with a high-frequency convolution kernel to obtain the gradient information of the image and capture the texture details of the object. This process is expressed as:

[0033]

[0034] Among them, F C ′ is the output visible light feature map, ▽ refers to the gradient operator, PWConv(·) represents the pointwise convolution operation, and C(·) represents the Concat splicing operation. represents element-wise summation. This part aggregates the learnable convolution features with gradient magnitude information.

[0035] Here, the Sobel operator is used to calculate the gradient magnitude. The Sobel gradient operator used is:

[0036]

[0037] Among them, G X and G Y are the gradient information obtained from the feature map F in the horizontal and vertical directions, respectively.

[0038] (2) In the processing of the infrared branch, use the coordinate attention mechanism to introduce spatial coordinate information to enhance the feature representation, capture the important spatial positions and inter-channel dependencies in the feature map, and further extract the texture features of the infrared image;

[0039] (3) In the processing of the visible light branch, use the C3CrossCovn module composed of cross convolutions for feature extraction and multi-level fusion to further extract the texture features of the visible light image; this module consists of two standard convolutional layers arranged in a cross pattern. The first cross convolution uses a 1×k-sized kernel with a horizontal stride of 1 and a vertical stride of s, and the second uses a k×1 kernel with a stride of s in both dimensions.

[0040] Further, in step 7, a fused feature map is generated as follows:

[0041] F O = S2(O), F C = S2(C)

[0042]

[0043] where O is the infrared image, C is the visible light image, is the gradient texture feature of the infrared image, is the gradient texture feature of the visible light image, S i (.) represents the function of the i-layer Conv, G(.) represents the function of the GFAM module, and D(.) represents the function of the DFFM module; the fused feature map is then passed through the subsequent network of the original yolov5-s model for object classification and localization.

[0044] Further, the dual-modal feature fusion is specifically as follows:

[0045] (1) Perform a preliminary element-wise addition fusion on the input features X and Y of the two modalities, and then the WGM module processes them through two different dimensions respectively;

[0046] (2) In the importance analysis of the left channel dimension, global average pooling is used to compress the feature map of size C×H×W into C×1×1, and two lightweight PW convolutions are used to learn the attention information of the channel dimension; in the attention analysis of the right spatial dimension, the PW convolution is used to reduce the dimension of the input feature map, and further two dilated convolutions with a 3×3 convolution kernel size are used to extract feature information, and the feature map is mapped to 1×H×W through the PW convolution;

[0047] (3) Use the Sigmoid activation function to make the output value between 0 and 1 to obtain the fusion weight value W of feature X, where the weight of feature Y is 1 - W; adaptively adjust the weights W and 1 - W to provide the learning ability of the model during backpropagation;

[0048] (4) Use the broadcast mechanism to enable the attention information on both sides to be added, multiply the weights with features X and Y respectively and then add them element-wise to generate the fused feature Z; among them, the fused feature Z is a feature sequence obtained by flattening the feature map obtained by the above feature extraction process of the RGB image and optimizing it.

[0049] Further, step 8 is specifically as follows:

[0050] (1) Process the fused feature Z through the encoder of the autoencoder to obtain the style variable α. Among them, the style variable α is a low-dimensional representation obtained by the encoder processing the action sequence and joint positions of the fused feature Z.

[0051] (2) Fuse the style variable α with the feature Z to obtain a comprehensive feature sequence. Among them, the feature sequence is a series of feature vectors containing information such as joint positions.

[0052] (3) Use the decoder of the autoencoder to reconstruct the comprehensive feature sequence using the L1 loss, which is the sum of the absolute values of the differences between the corresponding elements in the two sequences, and accurately model the action sequence to generate a predicted action sequence. Among them, the action sequence is a series of action instructions arranged in chronological order and consists of a series of feature vectors containing information such as joint positions. The L1 loss function used is:

[0053]

[0054] Among them, Y i is the input actual value, is the predicted value of the model, and n is the number of samples.

[0055] A robotic arm control system based on imitation learning, including a data collection layer and an ACT algorithm layer for multimodal image fusion;

[0056] The main function of the data collection layer is to collect the teaching data of the operator through the VR interaction module, including the pose information of the robotic arm, and convert this data into executable instructions for the robotic arm through coordinate transformation and kinematic analysis. At the same time, this layer is also responsible for integrating the teaching data and images of different scenarios to construct a complex scenario dataset;

[0057] The main function of the ACT algorithm layer for multimodal image fusion is to process the multimodal images of the data collection layer, extract and fuse features through the gradient fusion attention module and the bimodal feature fusion module, and finally use the ACT algorithm to generate the predicted action sequence of the robotic arm and drive the robotic arm to execute tasks through the motion control strategy.

[0058] Furthermore, the data collection layer includes a robotic arm motion analysis module, a coordinate transformation module, a VR interaction module, and a complex scenario dataset creation module; the robotic arm motion analysis module uses the actual model of the robot to establish a D-H coordinate model of the robotic arm, and combines the D-H parameters of each link and the expression of the coordinate system transformation relationship between adjacent links to obtain the forward and inverse kinematic relationships of the robotic arm; the coordinate transformation module uses the coordinate system transformation matrix to realize the transformation from the camera coordinate system to the VR global coordinate system, so as to achieve an accurate mapping of the operator's actions in the VR environment to the robotic arm motion; the VR interaction module uses the VR headset and VR handle to adjust the state of the robotic arm according to the task to simplify the teaching process; the complex scenario dataset creation module is responsible for integrating and processing the teaching data and images of different scenarios during the teaching process and transmitting them to the ACT algorithm layer for multi-modal image fusion.

[0059] Furthermore, the ACT algorithm layer for multi-modal image fusion includes a gradient fusion attention module, a bimodal feature fusion module, an ACT algorithm module, and a motion control strategy module; the gradient fusion attention module captures the image texture information through a gradient operator, and uses a residual-like structure to realize the interaction of the gradient information of the two branches. At the same time, a coordinate attention mechanism is introduced into the infrared branch to enable the model to better understand and utilize the complementary information of the infrared and visible light modalities; the bimodal feature fusion module is used to fuse the infrared and visible light features extracted above and transmit the fused feature map to the ACT algorithm module; the ACT algorithm module is used to generate a predicted robotic arm motion plan according to the generated fused feature map and the training of the teaching data; the motion control strategy module is used to convert the predicted robotic arm motion plan into a specific robotic arm motion plan, so as to drive the robotic arm to autonomously complete the task.

[0060] The beneficial effects of the present invention are as follows:

[0061] Through the innovative VR teleoperation data collection method and the ACT algorithm for multi-modal image fusion, the present invention solves the deficiencies of the existing robotic arm control methods in terms of data collection and algorithm performance. The present invention uses multi-modal image fusion to combine the information of the infrared and visible light modalities, giving full play to the feature advantages extracted by different modalities. The low-cost and easy-to-operate teaching data collection and the ACT algorithm for multi-modal image fusion significantly improve the robustness and accuracy in complex lighting environments, and the robotic arm can efficiently and smoothly complete tasks autonomously. Description of the Drawings

[0062] Figure 1 It is a schematic diagram of the structural composition of the present invention;

[0063] Figure 2 It is a schematic diagram of the robotic arm control experimental platform;

[0064] Figure 3 It is the overall framework diagram of the VR teleoperation data collection method;

[0065] Figure 4 It is the D-H coordinate model diagram of the robotic arm;

[0066] Figure 5 It is the schematic diagram of the conversion between VR coordinates and robot coordinates;

[0067] Figure 6 It is the main framework diagram of the improved feature extraction;

[0068] Figure 7 It is the framework diagram of the gradient fusion attention module;

[0069] Figure 8 It is the framework diagram of the bimodal feature fusion module;

[0070] Figure 9 It is the overall framework diagram of the ACT algorithm;

[0071] Figure 10 It is the schematic diagram of the robotic arm motion control strategy;

[0072] Figure 11 It is the test diagram of the robotic arm's autonomous fruit picking task under different lighting conditions;

[0073] Figure 12 It is the test diagram of the robotic arm's autonomous door opening task under different lighting conditions;

[0074] Figure 13 It is the test diagram of the robotic arm's autonomous desktop storage task under different lighting conditions. Specific implementation manners

[0075] The following further describes the present invention with reference to the accompanying drawings. It should be noted that the described embodiments are only partial embodiments of the present invention, not all of them. The description of the embodiments of the present invention is only for illustration and should not be regarded as a limitation on the present invention and its application scope. Other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention:

[0076] Example 1:

[0077] As shown in the attached Figure 1 figure, this example provides a robotic arm control method and system based on imitation learning, and the system includes:

[0078] Data collection layer: Through the VR interaction module, the operator uses the handle to control the robotic arm to complete tasks, obtain teaching data, and use the complex scene dataset production module to provide corresponding data to the ACT algorithm layer for multi-modal image fusion. The coordinate transformation module maps the handle pose information to the robotic arm coordinate system and uses the robotic arm kinematic analysis module to solve the joint angle information of the robotic arm to ensure the precise movement of the robotic arm.

[0079] Specifically, it includes a robotic arm motion analysis module, a coordinate transformation module, a VR interaction module, and a complex scene dataset production module; the robotic arm motion analysis module uses the actual model of the robot to establish the D-H coordinate model of the robotic arm, combines the D-H parameters of each link and the expression of the adjacent link coordinate system transformation relationship, and obtains the forward and inverse kinematic relationships of the robotic arm; the coordinate transformation module uses the coordinate system transformation matrix to realize the transformation from the camera coordinate system to the VR global coordinate system to achieve an accurate mapping of the operator's actions in the VR environment to the robotic arm movement; the VR interaction module uses the VR headset and VR handle to adjust the state of the robotic arm according to the task to simplify the teaching process; the complex scene dataset production module is responsible for integrating and processing the teaching data and images of different scenes during the teaching process and transmitting them to the ACT algorithm layer for multi-modal image fusion.

[0080] ACT algorithm layer for multi-modal image fusion: The feature extraction backbone module extracts and fuses features from the image, and uses the gradient fusion attention module and the bimodal feature fusion module to optimize the feature fusion process, effectively increasing the detection accuracy. Further, the fused features are input into the ACT algorithm module to obtain the predicted action parameters, and the motion control strategy is used to drive the robotic arm to autonomously complete the task.

[0081] Specifically, it includes a feature extraction backbone module, a gradient fusion attention module, a bimodal feature fusion module, an ACT algorithm module, and a motion control strategy; the feature extraction backbone module is used to receive images of two modalities, infrared and visible light, as inputs, and fuse the feature maps of different modalities to obtain a fused feature map; the gradient fusion attention module captures the image texture information through the gradient operator, uses the class residual structure to realize the interaction of the gradient information of the two branches, and introduces the coordinate attention mechanism in the infrared branch to enable the model to better understand and utilize the complementary information of the infrared and visible light modalities; the bimodal feature fusion module is used to fuse the features extracted by the above two modules and transmit the fused feature map to the ACT algorithm module; the ACT algorithm module is used to generate the predicted robotic arm motion plan according to the generated fused feature map and the training of the teaching data; the motion control strategy is used to convert the predicted robotic arm motion plan into a specific robotic arm motion plan, thereby driving the robotic arm to autonomously complete the task.

[0082] As attached Figure 2As shown in the figure, both types of cameras are installed at the front end of the robotic arm to ensure that they have an overlapping field of view angle, that is, they are installed forward. Specifically, the visible light camera is installed on the right side of the front end of the robotic arm, and the infrared camera is installed on the left side of the visible light camera. Such a layout design ensures that there is sufficient overlapping field of view angle between different cameras, so as to be able to cover a wider detection range and improve the accuracy of target detection. Through this installation method, data from different cameras can be effectively fused to achieve precise detection and tracking of targets in complex scenarios. The navigation trolley and the robotic arm are combined as an overall experimental platform, which can power the robotic arm while enabling it to move, and then complete more complex tasks.

[0083] As shown in the appendix Figure 3 As shown, a robotic arm control method based on imitation learning provided in this embodiment, the data collection layer, includes the following steps:

[0084] Step S101: Set up the routing and VR device, the IP address and port number of the PC host server, query the robot CAN device ID, and set the PC host ID to be the same. After the server and CAN communication configuration are completed, the robot is powered on and runs;

[0085] Step S102: The operator observes the operating state of the robotic arm through the VR headset or remote monitoring device, and uses the VR handle to adjust the state of the robotic arm according to the picking task;

[0086] Step S103: As shown in the appendix Figure 4 As shown, establish a D-H coordinate model of the robotic arm according to the actual model of the robot and specific DH parameters;

[0087] Combined with the D-H parameters of each link and the expression of the transformation relationship between adjacent link coordinate systems Obtain the kinematic relationship between adjacent links, and calculate through the following formula:

[0088]

[0089] Among them, i is the robotic arm joint number, αi-1 is the link twist angle, d is the link offset, ai-1 is the link length, θ is the joint angle, is the transformation relationship matrix between adjacent joint coordinate systems, is the transformation relationship matrix of the robotic arm end relative to the base;

[0090] Step S104: Use the cyclic coordinate descent method (CCD algorithm) to solve for joint i, let P i be the coordinate origin of joint i, P e be the position of the robotic arm end, and P t be the position of the target point. Let the vector from P i to P e be Ve-i , P i From to P t The vector is V t-i . Then the vector V e-i And the vector V t-i The included angle θ i (Rotation angle) and the normal vector (Rotation direction) are:

[0091]

[0092] Step S105: Use quaternions to represent the poses of each joint. Let n be the axis of rotation, and define As a unit vector, rotating by θ around n, it is represented by quaternions as:

[0093]

[0094] Among them, q i Represents the azimuth angle of joint i in the reference coordinate system. Rotate q i The azimuth after rotating by θ around n is q iq , and the joint angles of each joint can be calculated.

[0095] Step S106: As shown in the appendix Figure 5 , perform the conversion between VR coordinates and robot coordinates. The transformation matrix Between the binocular camera coordinate system and the robot base coordinate system is obtained through hand-eye calibration. The transformation matrix From the camera coordinate system to the VR global coordinate system is divided into: the transformation matrix From the camera coordinate system to the VR glasses coordinate system The transformation from the VR glasses coordinate system to the VR global coordinate system

[0096]

[0097] The transformation from the VR glasses coordinate system to the VR global coordinate system Is calculated by the following formula:

[0098]

[0099] Among them, R i (θ) is the augmented matrix composed of the rotation matrix obtained by rotating by θ around the i-axis and the position coordinates.

[0100] The transformation matrix From the camera coordinate system to the VR glasses coordinate system is calculated by the Rodrigues formula:

[0101]

[0102] Among them, gc , g g are the gravity vectors of the camera and the VR glasses, and n i is n x , n y , n z , g c , g g is the common rotation axis when the two coordinate systems coincide, and θ is the included angle between the two vectors.

[0103]

[0104] Among them, the identity matrix A provides the initial state of rotation, and B actually encodes the direction information of the rotation axis. and are known, and the transformation matrix from the camera coordinate system to the VR global coordinate system can be solved. Thus, the pose information of the target point in the VR handle coordinate system can be transmitted to the robotic arm controller through the transformation matrix, and Steps S103, 104, and 105 are called to perform kinematic analysis on the input pose, and the robotic arm is driven in combination with the robotic arm Python SDK.

[0105] Step S107: The operator observes the running state of the robotic arm in real time through the VR headset or the remote monitoring device for streaming, adjusts the VR handle operation to control the robotic arm to complete the picking task, and the PC host continuously collects the teaching data to make a data set.

[0106] As shown in the appendix Figure 6 , the feature extraction backbone module in the ACT algorithm layer of the multimodal image fusion includes the following steps:

[0107] Step S201: Use two convolutional layers to perform shallow feature extraction on the infrared and visible light images {O, C} of two modalities to obtain two feature maps {F O , F C} of different modalities.

[0108] Step S202: Use the GFAM module to obtain a feature map with modality characteristics through the interaction of the gradient information of the two branches and the feature extraction for the characteristics of each modality.

[0109] Step S203: Obtain the fused feature map through the DFFM module. After that, it passes through the subsequent network of the yolov5-s original model for target classification and localization. The process of obtaining the fused feature map can be described as:

[0110] F O = S2(O), F C = S2(C)(12)

[0111]

[0112] Among them, S i (·) represents the function of the i-th layer of Conv, G(·) represents the function of the GFAM module, and D(·) represents the function of the DFFM module.

[0113] As shown in the attached Figure 7 In the ACT algorithm layer of the multi-modal image fusion shown, the gradient fusion attention module includes the following steps:

[0114] Step S301: Use a class residual structure for gradient fusion interaction, where the mainstream part uses ordinary CBL convolution combinations, including a 3×3 convolutional layer, a regularization BN layer, and a LeakyRelu activation function. The residual part uses two-way gradient operator convolutions and a point-wise convolution (Point-wise Conv, PW). In the gradient operator convolution, a gradient operator is introduced to convolve the input features with a high-frequency convolutional kernel to obtain the gradient information of the image and capture the texture details of the object. This process can be expressed as:

[0115]

[0116] Among them, F′ C is the output visible light feature map, ▽ refers to the gradient operator, PWConv(·) represents the point-wise convolution operation, C(·) represents the Concat splicing operation, represents element-wise summation. This part aggregates learnable convolutional features with gradient magnitude information.

[0117] Here, the Sobel operator is used to calculate the gradient magnitude, and the used Sobel gradient operator is:

[0118]

[0119] Among them, G X and G Y are the gradient information obtained from the feature map F in the horizontal and vertical directions respectively;

[0120] Step S302: In the processing of the infrared branch, use the Coordinate Attention (CA) mechanism to introduce spatial coordinate information to enhance the feature representation, capture the important spatial positions and inter-channel dependencies in the feature map, and further extract the texture features of the infrared image.

[0121] Step S303: In the visible light branch processing, use the C3CrossCovn module composed of CrossConv for feature extraction and multi-level fusion to further extract the texture features of the visible light image. This module consists of two standard convolutional layers arranged in a cross pattern. The first cross convolution uses a kernel of size 1×k, with a horizontal stride of 1 and a vertical stride of s. The second uses a kernel of size k×1, with a stride of s in both dimensions.

[0122] As shown in the Figure 8 appendix, the dual-modal feature fusion module in the ACT algorithm layer of the multi-modal image fusion includes the following steps:

[0123] Step S401: Perform a preliminary element-wise addition fusion on the input features X and Y of the two modalities, and then pass them through the WGM module for processing in two different dimensions respectively.

[0124] Step S402: In the importance analysis of the left channel dimension, use global average pooling to compress the feature map of size C×H×W into C×1×1, and use two lightweight PW convolutions to learn the attention information of the channel dimension. In the attention analysis of the right spatial dimension, use PW convolution to perform dimensionality reduction on the input feature map, and further use two dilated convolutions with a kernel size of 3×3 to extract feature information, and map the feature map to 1×H×W through PW convolution.

[0125] Step S403: Use the Sigmoid activation function to make the output value between 0 and 1 to obtain the fusion weight value W of feature X, where the weight of feature Y is 1 - W. Adaptively adjust the weights W and 1 - W to provide the learning ability of the model during backpropagation.

[0126] Step S404: Use the broadcast mechanism to enable the addition of the attention information on both sides, multiply the weights with features X and Y respectively and then perform element-wise addition to generate the fusion feature Z. Among them, the fusion feature Z is obtained by flattening the feature map obtained by the above feature extraction processing of the RGB image and optimizing to obtain a feature sequence.

[0127] As shown in the Figure 9 appendix, the ACT algorithm module in the ACT algorithm layer of the multi-modal image fusion includes the following steps:

[0128] Step S501: Process the fusion feature Z through the encoder of the autoencoder (CVAE) to obtain the style variable α. Among them, the style variable α is a low-dimensional representation obtained by processing the action sequence and joint positions of the fusion feature Z through the encoder.

[0129] Step S502: Integrate the style variable α with the feature Z to obtain a comprehensive feature sequence. The feature sequence is a series of feature vectors containing information such as joint positions.

[0130] Step S503: Use the decoder of the convolutional variational autoencoder (CVAE) to reconstruct the comprehensive feature sequence using the L1 loss, which is the sum of the absolute values of the differences between the corresponding elements in the two sequences, to accurately model the action sequence and generate a predicted action sequence. The action sequence is a series of action instructions arranged in chronological order and consists of a series of feature vectors containing information such as joint positions. The L1 loss function used is:

[0131]

[0132] where Y i is the actual input value, is the predicted value of the model, and n is the number of samples;

[0133] Step S504: As shown in the appendix Figure 10 transmit the predicted action sequence to the motion control strategy part to enable the robotic arm to autonomously complete the task.

[0134] Example 2:

[0135] This embodiment provides an experimental test of a robotic arm control method based on imitation learning:

[0136] This experiment selects common application scenarios that may occur in reality for actual interaction testing of the ACT algorithm. It includes three task scenarios: fruit picking, autonomous door opening, and desktop organization, and tests the completion effects of each task under sufficient light and insufficient light conditions.

[0137] Both fruit picking and autonomous door opening require cooperation with a navigation cart. After the cart moves to the target position, the task is carried out. For fruit picking, it mainly tests the flexibility of the robotic arm to pick the fruit without damaging the fruit and avoiding the tree trunk and put it into the collection bucket. The autonomous door opening task requires higher precision of the robotic arm because the doorknob is a shiny stainless steel product with a small volume. In this paper, the robotic arm can compensate for the motion error caused by the doorknob by introducing a multi-modal image fusion algorithm for visual servoing. The items on the desktop for organization are a toy pear, an alcohol bottle, and a remote control handle, representing the organization of a spherical object, a cylindrical object, and an irregular object respectively, to verify the applicability when grasping different objects. The above performances are all reflected in the form of task success rates.

[0138] Appendix Figure 11For fruit picking demonstration, *1 is the combination of a robotic arm and a navigation cart, and *2 is the autonomous navigation of the cart towards the fruit tree. *3--*7 are the sub-actions of the entire picking process under sufficient light: initial state, approaching the fruit, grasping the fruit, picking the fruit, fruit collection, and reset; *8--*12 are the entire picking process under insufficient light, and the sub-actions are the same as those under sufficient light. Attachment Figure 12 For autonomous door opening demonstration, the sub-actions of *1--*6 are respectively the initial state, approaching the doorknob, grasping the doorknob, pulling down the doorknob, opening the door, and reset. *7--*12 are the door opening process under insufficient light, and the sub-actions are the same as those under sufficient light. Attachment Figure 13 For desktop storage demonstration, the sub-actions of *1--*6 are respectively the initial state, approaching the storage object, storing spherical objects, storing cylindrical objects, storing irregular objects, and closing the storage box. *7--*12 are the storage process under insufficient light, and the sub-actions are the same as those under sufficient light.

[0139] Four imitation learning algorithms, namely training VINN and Diffusion Policy, the original ACT and the multi-modal image fusion ACT in this paper, are respectively deployed in the robotic arm to perform the above three tasks. Among them, Diffusion Policy uses the DDIM scheduler and image enhancement to improve the inference speed and fitting discrimination degree. For VINN, the weights are adjusted to balance the emphasis on visual and action features and it is also optimized with fast actions. While ACT and the improved ACT in this paper calculate the sampling probability according to the dataset size, so that the data in the source domain and the target domain can be evenly mixed, thereby constructing a data loader that combines dynamic and static data for collaborative training. The success rate of each link of each task is obtained by evaluating and calculating 50 executions. The execution of the four algorithms in the three tasks is shown in Table 1.

[0140] Table 1 Comparative experiments for performing three tasks

[0141]

[0142] Generally speaking, in sufficient sunlight, Diffusion Polic, ACT, and multi-modal image fusion ACT have good success rates in each task link, and ACT and multi-modal image fusion ACT almost reach more than 90%. However, when the light is insufficient, except for the algorithm in this paper, their success rates all decrease by 7% to 30%, generally about 10%. The success rates of VINN and DiffusionPolic in actions such as picking fruits are less than 40%, which are in an unusable state. The success rate of the algorithm in this paper decreases by 0% to 6% when the light is insufficient, and nearly half of the action success rates do not decrease.

[0143] During the fruit picking task, the success rates of the VINN and Diffusion Polic algorithms in picking fruits are generally low. We observed that after grasping the fruit, when the robotic arm moves downward to pick the fruit but fails to pick it, the action will pause. After waiting for a period of time, it may continue to pick the fruit until it is picked or remain paused indefinitely. We believe this is due to the combined errors of VINN and Diffusion Polic themselves, and even with action block optimization, this impact cannot be completely offset. The success rates of the fruit collection actions are generally high. The analysis shows that because a large red bucket is used as the collection device, its features are obvious and the training effect is good.

[0144] In the autonomous door opening task, the success rates of pulling down the doorknob are generally low. We observed that when the robotic arm pulls down the doorknob, the actuator will slip off the doorknob. When this slipping occurs, the robotic arm often grabs the doorknob at a relatively edge position. The main reasons are that the doorknob is smooth, reflective, and small in volume, and the number of teleoperation demonstrations is too small. Eventually, the grasping position of the doorknob in the previous action is not good, and it slips off when pulled down. Moreover, this link is greatly affected by light, and the success rate generally drops by more than 14%.

[0145] In the autonomous door opening task, the success rates of pulling down the doorknob are generally low. We observed that when the robotic arm pulls down the doorknob, the actuator will slip off the doorknob. When this slipping occurs, the robotic arm often grabs the doorknob at a relatively edge position. The main reasons are that the doorknob is smooth, reflective, and small in volume, and the number of teleoperation demonstrations is too small, which eventually leads to a poor grasping position of the doorknob and it slips off when pulled down.

[0146] The embodiments described in this invention patent are described one by one in a progressive manner. Each embodiment mainly elaborates on its differences from other embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, refer to the description in the method section. For the disclosed embodiments, those skilled in the art can implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

[0147] In summary, the present invention provides a manipulator control method based on imitation learning. This method uses a VR device for teleoperation to collect teaching data, adopts infrared-visible light feature information fusion for image data processing, and combines the ACT algorithm for feature extraction and processing to improve the motion accuracy and robustness of the manipulator in complex environments. Through detailed embodiments and technical solutions, the present invention demonstrates its feasibility and effectiveness in practical applications, and has broad application prospects and market potential.

[0148] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A manipulator control method based on imitation learning, characterized in that: The specific steps are as follows: Step 1: Configure the robot and run it. Collect the operator's teaching data through the VR interaction module, including the pose information of using the handle to control the robotic arm to complete tasks; Step 2: Establish a D-H coordinate model of the robotic arm using the actual model of the robot. Combine the D-H parameters of each link and the expression of the transformation relationship between adjacent link coordinate systems to obtain the forward and inverse kinematic relationships of the robotic arm; Step 3: Map the VR handle pose information to the base coordinate system of the robotic arm. Combine the cyclic coordinate descent method to solve the joint angles of the robotic arm, and drive the movement of the robotic arm through the robotic arm Python SDK; Step 4: Use the binocular camera to collect environmental image data. Calculate the transformation matrix from the camera coordinate system to the VR global coordinate system through hand-eye calibration and the Rodrigues formula, and synchronously construct a complex scene dataset containing multi-modal images and teaching actions; Step 5: Perform shallow feature extraction on the infrared image and visible light images of the two modalities respectively to generate an infrared feature map and a visible light feature map; Step 6: Extract the gradient texture features of the infrared and visible light images through the gradient fusion attention module, and combine the coordinate attention mechanism to enhance the complementary information of the infrared and visible light modalities; Step 7: Perform cross-modal fusion on the infrared and visible light features through the dual-modal feature fusion module to generate a fusion feature map; Step 8: Input the fusion feature map into the ACT algorithm module, and use the conditional variational autoencoder to generate the predicted action sequence of the robotic arm; Step 9: Transmit the predicted action sequence to the motion control strategy part to enable the robotic arm to autonomously complete tasks.

2. The robotic arm control method based on imitation learning according to claim 1, characterized in that: The above step 2 is implemented by a robotic arm kinematic analysis module. A D-H coordinate model of the robotic arm is established based on the actual model of the robot and specific DH parameters. Combining the D-H parameters of each connecting rod and the expression of the transformation relationship between adjacent connecting rod coordinate systems The kinematic relationship between adjacent connecting rods is obtained, and the formula is as follows: where i is the serial number of the robotic arm joint, α i-1 is the link twist angle, d is the link offset, a i-1 is the link length, θ is the joint angle, is the transformation relation matrix of adjacent joint coordinate systems, is the transformation relation matrix of the robotic arm end relative to the base.

3. The robotic arm control method based on imitation learning according to claim 1, characterized in that: Step 4 obtains the transformation matrix from the camera coordinate system to the robot base coordinate system through hand-eye calibration Calculate the transformation matrix from the camera coordinate system to the VR global coordinate system through the Rodrigues formula The formula is as follows: Among them, is the transformation matrix from the camera coordinate system to the VR glasses coordinate system, is the matrix for transforming the VR glasses coordinate system to the VR global coordinate system; Among them, Among them, R i (θ) is an augmented matrix composed of a rotation matrix obtained by rotating by an angle θ about the i-axis and a position coordinate; Calculated by Rodrigues' formula where g c and g g are the gravity vectors of the camera and the VR glasses, n i is n x , n y , and n z , which is the common rotation axis when the two coordinate systems of g c and g g coincide; θ is the included angle between the two vectors. Among them, the identity matrix A provides the initial state of rotation, and B actually encodes the direction information of the rotation axis. and Given this, the transformation matrix from the camera coordinate system to the VR global coordinate system can be solved.

4. A robotic arm control method based on imitation learning according to claim 1, characterized in that: The specific content of Step 6 is as follows: (1) Use a class residual structure with gradient fusion interaction. The mainstream part uses ordinary CBL convolution combinations, including a 3×3 convolutional layer, a regularization BN layer, and a LeakyRelu activation function. The residual part uses two-way gradient operator convolutions and a pointwise convolution. In the gradient operator convolution, a gradient operator is introduced to convolve the input features with a high-frequency convolutional kernel to obtain the gradient information of the image and capture the texture details of the object. This process is expressed as: Among them, F C ′ is the output visible light feature map, refers to the gradient operator, PWConv(·) represents the pointwise convolution operation, and C(·) represents the Concat splicing operation, represents element-wise summation. This part aggregates learnable convolution features with gradient magnitude information. Here, the Sobel operator is used to calculate the gradient magnitude, and the Sobel gradient operator used is: Among them, G X and G Y are the gradient information obtained from the feature map F in the horizontal and vertical directions respectively; (2) In the infrared branch processing, use the coordinate attention mechanism to introduce spatial coordinate information to enhance the feature representation, capture the important spatial positions and dependencies between channels in the feature map, and further extract the texture features of the infrared image; (3) In the visible light branch processing, use the C3CrossCovn module composed of cross convolutions for feature extraction and multi-level fusion, and further extract the texture features of the visible light image; this module consists of two standard convolutional layers arranged in a cross pattern. The first cross convolution uses a 1×k kernel, with a horizontal stride of 1 and a vertical stride of s. The second uses a k×1 kernel with a stride of s in both dimensions.

5. A manipulator control method based on imitation learning according to claim 1, characterized in that: The fused feature map is generated in the step 7 as follows: F O = S2(O), F C = S2(C) Among them, О is the infrared image, and С is the visible light image. is the gradient texture feature of the infrared image. is the gradient texture feature of the visible light image, S i (.) represents the role of the i-th layer of Conv, G(.) represents the role of the GFAM module, and D(.) represents the role of the DFFM module; the fused feature map is obtained. After that, it passes through the subsequent network of the original yolov5-s model for object classification and localization.

6. The robotic arm control method based on imitation learning according to claim 5, characterized in that: The specific content of the dual-modal feature fusion is as follows: (1) Perform preliminary element-wise addition fusion on the features X and Y of the two input modalities, and then the WGM module makes them go through two different-dimensional processes respectively; (2) In the importance analysis of the left channel dimension, global average pooling is adopted to compress the feature map of size C×H×W into C×1×1, and two lightweight PW convolutions are used to learn the attention information of the channel dimension; In the attention analysis of the right spatial dimension, the PW convolution is used to perform dimensionality reduction on the input feature map, and two dilated convolutions with a kernel size of 3×3 are further used to extract feature information, and the feature map is mapped to 1×H×W through the PW convolution; (3) The Sigmoid activation function is used to make the output value between 0 and 1 to obtain the fusion weight value W of feature X, where the weight of feature Y is 1 - W; the weights W and 1 - W are adaptively adjusted to provide the learning ability of the model during backpropagation; (4) The broadcast mechanism is used to enable the addition of the attention information on both sides. The weights are multiplied by features X and Y respectively and then added element by element to generate the fused feature Z; among them, the fused feature Z is obtained by flattening the feature map obtained by the above-mentioned feature extraction process of the RGB image and optimizing to obtain a feature sequence.

7. A manipulator control method based on imitation learning according to claim 1, characterized in that: The specific content of step 8 is as follows: (1) The fused feature Z is processed by the encoder of the autoencoder to obtain the style variable α; among them, the style variable α is a low-dimensional representation obtained by the encoder processing the action sequence and joint positions of the fused feature Z; (2) The style variable α is fused with the feature Z to obtain the comprehensive feature sequence; among them, the feature sequence is a series of feature vectors containing information such as joint positions; (3) Using the decoder of the autoencoder, the L1 loss is used for reconstruction of the comprehensive feature sequence. It is the sum of the absolute values of the differences between the corresponding elements in the two sequences, and the action sequence is accurately modeled to generate the predicted action sequence; among them, the action sequence is a series of action instructions arranged in chronological order and consists of a series of feature vectors containing information such as joint positions; the L1 loss function used is: Among them, Y i is the actual value of the input, is the predicted value of the model, and n is the number of samples.

8. A robotic arm control system based on imitation learning, characterized in that: It includes a data collection layer and an ACT algorithm layer for multimodal image fusion; The main function of the data collection layer is to collect the teaching data of the operator through the VR interaction module, including the pose information of the robotic arm, and convert this data into executable instructions for the robotic arm through coordinate transformation and kinematic analysis; at the same time, this layer is also responsible for integrating the teaching data and images of different scenarios to construct a complex scenario dataset; The main function of the ACT algorithm layer for multimodal image fusion is to process the multimodal images of the data collection layer, extract and fuse features through the gradient fusion attention module and the bimodal feature fusion module, and finally generate the predicted action sequence of the robotic arm using the ACT algorithm and drive the robotic arm to execute tasks through the motion control strategy.

9. A robotic arm control system based on imitation learning according to claim 8, characterized in that: The data collection layer includes a robotic arm motion analysis module, a coordinate transformation module, a VR interaction module, and a complex scenario dataset production module; the robotic arm motion analysis module uses the actual model of the robot to establish the D-H coordinate model of the robotic arm, and combines the D-H parameters of each link and the expression of the coordinate system transformation relationship between adjacent links to obtain the forward and inverse kinematic relationships of the robotic arm; The coordinate transformation module uses a coordinate system transformation matrix to achieve the transformation from the camera coordinate system to the VR global coordinate system, so as to accurately map the actions of the operator in the VR environment to the movement of the robotic arm; the VR interaction module uses a VR headset and a VR handle to adjust the state of the robotic arm according to the task, so as to simplify the teaching process. The complex scene dataset production module is responsible for integrating and processing the teaching data and images of different scenes during the teaching process, and transmitting them to the ACT algorithm layer for multi-modal image fusion.

10. A robotic arm control system based on imitation learning according to claim 8, characterized in that: The ACT algorithm layer for multi-modal image fusion includes a gradient fusion attention module, a bimodal feature fusion module, an ACT algorithm module, and a motion control strategy module; the gradient fusion attention module captures the image texture information through a gradient operator, and uses a residual-like structure to realize the interaction of the gradient information of the two branches. At the same time, a coordinate attention mechanism is introduced into the infrared branch, so that the model can better understand and utilize the complementary information of the infrared and visible light modalities; the bimodal feature fusion module is used to fuse the infrared and visible light features extracted above, and transmit the fused feature map to the ACT algorithm module; the ACT algorithm module is used to generate a predicted robotic arm motion plan according to the generated fused feature map and the training of the teaching data; the motion control strategy module is used to convert the predicted robotic arm motion plan into a specific robotic arm motion plan, so as to drive the robotic arm to autonomously complete the task.

Citation Information

Patent Citations

  • Simulation learning fruit picking method and device based on multi-modal information fusion

    CN116985132A

  • Infrared-visible light image fusion-based integrated management and control method for grid field operation

    WO2024183245A1

Cited By

  • Picking robot man-machine interaction system and method based on machine vision

    CN120959045A

  • Method and system for staged fusion control of bunched tomato picking robot

    CN122480999A