Manipulator Control Method and System Based on Improved Swin-Unet

The SU-Grasp model addresses the limitations of existing 3D vision algorithms by integrating dual-stream encoding in Swin-Unet to enhance depth information utilization and perception, improving grasping accuracy and reducing computational complexity.

CN119600095BActive Publication Date: 2025-07-15YUNNAN YUNLING EXPRESSWAY BRIDGE ENG CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411436248.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2024-08-12
Filing Date
2024-10-15
Publication Date
2025-07-15
Estimated Expiration
2044-10-15

AI Technical Summary

Technical Problem

The existing three-dimensional visual grasping algorithms are insufficiently utilized in complex environments, and the local and global perception capabilities are weak, resulting in insufficient accuracy of robotic arm grasping.

Method used

The improved Swin-Unet model is used to construct a capture pose prediction model SU-Grasp based on the three-dimensional visual dual-stream coding strategy. By converting RGB color images and depth images into normal vector angle images for stitching, the dual-stream encoder is used to extract multi-scale features, and the information interaction is enhanced through cross-modal fusion, solving the problem of insufficient utilization of depth information.

Benefits of technology

It improves the accuracy and efficiency of the robotic arm in complex environments, enhances local and global perception capabilities, reduces algorithm complexity and speeds up the training process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119600095B_ABST
    Figure CN119600095B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical fields of computer vision and robotic arm automation, and discloses a robotic arm control method and system based on an improved Swin-Unet. The method comprises the following steps: S0, obtaining an RGB color image and a depth image of a target working scenario; S1, converting the depth image into a normal vector angle image, and performing a splicing operation on the two to obtain a spatial image; S2, using a grasping model to analyze the RGB color image and the spatial image to obtain grasping parameters; the grasping model is constructed based on an improved Swin-Unet and comprises: an encoding part, respectively extracting features of multiple set scales of the RGB color image and the spatial image and performing feature fusion; a decoding part, gradually upsampling the fused features of the smallest scale to the input resolution, and fusing the fused features of the corresponding scales through skip connections, so as to predict and obtain the grasping parameters; S3, controlling a robotic arm to grasp a target object based on the grasping parameters. The present invention can accurately determine the position information of a target object and complete the grasping in a multi-object cluttered scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and robotic arm automation, and particularly relates to a robotic arm control method and system based on an improved Swin-Unet. Background Art

[0002] In recent years, with the continuous development of automation technology, intelligent robots are gradually expanding from the industrial field to all walks of life. Among them, the automated application of civil engineering robots has become an emerging topic recently. By extracting and segmenting the visual image features collected, automated operation of the robotic arm is realized, such as rock drilling robots, hoisting robotic arms, building decoration robots, and so on. However, most traditional robots complete fixed assembly line tasks through pre-programming or manual teaching, and it is difficult to adapt to more complex construction site environments. Grasping an object, as an important function of a robot, its accuracy and practicability to a certain extent reflect the intelligent level of the robot. Studying reliable and efficient grasping algorithms will become the key to improving the automation level of construction robots.

[0003] According to the type of input data, the existing end-to-end grasping prediction algorithms can be divided into the following three types:

[0004] (1) Grasping pose detection method based on two-dimensional images: This method requires pre-acquiring the geometric attributes of the object, and highly depends on the simplification of the physical model and the assumption of a fully observable environment, resulting in limited application scenarios, single grasping object, and poor generalization. Compared with two-dimensional color images, three-dimensional vision can additionally reflect the spatial depth information of the target object, and can also identify objects without texture and with occlusion. For the disordered grasping scenario of scattered objects, where the heights and volumes of each target are different, a three-dimensional vision grasping system must be used.

[0005] (2) Grasping pose detection method based on point cloud segmentation: This method can help the robotic arm more accurately identify specific objects, but it is necessary to first obtain high-precision point cloud data, which has high requirements for hardware devices and is currently not convenient to popularize to daily robotic arm grasping tasks.

[0006] (3) Grasping pose distribution map prediction method based on RGBD three-dimensional vision: For an input RGBD image, a model is constructed to output a set of prediction values for each pixel therein, including grasping confidence and grasping parameters (such as grasping depth, width, and rotation angle). This method can achieve good results, but there are still the following deficiencies:

[0007] (a) To find the graspable area, the robot not only has to focus on the local geometric information of the object, but also consider its entire visual appearance. Especially in a cluttered unstructured environment, dealing with different shapes, positions, and the spatial relationships between objects is the difficulty of grasping detection.

[0008] (b) Existing 3D vision algorithms often regard the depth map as a grayscale image, splice it with the three color channels to form an input vector of RGBD four channels, and apply various 2D vision algorithms on this basis, such as Unet, ViT, Swin Transformer, etc., to use depth information to assist in entity segmentation. However, in fact, traditional 2D vision algorithms may not be suitable for direct application in the processing of depth images. There are certain differences in the representation of semantic information between RGB color images and depth images. The depth image captures depth information and a two-dimensional grid structure, and contains implicit local set features and directional information, which is more like a complementary relationship with RGB color images. In addition, the acquisition of depth images is generally through the camera emitting infrared lasers and receiving the corresponding reflected signals, which is easily interfered by environmental light, smooth surfaces of objects, etc., so there are more noises and missing values. At this time, directly splicing with RGB may have the opposite effect.

[0009] Therefore, how to make full use of depth information to improve the performance of pose estimation is still a difficult and under-researched problem. Summary of the Invention

[0010] Based on the existing problems, the present invention provides a manipulator control method and system based on an improved Swin-Unet, which can solve the problems of insufficient utilization of depth information and weak local and global perception ability during 3D vision grasping of the manipulator.

[0011] The technical solution adopted by the present invention is as follows:

[0012] The manipulator control method based on an improved Swin-Unet adopts the architecture of the improved Swin-Unet model, and constructs a grasping pose prediction model framework SU-Grasp based on a three-dimensional vision dual-stream encoding strategy based on a four-layer Swin-Unet network; the entire network architecture includes 5 modules, namely a point cloud preprocessing module, a dual-stream encoder, a bottleneck layer, a decoder, and a skip connection layer;

[0013] The manipulator control method based on an improved Swin-Unet includes the following steps:

[0014] Step S0: Obtain the RGBD image of the target working scene, that is, the data containing the RGB color image and the depth image, and use it as the model input;

[0015] Step S1: Convert the depth image into a normal vector angle image, and splice the depth image and the normal vector angle image to obtain a spatial image;

[0016] Step S2: Analyze the RGB color image and the spatial image using a grasping model to obtain grasping parameters; regard the RGB two-dimensional image as the color semantic modality, and regard the depth map D and the normal angle image I as the spatial semantic modalities, which are used as the inputs of the SU-Grasp network;

[0017] Step S3: Control the robotic arm to grasp the target object based on the grasping parameters.

[0018] Furthermore, in the architecture of the Swin-Unet model, the encoder, bottleneck layer, and decoder are constructed based on the Swin Transformer Block model; the Swin-Unet network structure divides the input image into non-overlapping image patches, which are input into the encoder to learn deep feature representations; after extracting the context features, the decoder with an image patch expansion layer performs upsampling and is connected to the same-level encoder through skip connections to gradually restore the spatial resolution of the feature map for pixel-level segmentation prediction; use the Swin-Unet model as the basic network framework of the SU-Grasp model to obtain both local information and long-range connections at the same time;

[0019] During the implementation of the SU-Grasp model, three grasping outputs are connected in parallel to the top of the decoder, representing the grasping confidence Q, grasping width W, and grasping angle Θ respectively, and each output is a heat map with the same size as the input RGBD image.

[0020] Furthermore, in step S0, the data includes the Cornell dataset and the Jacquard dataset; step S0 also includes preprocessing the depth image, and the preprocessing includes plane segmentation operations based on the Random Sample Consensus (RANSAC) algorithm and outlier removal operations based on radius-nearest neighbor search.

[0021] Furthermore, in step S1, the normal vector angle image is introduced, and the 3D vision is divided into two semantic modalities: color modality and spatial modality, and the information interaction between the two is enhanced through a multi-modal fusion method; during the grasping process in the spatial modality: first, determine the overall position and pose of the object to be grasped; second, provide local geometric information on the object surface to assist in making decisions on the most stable grasping points.

[0022] Furthermore, in step S1, call the Open3D library to convert the preprocessed depth map into a normal vector angle image; the calculation process of the normal vector angle image is as follows:

[0023] Calculate the unit normal vector Ν(n x ,n y ,n z ) corresponding to the object surface for each pixel (u, v) in the depth image D;

[0024] Determine the angles between N and the three axes X, Y, and Z of the camera coordinate system;

[0025] Let the three axes X, Y, and Z be represented as vectors x(1, 0, 0), y(0, 1, 0), z(0, 0, 1) respectively, and the angles between Ν(n x , n y , n z ) and these three axis vectors can be expressed as:

[0026] a x = arccos(N · x)

[0027] a y = arccos(N · y)

[0028] a z = arccos(N · z)

[0029] Wherein, N is the unit normal vector of the object surface corresponding to the current pixel; x, y, z represent the unit normal vectors in the camera coordinate system that are in the same direction as the positive directions of the three coordinate axes; a x , a y , a z correspond to the angles between N and the x, y, z axes; the arccos function is used to calculate the cosine value of the angle between two vectors.

[0030] Normalize the angles (a x , a y , a z ) to the range of 0 to 255 for creating a three-channel normal vector angle image I nrm , then each pixel has the normalized (a x , a y , a z ) as its value.

[0031] Furthermore, in step S1, introduce the normal vector angle image into the robotic arm grasping task to explicitly express the hidden semantics of the depth information in the directionality of the object surface; regard the depth image and the RGB image as two modalities, the former describes the depth difference and structural information between different objects, and the latter reflects the appearance information including the color, shape, and boundary of a scene; adopt a two-stream structure to exploit the respective advantages of different modalities and their complementary properties;

[0032] The dual-stream structure uses two encoders to process inputs of different modalities respectively, and then uses a network architecture to perform cross-modal information fusion; the depth image D and the normal vector angular image are regarded as the same modality and spliced into a four-channel spatial image vector, recorded as a DI image; the color information and spatial information are first encoded separately, and then the extracted information is integrated through cross-modal fusion to enhance the semantic representation of the object position and shape of the two, and finally complete the subsequent decoding operation.

[0033] Furthermore, the grasping model in step S2 is constructed based on the improved Swin-Unet, including: the encoding part extracts multiple set scale features of the RGB color image and the spatial image respectively to obtain color features and spatial features; then, the color features and spatial features of the same scale are feature-fused to obtain a plurality of first fused features; then, the first fused features of the minimum set scale are extracted to obtain second fused features; the decoding part predicts and obtains grasping parameters according to the first fused features and the second fused features;

[0034] In step S2, the multi-level structure of the Swin Transformer decoder is used to transform the encoder. The specific logic of the algorithm is as follows:

[0035] Set the hyperparameter C of the input→stage2 part of the original Swin-Unet encoder to 1 / 2 of the original one, and copy the structure to get two sub-encoders;

[0036] The RGB color image and the spatial image obtained by concatenating the depth image and the normal vector angle image are input into the two sub-encoders respectively; after one convolution operation, the data size in the two sub-encoders becomes

[0037] The two sub-encoders complete stage 1 and 2 respectively, and up-sample through two Swin Transformer Blocks. At this time, the feature dimension of the two outputs is 2C. Before performing stage 3 calculation, the data in the two sub-encoders are concatenated to restore the embedding dimension of the current feature map to 4C.

[0038] Continue to complete the calculation of stage3 and bottleneck layer in the encoder;

[0039] In the decoder, the jump connection operation maintains the first layer consistent with the original Swin-Unet structure, and concatenates the outputs of the two same-level sub-encoders and the decoder feature maps of the previous level in the second and third layers, as shown in the following formula:

[0040]

[0041] In the formula, respectively represent the input of the nth-level decoder and the output of the (n - 1)th-level decoder; and respectively represent the output of the two-stream sub-encoders corresponding to the nth-level decoder;

[0042] Select stage3 as the information fusion position of the two-stream structure; respectively extract features of multiple set scales of the RGB color image and the spatial image through two sub-encoders based on the Swin-Unet structure to obtain color features and spatial features; the lengths of the patch embedding vectors of the two sub-encoders are kept the same.

[0043] Furthermore, the process of grasping parameters in step S2 includes: upsample the second fusion feature of the smallest scale to the input resolution level by level; in each level of upsampling operation, fuse the first fusion feature or the second fusion feature of the corresponding scale through skip connections, and then predict and obtain the grasping parameters;

[0044] The grasping parameters include the grasping position, and the grasping confidence, grasping width, and grasping angle at the grasping position; the output of the SU-Grasp network is the heatmaps (Q, W, Θ) of the grasping center point confidence, grasping width, and grasping angle; when determining the actual grasping position, retrieve the position with the highest grasping confidence through the grasping quality heatmap Q, defined as:

[0045]

[0046] In the formula, represents the grasping position with the highest grasping confidence, that is, the best grasping position recommended by the model; argmax pos represents finding the independent variable that makes the subsequent expression take the maximum value; Q represents the grasping confidence heatmap;

[0047] Then, extract the predicted width w and angle θ at the corresponding position from the predicted heatmaps of the grasping width W and the grasping angle Θ;

[0048] Regard the grasping pose estimation as a regression problem, that is, establish a mapping relationship F∶ through the SU-Grasp model where I represents the input information, specifically referring to the RGBDI nrm seven-channel vector, while represents the predicted heatmaps (Q, W, Θ) output by the model;

[0049] Based on minimizing the prediction distance, define the loss function as:

[0050]

[0051] In the formula, represents the total loss value; N represents the number of samples; m represents different output metrics; w m represents the weight of each predicted heatmap loss value, with a default value of 1 for all; is the predicted value of the model, representing the prediction result of the i-th sample under the metric m; is the true label data, representing the actual result of the i-th sample under the metric m; ||····|| represents the norm of the vector, usually the L2 norm, indicating the gap between the predicted value and the true value.

[0052] Furthermore, step S3 includes: retrieving the grasping position with the highest grasping confidence to obtain the target grasping position; extracting the grasping width and grasping angle of the target grasping position to obtain the target grasping width and target grasping angle; controlling the robotic arm to grasp the target object according to the target grasping position, target grasping width, and target grasping angle.

[0053] A robotic arm control system based on the improved Swin-Unet, based on the above robotic arm control method, includes:

[0054] An input module that acquires the RGB image and depth image of the target working scenario;

[0055] A processing module that converts the depth image into a normal vector angle image and performs a splicing operation on the depth image and the normal vector angle image to obtain a spatial image;

[0056] A grasping model that analyzes the RGB image and the spatial image using the grasping model to obtain grasping parameters;

[0057] A control module that controls the robotic arm to grasp the target object based on the grasping parameters;

[0058] Among them, the grasping model is constructed based on the improved Swin-Unet and includes:

[0059] An encoding part that respectively extracts the features of multiple set scales of the RGB color image and the spatial image to obtain the image color features and spatial features, and performs feature fusion on the image color features and spatial features of the same scale to obtain several first fusion features, and then performs feature extraction on the first fusion features of the smallest set scale to obtain second fusion features; A decoding part that predicts the grasping parameters according to the first fusion features and the second fusion features.

[0060] The beneficial effects of the present invention are:

[0061] (1) Aiming at the problem of insufficient spatial understanding ability in the existing robotic arm vision grasping algorithm, the present invention introduces a two-stream encoding mechanism based on sliding window attention and spatial semantic enhancement into the U-shaped network model. Two independent sub-encoders are used to extract multi-scale features of RGB color images and DI spatial images respectively, and cross-modal information fusion is achieved through concatenation. Compared with directly inputting seven-channel data, the two-stream encoder structure can first separately extract the shallow semantic features of each modality in a way similar to "hierarchical management", and then perform more efficient fusion on the summarized information, which can also reduce the algorithm complexity and enable parallel computing to accelerate training under the same parameter scale.

[0062] (2) The present invention divides 3D vision into two semantic modalities, namely the color modality and the spatial modality. By introducing the normal vector angle image as prior information of the spatial semantic modality into the network input to explicitly express the hidden semantics of the object surface directionality, the network with poor interpretability can directly obtain depth information, thus simplifying the network learning task. In addition, the normal vector angle image can significantly reflect the direction change of the object edge on the corresponding coordinate axis. Combining with the depth map, it can quickly point out the spatial positions of different objects, and can solve the occlusion problem in the scene to a certain extent, enhancing the accuracy of model prediction.

[0063] (3) The present invention changes the encoder in the Swin-Unet backbone network to a two-stream structure, which solves the problem of insufficient fusion of RGB color and spatial information caused by introducing the normal vector angle image. If the RGB color image and the DI spatial image are directly concatenated into a seven-channel initial vector and input into the original Swin-Unet network, due to the large differences in data distribution and semantic information between different modalities, the network cannot well distinguish the information of each channel, resulting in a regression of the model effect or even non-convergence. By introducing the two-stream structure, the two sub-encoders first separately perform semantic understanding within the modality on the RGB color information and the DI spatial information, which can avoid interference between the two due to premature fusion. Brief Description of the Drawings

[0064] Figure 1 is the Swin-Unet network architecture diagram of the present invention;

[0065] Figure 2 is the depth map of a sample of the present invention and the grayscale map of the normal vector angle;

[0066] Figure 3 is the schematic diagram of the SU-Grasp two-stream multi-modal fusion strategy of the present invention. Detailed Embodiment

[0067] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0068] Embodiment 1

[0069] This embodiment provides a manipulator control method and system based on an improved Swin-Unet, aiming to solve the problems of insufficient utilization of depth information and weak local and global perception ability during three-dimensional visual grasping of the manipulator.

[0070] The architecture of the improved Swin-Unet model adopted by this method is as follows:

[0071] As Figure 1 shown, a grasping pose prediction model framework SU-Grasp based on a three-dimensional visual two-stream encoding strategy is constructed based on a four-layer Swin-Unet network. Other-layer Swin-Unet networks can also be used according to specific situations. The entire network architecture includes 5 modules, namely a point cloud preprocessing module, a two-stream encoder, a bottleneck layer, a decoder, and a skip connection layer.

[0072] Swin-Unet network structure: The Swin-Unet model is a U-shaped architecture based on the Swin Transformer model. It consists of an encoder, a bottleneck layer, a decoder, and skip connections. Among them, the encoder, the bottleneck layer, and the decoder are all constructed based on the Swin Transformer Block model. This network structure divides the input image into non-overlapping image patches, which are input into the encoder to learn deep feature representations. After extracting the context features, they are upsampled by the decoder with an image patch expansion layer and connected to the same-level encoder through skip connections to gradually restore the spatial resolution of the feature map for pixel-level segmentation prediction. Using the Swin-Unet model as the basic network framework of the SU-Grasp model can obtain both local information (such as object contours and long-range connections) and long-range connections (such as the relationships between different visual entities in a cluttered environment).

[0073] In the implementation of the SU-Grasp model, to achieve end-to-end parallel two-finger gripper grasping prediction, three grasping outputs are connected in parallel to the top of the decoder, representing the grasping confidence Q, the grasping width W, and the grasping angle Θ respectively. Each output is a heatmap of the same size as the input RGBD image. Using this pixel-level grasping representation method can obtain the best grasping pose in the global visual scene in one forward propagation, thus avoiding the need to generate multiple grasping candidate objects and saving computational costs.

[0074] This robotic arm control method based on the improved Swin-Unet mainly focuses on two issues: multi-scale feature extraction and the efficient fusion of spatial information and color information. For the former, the Swin-Unet model of the object detection algorithm is adapted for the robotic arm grasping task, and the multi-scale is improved through the multi-level Swin Transformer Block model and the U-shaped structure with skip connections; for the latter, on the one hand, a normal vector angle image is constructed to provide prior information in the scene, and on the other hand, a two-stream structure is used to achieve the separate extraction and cross-fusion of color information and spatial information.

[0075] This robotic arm control method based on the improved Swin-Unet includes the following steps:

[0076] Step S0: Obtain the RGBD image of the target working scene, that is, the data containing the RGB color image and the depth image, and use it as the model input. Specifically, the dataset preparation is as follows:

[0077] In the robotic arm planar grasping task, the commonly used datasets include Cornell and Jacquard. The Cornell dataset contains 885 pictures of 244 common daily items. Each sample shows different poses of different items and is accompanied by grasping box annotations. There are a total of 8019 grasping boxes marked in this dataset, among which 5110 are feasible grasping boxes and 2909 are infeasible grasping boxes. The Jacquard dataset includes: RGB scene images rendered using Blender, depth images obtained through existing stereo vision algorithms, text files with grasping annotations, and binary background mask images that can be used for data augmentation.

[0078] Furthermore, step S0 also includes preprocessing the depth image. This preprocessing includes plane segmentation operations based on the Random Sample Consensus (RANSAC) algorithm and outlier removal operations based on radius-nearest neighbor search.

[0079] Step S1: Convert the depth image into a normal vector angle image, and perform a splicing operation on the depth image and the normal vector angle image to obtain a spatial image.

[0080] During the process of step S1, a normal vector angle image is introduced. The 3D vision is divided into two semantic modalities: color modality and spatial modality. An attempt is made to enhance the information interaction between the two through the method of multimodal fusion. Combining the process of human grasping an object, the spatial modality plays two key roles in the grasping process: First, it helps to determine the overall position and orientation of the object to be grasped; Second, it provides local geometric information of the object surface to assist in making decisions on the most stable grasping points. In order to avoid letting the network learn the hidden semantics in the depth information by itself, in this embodiment, it is directly expressed explicitly through the normal vector angle image, thereby simplifying the learning task. Therefore, the normal vector angle image is added to the network input as prior information of the spatial semantic modality.

[0081] Furthermore, call the Open3D library to convert the preprocessed depth map into a normal vector angle image. To calculate the normal vector angle image, for each pixel (u, v) in the depth image D, calculate the unit normal vector Ν(n x ,n y ,n z ) corresponding to the object surface at this pixel. Then, determine the angles between N and the three axes XYZ of the camera coordinate system. Let the three axes XYZ be represented as vectors x(1, 0, 0), y(0, 1, 0), z(0, 0, 1) respectively. Then the angles between Ν(n x ,n y ,n z ) and these three axis vectors can be expressed as:

[0082] a x =arccos(N·x)

[0083] a y =arccos(N·y)

[0084] a z =arccos(N·z)

[0085] In the formula, N is the unit normal vector of the object surface corresponding to the current pixel; x, y, z represent the unit normal vectors in the camera coordinate system that are in the same direction as the positive directions of the three coordinate axes; a x 、a y 、a z correspond to the angles between N and the x, y, z axis directions; the arccos function is used to calculate the cosine value of the angle between two vectors.

[0086] Finally, normalize the angles (a x ,a y ,a z ) to the range of 0 to 255 for creating a three-channel normal vector angle image I nrm , then each pixel has the normalized (a x ,a y ,az ) as its value.

[0087] Figure 2 is the visualization result of the grayscale image of the depth map and the normal vector angle of a certain sample, where Figure a is the original scene image, and Figure b is a x single-channel grayscale image, and Figure c is a z single-channel grayscale image, and Figure d is a y single-channel grayscale image. It can be found that the normal vector angle image can significantly reflect the direction change of the object edge on the corresponding coordinate axis. Combined with the depth map, it can quickly indicate the spatial positions of different objects, and to a certain extent, it can solve the occlusion problem in the scene. The ablation experiment proves that the normal vector angle image can effectively accelerate the model convergence process and help the model better understand the spatial features.

[0088] In order to more efficiently fuse RGB color information and depth information, the encoder part of the Swin-Unet model in this embodiment has made innovative adjustments.

[0089] First, the normal vector angle image is introduced into the robotic arm grasping task to explicitly express the hidden semantics of the depth information on the object surface. Second, the depth image and the RGB image are regarded as two modalities. The former describes the depth difference and structural information between different objects, while the latter reflects the appearance information such as the color, shape, and boundary of a scene. Therefore, this embodiment adopts the two-stream structure widely used in the field of salient object detection to explore the respective advantages of different modalities and their complementary properties.

[0090] The two-stream structure refers to a network architecture that uses two encoders to process the inputs of different modalities respectively, and then performs information fusion through cross-modal methods. Considering the semantic similarity of the depth image D and the normal vector angle image I nrm in representing spatial information, they are regarded as the same modality and concatenated into a four-channel spatial image vector (denoted as the DI image). After the color information and the spatial information are encoded respectively, the extracted information is integrated through cross-modal fusion methods to enhance their semantic representation of the object position and shape, and finally the subsequent decoding operation is completed.

[0091] Step S2: Analyze the RGB color image and the spatial image by using the grasping model to obtain the grasping parameters. The RGB two-dimensional image is regarded as the color semantic modality, and the depth map D and the normal vector angle image I are regarded as the spatial semantic modalities, which are used as the inputs of the SU-Grasp network.

[0092] The grasping model is constructed based on the improved Swin-Unet and includes:

[0093] The encoding part extracts multiple set-scale features of the RGB color image and the spatial image respectively to obtain color features and spatial features. Then, the color features and spatial features of the same scale are fused to obtain several first fusion features. Next, the first fusion features of the smallest set scale are extracted to obtain second fusion features. The decoding part predicts the grasping parameters based on the first fusion features and the second fusion features.

[0094] Furthermore, the process of predicting the grasping parameters includes: upsampling the second fusion features of the smallest scale step by step to the input resolution. In each upsampling operation, the first fusion features or the second fusion features of the corresponding scale are fused through skip connections, and then the grasping parameters are predicted.

[0095] Considering that SU-Grasp is a hierarchical U-shaped network and after comprehensive consideration of aspects such as algorithm complexity and code implementation difficulty, the multi-level structure of the Swin Transformer decoder is used to innovatively transform the current encoder, such as Figure 3 shown, the specific logic of the algorithm is as follows:

[0096] (1) Set the hyperparameter C (i.e., the length of each image patch embedding vector) in the "input→stage2" part of the original Swin-Unet encoder to 1 / 2 of the original, and copy this structure to obtain two sub-encoders, as Figure 3 shown by the red box in.

[0097] (2) Input the RGB color image and the DI spatial image into the two sub-encoders respectively. After one convolution operation, the data sizes in the two sub-encoders both become

[0098] (3) Each of the two sub-encoders completes steps 1 and 2, and upsampling is achieved through two Swin Transformer Blocks. At this time, the feature dimensions output by both are 2C. Before performing the stage3 calculation, the data in the two sub-encoders are concatenated to restore the embedding dimension of the current feature map to 4C.

[0099] (4) Continue to complete the calculations of stage3 and the bottleneck layer in the encoder.

[0100] (5) In the decoder, the skip connection operation keeps the first layer consistent with the original Swin-Unet structure, and at the same time, in the second and third layers, the outputs of the two sibling sub-encoders and the feature map of the previous-level decoder need to be concatenated, as shown in the following formula:

[0101]

[0102] Wherein, respectively represent the input of the nth - stage decoder and the output of the (n - 1)th - stage decoder; and respectively represent the outputs of the two - stream sub - encoders corresponding to the nth - stage decoder.

[0103] Figure 3 The shown scheme selects stage3 as the information fusion position of the two - stream structure. Of course, the cross - modal timing can also be advanced or delayed according to needs, which benefits from the hierarchical structure of the Swin - Unet model. Further, color features and spatial features are obtained by respectively extracting features of multiple set scales from the RGB color image and the spatial image, which is achieved by constructing two sub - encoders based on the Swin - Unet structure. In addition, the lengths of the patch embedding vectors of the two sub - encoders are kept the same.

[0104] Further, the grasping parameters include the grasping position, and the grasping confidence, grasping width, and grasping angle at the grasping position. The output of the SU - Grasp network is the heatmaps (Q, W, Θ) of the grasping center - point confidence, grasping width, and grasping angle. When determining the actual grasping position, the position with the highest grasping confidence can be obtained by retrieving the grasping quality heatmap Q, defined as:

[0105]

[0106] Wherein, represents the grasping position with the highest grasping confidence, that is, the best grasping position recommended by the model; argmax pos represents finding the independent variable that makes the following expression take the maximum value; Q is the heatmap representing the grasping confidence.

[0107] Then, the predicted width w and angle θ at the corresponding position are extracted from the predicted heatmaps of the grasping width W and grasping angle Θ.

[0108] In this embodiment, the grasping pose estimation is regarded as a regression problem, that is, a mapping relationship F∶ is established through the SU - Grasp model where I represents the input information, specifically referring to the RGBDI nrm seven - channel vector, and represents the predicted heatmaps (Q, W, Θ) output by the model.

[0109] Based on the idea of minimizing the prediction distance, the loss function is defined as:

[0110]

[0111] Wherein, represents the total loss value; N represents the number of samples; m represents three different output metrics: grasping confidence, grasping width, and grasping rotation angle; w m represents the weight of each predicted heatmap loss value, and the default value is 1 for all; is the predicted value of the model, representing the prediction result of the i-th sample under the metric m; is the true label data, representing the actual result of the i-th sample under the metric m; ||····|| represents the norm of the vector, usually the L2 norm, indicating the gap between the predicted value and the true value.

[0112] Step S3: Control the robotic arm to grasp the target object based on the grasping parameters.

[0113] Further, Step S3 includes:

[0114] Retrieve the grasping position with the highest grasping confidence to obtain the target grasping position; extract the grasping width and grasping angle of the target grasping position to obtain the target grasping width and target grasping angle; control the robotic arm to grasp the target object according to the target grasping position, target grasping width, and target grasping angle.

[0115] Aiming at the problem of insufficient spatial understanding ability in the existing robotic arm visual grasping algorithm, this embodiment proposes a robotic arm control method based on the improved Swin-Unet. This method introduces a dual-stream encoding mechanism based on sliding window attention and spatial semantic enhancement into the U-shaped network model. The accuracy rates on the Cornell and Jacquard datasets reach 98.1% and 95.2% respectively, which are superior to the existing algorithms in terms of training efficiency and training time, indicating that the improved algorithm has stronger learning ability in spatial understanding. Through simulation tests, it is found that the model can still accurately judge the position information of the current significant object and complete the grasping in the face of a multi-object cluttered scene, and the accuracy rate reaches 83%.

[0116] Embodiment 2

[0117] Based on the robotic arm control method based on the improved Swin-Unet provided in the embodiment, this embodiment proposes a robotic arm control system based on the improved Swin-Unet, including:

[0118] An input module that acquires the RGB image and depth image of the target working scene;

[0119] A processing module that converts the depth image into a normal vector angle image and performs a splicing operation on the depth image and the normal vector angle image to obtain a spatial image;

[0120] A grasping model that analyzes the RGB image and the spatial image using the grasping model to obtain grasping parameters;

[0121] A control module that controls the robotic arm to grasp the target object based on grasping parameters;

[0122] Among them, the grasping model is constructed based on the improved Swin-Unet and includes:

[0123] An encoding part that respectively extracts the features of multiple set scales of the RGB color image and the spatial image to obtain image color features and spatial features, and performs feature fusion on the image color features and spatial features of the same scale to obtain a number of first fusion features, and then performs feature extraction on the first fusion features of the smallest set scale for at least one scale to obtain second fusion features;

[0124] A decoding part that predicts the grasping parameters according to the first fusion features and the second fusion features

[0125] In summary, SU-Grasp uses independent encoders to respectively extract multi-scale features of the RGB color image and the DI spatial image, and realizes cross-modal information fusion through serial splicing. Compared with directly inputting seven-channel data, the dual-stream encoder structure can first separately extract the shallow semantic features of each modality in a way similar to "hierarchical management", and then perform more efficient fusion on the summarized information. It can also reduce the algorithm complexity and achieve parallel computing to accelerate training under the same parameter level.

[0126] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. An arm control method based on an improved Swin-Unet, characterized in that: The Swin-Unet network based on a four-layer structure is used to construct a grasping pose prediction model SU-Grasp based on a three-dimensional vision dual-stream encoding strategy. The improvements of Swin-Unet include: setting the hyperparameter C of the part from input to stage2 in the original Swin-Unet encoder to 1 / 2 of the original, and replicating this structure to obtain two sub-encoders. Each of the two sub-encoders completes stage1 and 2, and upsampling is achieved through two Swin Transformer Blocks. The feature dimensions output by the two sub-encoders are 2C. Before performing the stage3 calculation, the data in the two sub-encoders are concatenated to restore the embedding dimension of the current feature map to 4C. In the decoder, the skip connection operation maintains the first layer consistent with the original Swin-Unet structure, and in the second and third layers, the outputs of the two sibling sub-encoders and the feature map of the previous-level decoder are concatenated. The manipulator control method based on the improved Swin-Unet includes the following steps: Step S0: Obtain the RGBD image of the target working scenario, where the RGBD image contains the data of the RGB color image and the depth image. Step S1: Convert the depth image into a normal vector angle image, and perform a concatenation operation on the depth image and the normal vector angle image to obtain a spatial image. Step S2: Consider the RGB two-dimensional image as the color semantic modality, and consider the depth map D and the normal vector angle image I nrm as the spatial semantic modality. Use the color semantic modality image and the spatial semantic modality image as the inputs of two sub-encoders in the SU-Grasp model respectively to obtain the grasping parameters; Step S3: Control the manipulator to grasp the target object based on the grasping parameters.

2. The manipulator control method based on the improved Swin-Unet according to claim 1, characterized in that: In the architecture of the Swin-Unet model, the encoder, bottleneck layer, and decoder are constructed based on the Swin Transformer Block model. The Swin-Unet network structure divides the input image into non-overlapping image patches, which are input into the encoder to learn deep feature representations. After extracting the context features, upsampling is performed by the decoder with an image patch expansion layer, and it is connected to the sibling encoder through skip connections to gradually restore the spatial resolution of the feature map for pixel-level segmentation prediction. Using the Swin-Unet model as the basic network framework of the SU-Grasp model to obtain both local information and long-range connections. During the implementation of the SU-Grasp model, three grasping outputs are connected in parallel to the top of the decoder, representing the grasping confidence heatmap Q, the grasping width heatmap W, and the grasping angle heatmap Θ respectively. Each output is a heatmap with the same size as the input RGBD image.

3. The robotic arm control method based on the improved Swin-Unet according to claim 1, characterized in that: In step S0, the data includes the Cornell dataset and the Jacquard dataset. Step S0 also includes preprocessing the depth image, and the preprocessing includes plane segmentation operation based on the Random Sample Consensus (RANSAC) algorithm and outlier removal operation based on radius-nearest neighbor search.

4. The robotic arm control method based on the improved Swin-Unet according to claim 1, characterized in that: In step S1, the normal vector angle image is introduced, and the three-dimensional vision is divided into two semantic modalities: color modality and spatial modality. The information interaction between the two is enhanced through the method of multi-modal fusion. During the grasping process in the spatial modality: First, determine the overall position and pose of the object to be grasped. Second, provide local geometric information on the object surface to assist in deciding the most stable grasping point.

5. The robotic arm control method based on the improved Swin-Unet according to claim 1, characterized in that: In step S1, the Open3D library is called to convert the preprocessed depth map into a normal vector angle image. The calculation process of the normal vector angle image is as follows: Calculate the unit normal vector Ν(n x ,n y ,n z ) corresponding to the object surface for each pixel (u, v) in the depth image D; Determine the angles between N and the three axes X, Y, and Z of the camera coordinate system; Let the X, Y, and Z axes be represented as vectors x(1, 0, 0), y(0, 1, 0), and z(0, 0, 1) respectively. The angles between Ν(n x , n y , n z ) and these three axis vectors can be expressed as: a x = arccos(N·x) a y = arccos(N·y) a z = arccos(N·z) Where, N is the unit normal vector of the object surface corresponding to the current pixel; x, y, z represent the unit normal vectors in the camera coordinate system that are in the same direction as the positive directions of the three coordinate axes; a x , a y , a z correspond to the angles of N with the x, y, z axes; the arccos function is used to calculate the cosine value of the angle between two vectors; Normalize the angles (a x , a y , a z ) to the range of 0 to 255 for creating a three-channel normal vector angle image I nrm . Then each pixel has the normalized (a x , a y , a z ) as its value.

6. The robotic arm control method based on the improved Swin-Unet according to claim 1, wherein: In step S1, the normal vector angle image is introduced into the robotic arm grasping task to explicitly represent the hidden semantics of the depth information in the directionality on the object surface. The depth image and the RGB image are regarded as two modalities. The former describes the depth difference and structural information between different objects, while the latter reflects the appearance information including the color, shape, and boundaries of a scene. A two-stream structure is adopted to exploit the respective advantages of different modalities and their complementary properties; The two-stream structure uses two encoders to process the inputs of different modalities respectively, and then a network architecture for cross-modal information fusion; Regarding the depth image D and the normal vector angle image I nrm as the same modality, and concatenating them into a four-channel spatial image vector, denoted as DI nrm image; the color information and the spatial information are first encoded separately, and then the extracted information is integrated through cross-modal fusion to enhance the semantic representation of the object's position and shape by both, and finally the subsequent decoding operation is completed.

7. The manipulator control method based on the improved Swin-Unet according to claim 6, wherein: The grasping model in step S2 is based on the improved Swin-Unet and includes: the encoding part extracts multiple set-scale features of the RGB color image and the spatial image respectively to obtain color features and spatial features; then, the color features and spatial features of the same scale are fused to obtain several first fusion features; then, the first fusion features of the smallest set scale are extracted to obtain second fusion features; the decoding part predicts the grasping parameters according to the first fusion features and the second fusion features.

8. The manipulator control method based on the improved Swin-Unet according to claim 7, wherein: The process of the grasping parameters in step S2 includes: upsampling the second fusion features of the smallest scale step by step to the input resolution; in each upsampling operation, the first fusion features or the second fusion features of the corresponding scale are fused through skip connections, and then the grasping parameters are predicted; The grasping parameters include the grasping position, and the grasping confidence, grasping width, and grasping angle of the grasping position; the output of the SU-Grasp network is the heat maps (Q, W, Θ) of the grasping center point confidence, grasping width, and grasping angle; when determining the actual grasping position, the position with the highest grasping confidence is obtained by retrieving the grasping confidence heat map Q, defined as: wherein, represents the grasping position with the highest grasping confidence, that is, the best grasping position recommended by the model; argmax pos represents finding the independent variable that maximizes the following expression; Q represents the grasping confidence heat map; Then, the predicted width w and angle θ at the corresponding position are extracted from the predicted heat maps of the grasping width W and the grasping angle Θ; Regarding the grasping pose estimation as a regression problem, that is, establishing a mapping relationship F∶ through the SU-Grasp model where I represents the input information, specifically referring to RGBDI nrm a seven-channel vector, and represents the predicted heat map (Q, W, Θ) output by the model; Based on minimizing the predicted distance, the loss function is defined as: Wherein, represents the total loss value; N represents the number of samples; m represents three different output metrics of grasping confidence, grasping width, and grasping rotation angle; w m represents the weight of each predicted heatmap loss value, and the default value is 1 for all; is the predicted value of the model, representing the prediction result of the i-th sample under the metric m; is the true label data, representing the actual result of the i-th sample under the metric m; ||....|| represents the norm of the vector, usually the L2 norm, indicating the gap between the predicted value and the true value.

9. The robotic arm control method based on the improved Swin-Unet according to claim 7, characterized in that: Step S3 includes: retrieving the grasping position with the highest grasping confidence to obtain the target grasping position; extracting the grasping width and grasping angle of the target grasping position to obtain the target grasping width and target grasping angle; controlling the robotic arm to grasp the target object according to the target grasping position, target grasping width, and target grasping angle.

10. A robotic arm control system based on an improved Swin-Unet, based on the robotic arm control method based on the improved Swin-Unet according to any one of claims 1-9, characterized in that, Including: An input module that acquires the RGB image and the depth image of the target working scene; A processing module that converts the depth image into a normal vector angle image and performs a splicing operation on the depth image and the normal vector angle image to obtain a spatial image; A grasping model that analyzes the RGB image and the spatial image using the grasping model to obtain the grasping parameters; A control module that controls the robotic arm to grasp the target object based on the grasping parameters; Among them, the grasping model is based on the improved Swin-Unet and includes: The encoding part respectively extracts features of multiple set scales of an RGB color image and a spatial image to obtain image color features and spatial features, and fuses the image color features and spatial features of the same scale to obtain a number of first fusion features, and then performs feature extraction on the first fusion features of the smallest set scale at at least one scale to obtain second fusion features; the decoding part predicts grasping parameters according to the first fusion features and the second fusion features.