Mechanical arm visual grabbing method based on Sim2Real migration under limited space condition
By performing visual domain randomization training in the simulation environment, the problem of crawling robotic arms in space-constrained and crowded and messy scenarios is solved, and efficient and accurate target capture and improved model robustness is achieved.
Patent Information
- Application Number
- CN202510364949.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-26
AI Technical Summary
In the case of space constrained and crowded and messy, it is difficult for robotic arm visual grasping technology to achieve efficient and accurate target grasping, and the existing methods are not effective in this environment.
Using a method based on Sim2Real migration, a variety of different simulation environments are generated by randomized visual domain in the simulation environment, and the grabbing model of the robotic arm is trained so that it can grasp the target object in a complex environment.
The goal grabbing success rate and action efficiency of the robotic arm in space-constrained scenarios are improved, and the robustness of the model and adaptability to complex environments are enhanced.
Smart Images

Figure CN119974010A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robot arm visual grasping and Sim2Real technology, and in particular to a robot arm visual grasping method based on Sim2Real migration under space-constrained conditions. Background Art
[0002] With the continuous development of robotics technology, robot visual grasping technology, as one of the core functions of robot operation, has become an important research direction in the fields of industrial automation, intelligent logistics, service robots, etc. However, in practical applications, robot visual grasping technology still faces many challenges, especially in crowded and cluttered scenes, how to grasp the target object accurately and efficiently has become an urgent problem to be solved.
[0003] There are several challenges in the problem of grasping targets in crowded and cluttered scenes in the real world: first, the target object will be blocked and surrounded by other objects, and it is impossible to grasp the target object directly; second, the target objects have different colors and shapes, which require accurate identification and stable grasping; third, there is a reality gap between the simulation environment and the real world, which may cause the grasping performance in the real world to be lower than that in the simulation environment. At present, many studies use the "push and grasp" strategy for grasping in a single simulation environment, aiming to break up all objects through the "push" action in order to create enough grasping space for the target object. However, these methods are suitable for target grasping tasks in scenes with unrestricted workspace. When the workspace is restricted, the "push" action is invalid in scenes with restricted space because the target object is often tightly surrounded or blocked by surrounding debris. At the same time, since a single environment cannot meet the environmental changes in the real world, the grasping model trained in the simulation environment is not robust enough. Therefore, exploring new grasping strategies to achieve efficient and accurate grasping of targets in the real world with restricted space and crowded and cluttered conditions is a technical problem that needs to be solved urgently in this field.
[0004] Sim2Real transfer is an important challenge in the fields of robotics and computer vision, which involves transferring control policies or models learned in simulation environments to real-world applications.
[0005] Sim2Real migration, that is, the migration from simulation to reality, refers to the process in which the model, algorithm or control strategy trained or optimized in the simulation environment can be effectively run in the real world directly or with a small amount of adjustment. Due to the complexity and uncertainty of the real-world environment, direct training or optimization in the real environment is often costly and inefficient, so Sim2Real migration has become an important research direction.
[0006] Sim2Real migration has broad application prospects in robotics, autonomous driving, augmented reality and other fields. For example, in robotics, Sim2Real migration can be used to train robots that can perform complex tasks in the real world; in the field of autonomous driving, the simulation environment can be used to test and optimize the safety and reliability of autonomous driving algorithms; in the field of augmented reality, Sim2Real migration can be used to seamlessly integrate virtual objects with the real world.
[0007] Visual domain randomization is a key technology in Sim2Real migration, which aims to improve the generalization and robustness of the model by introducing diverse visual features in the simulation environment. The core idea of this method is to randomize the visual parameters in the simulation environment so that the model can handle unpredictable changes in the real world during training, such as lighting changes, camera perspective, object shape and position, image noise, etc. Summary of the invention
[0008] The technical problem to be solved by the present invention is to provide a robot arm visual grasping method based on Sim2Real migration under space-constrained conditions in view of the deficiencies of the above-mentioned prior art, and to realize the visual grasping of the robot arm based on Sim2Real migration.
[0009] In order to solve the above technical problems, the technical solution adopted by the present invention is: a robot arm visual grasping method based on Sim2Real migration under space-constrained conditions, comprising:
[0010] Build a task scene, select the visual domain randomization parameters of the simulation environment, and conduct a grasping simulation experiment in the task scene to determine the randomization parameters; the randomization parameters include the camera pose; the brightness, saturation and contrast of the RGB image captured by the camera; and the Gaussian noise of the image;
[0011] Set the randomization range for each randomization parameter to form a variety of different simulation environments;
[0012] The first-stage grasping model is trained in the generated random simulation environment to enable the robot arm to grasp all objects; the first-stage grasping model includes a first perception module and a grasping module; the first perception module extracts image features; the grasping module predicts the Q value of each pixel in the image;
[0013] The second-stage grasping model is trained in a cluttered and crowded simulation environment, so that the robot arm grasps the target objects in sequence from the cluttered and crowded simulation environment; the second-stage grasping model includes a second perception module and a grasping module; the second perception module extracts image features; the grasping module predicts the Q value of each pixel in the image;
[0014] The grasping model trained in the second stage is used to conduct target grasping tests in spatially constrained scenarios.
[0015] Furthermore, the task scene is constructed, the visual domain randomization parameters of the simulation environment are selected, and a grasping simulation experiment is performed in the task scene to determine the randomization parameters, including:
[0016] Multiple objects of different colors and shapes are randomly placed in a workspace, and the grasping task is to grasp all the objects;
[0017] According to the possible camera poses in the real environment, the 3D position and rotation angle of the camera are randomized to simulate different perspective changes and train the network model deployed on the robot arm to guide the robot arm to achieve the grasping task;
[0018] According to different lighting conditions in real environments, the visual features of the RGB images captured by the camera are adjusted to perform image enhancement and then the network model is trained;
[0019] According to the influence of sensor noise in the real environment, Gaussian noise is added to the RGB images taken by the camera during network model training to train the network model.
[0020] Furthermore, the randomization range is set for each randomization parameter to form a variety of different simulation environments, including:
[0021] Set the camera position distribution range and camera angle distribution range;
[0022] Set the range of visual feature changes of the RGB image captured by the camera;
[0023] Set the Gaussian noise added during image enhancement;
[0024] Each round of training is randomly sampled from the randomized parameter distribution range set above, so as to generate different parameter configurations to form simulation environments with different conditions, that is, each round of training of the network model is trained in a simulation environment with different conditions.
[0025] Furthermore, the camera position distribution range is set to: (±0.05, ±0.025, ±0.05) m, that is, the camera randomly selects the three-dimensional position coordinates in a cuboid with a length of 0.1 m, a width of 0.05 m, and a height of 0.1 m; the camera angle distribution range is set to: (0, ±0.1, ±0.1) rad, that is, the randomization range of the pitch angle and the yaw angle is ±0.1 rad;
[0026] The brightness, saturation, and contrast of the RGB image captured by the camera are set to a range of [0.7, 1.3], that is, ranging from 70% to 130% of the brightness, contrast, and saturation of the original image;
[0027] The Gaussian noise added to the RGB image captured by the camera is set to have a mean μ=0 and a variance σ=25.
[0028] Furthermore, the first-stage grasping model training is performed in the generated random simulation environment so that the robot arm grasps all objects, including:
[0029] Acquire RGB-D images of the scene from depth cameras randomly placed above the workspace;
[0030] Randomly change the brightness, saturation, and contrast of RGB-D images;
[0031] Add random Gaussian noise to the altered RGB-D image;
[0032] The RGB-D image with random Gaussian noise added is sent to the first perception module to extract image features, and then the grasping module is used to predict the Q value of each pixel, where the Q value represents the expected future reward obtained by grasping at the pixel. The first perception module uses a Densenet network to extract image features. The grasping module consists of convolution, normalization and activation layers, and is connected to the perception module through upsampling and skip connection operations to predict the Q value of the pixel.
[0033] Perform the grabbing action at the pixel point corresponding to the maximum Q value, and obtain the corresponding reward if the grabbing is successful;
[0034] If the set number of grasping actions has been performed, the first-stage grasping model training is stopped, and the trained first-stage grasping model is used as the initialization model for the robotic arm grasping task; otherwise, the RGB-D image of the scene is obtained again from the depth camera randomly placed above the workspace.
[0035] Furthermore, the training of the second-stage grasping model in a cluttered and crowded simulation environment so that the robot arm can grasp the target objects in sequence from the cluttered and crowded simulation environment includes:
[0036] S1: Randomly place M objects of different colors and shapes in a workspace, and the grasping task is to grasp all objects of one color; randomly select the color of the target object to be grasped and set the number of target objects to be grasped, and initialize the number of grasped target objects to 0;
[0037] S2: Obtain an RGB-D image of a random scene, change its brightness, contrast, saturation, and add random Gaussian noise; send the RGB-D image to the second perception module to extract image features, and then use the capture module to predict the Q value of each pixel; the second perception module uses the SE-Densenet network to extract image features;
[0038] S3: Obtain a binary mask image containing all target objects, and then use the connected component function to obtain a binary mask image containing only a single target object and its center position;
[0039] S4: Calculate the edge binary mask map and edge height map corresponding to the binary mask map of the target object, and then calculate the boundary occupancy rate of the target object;
[0040] S5: Select the target object with the smallest boundary occupancy rate as the current target to be grasped;
[0041] S6: If the boundary occupancy rate of the current target to be grasped is lower than the set threshold γ, the target is grasped; otherwise, the first search is performed with the target as the center to find obstacles that can be grasped within the set range around it;
[0042] S7: If there is an obstacle with a boundary occupancy rate lower than the set threshold value γ within the set range, the obstacle is captured; otherwise, the obstacle corresponding to the minimum boundary occupancy rate is taken as the center, and a second search is performed to capture the obstacle with the minimum boundary occupancy rate within the set range around it;
[0043] S8: If the capture is successful, you will receive corresponding rewards;
[0044] S9: adding 1 to the number of captured target objects;
[0045] S10: If the set number of grasping actions has been performed, stop training and save the second-stage grasping model; otherwise, determine whether there is still a target object in the workspace: if the target object is detected, loop through steps S2 to S10;
[0046] If the target object is not detected and the number of captured target objects is less than the total number of target objects to be captured, it means that the target object is completely blocked by other obstacles and is not detected. Then, all objects in the workspace whose height exceeds the preset threshold are regarded as potential objects to be captured, and the object corresponding to the maximum Q value is captured until the target object enters the detection range, and steps S2 to S10 are cycled.
[0047] If no target object is detected and the number of captured target objects is equal to the number of target objects to be captured, step S1 is re-executed to start a new round of training.
[0048] Furthermore, the SE-Densenet network connects the DenseNet-121 network and the attention mechanism SENet network: first, the DenseNet-121 network extracts the global features of the image, and then uses the global average pooling operation to compress the global characteristics output by the DenseNet-121 network; then, the number of channels of the feature map is reduced to one-fourth of the original through the first fully connected layer and activated using Relu, and then the number of channels is restored through the second fully connected layer to model the correlation between channels; then, the importance weights of each channel are normalized by the Sigmoid function to ensure that the weight value is between 0 and 1; finally, the normalized weights are applied to each channel of the feature map, giving each channel different importance, so that the SE-Densenet network can focus on key features when predicting.
[0049] On the other hand, the present invention also provides a robot arm visual grasping system based on Sim2Real migration under space-constrained conditions, including: a simulation environment construction module: setting a randomization range for each randomization parameter to form a variety of different simulation environments;
[0050] The first training module: trains the first-stage grasping model in the generated random simulation environment to enable the robot arm to grasp all objects;
[0051] Second training module: training the second-stage grasping model in a cluttered and crowded simulation environment, so that the robot arm can grasp the target objects in sequence from the cluttered and crowded simulation environment;
[0052] Test module: Use the grasping model trained in the second stage to conduct target grasping tests in space-constrained scenarios.
[0053] In a third aspect, the present application proposes an electronic device comprising: one or more processors, and a memory, wherein the memory is used to store instructions, and when the instructions are executed by the one or more processors, the one or more processors execute the robotic arm visual grasping method based on Sim2Real migration under spatially constrained conditions.
[0054] In a fourth aspect, the present application proposes a computer-readable storage medium storing executable instructions, which, when executed, enable a processor to execute the robotic arm visual grasping method based on Sim2Real migration under the spatially constrained conditions.
[0055] The beneficial effects of adopting the above technical solution are: the present invention provides a robot arm visual grasping method based on Sim2Real migration under space-constrained conditions. (1) The present invention adopts a phased training strategy. The first phase focuses on enabling the robot arm to have a certain grasping success rate for objects of different shapes and sizes, and initially master the basic grasping skills; the second phase focuses on coordinating the grasping sequence of targets and obstacles to improve the action efficiency and grasping success rate for the target, and ensure that the robot arm can accurately grasp the target object in a complex environment. This phased training strategy reduces the training difficulty of the task and improves the grasping performance of the grasping model.
[0056] (2) The target grasping framework based on obstacle removal provided by the present invention selects obstacles around the target to be grasped when the target object cannot be grasped directly, thereby reducing the complexity of the workspace and providing sufficient grasping space for the target. This not only improves the target grasping success rate, but also improves the action efficiency.
[0057] (3) In the second stage of training, the present invention proposes a new perception network module SE-DenseNet, embedding the SENet network into the DenseNet module, thereby enhancing the network's nonlinear modeling ability and feature representation ability. The SENet network adaptively learns channel attention weights to strengthen the network's response to channels related to object color, shape, position information and object edge contour information, thereby improving the sensitivity and perception ability to key feature information and further improving the grasping success rate.
[0058] (4) The present invention takes into account the actual gap between the simulation environment and the real world. When training the grasping model in the simulation environment, a visual domain randomization method based on expert prior knowledge is used to generate a variety of different random environments, so that the grasping model is exposed to a variety of different environmental conditions, further improving the robustness of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 A framework diagram of a robotic arm visual grasping method based on Sim2Real migration under space-constrained conditions provided by an embodiment of the present invention;
[0060] Figure 2 A diagram of the SE-DenseNet network structure for extracting image features provided by an embodiment of the present invention;
[0061] Figure 3 A schematic diagram of an obstacle search provided by an embodiment of the present invention, wherein (a) is a case where the target object boundary occupancy rate is lower than a threshold value, and (b) is a case where the target object boundary occupancy rate is higher than or equal to the threshold value;
[0062] Figure 4Schematic diagram of test scenarios with different space sizes provided in an embodiment of the present invention, where (a) is a normal working space and (b) is a restricted working space;
[0063] Figure 5 A comparison chart of the capture success rates before and after the introduction of the SENet network provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0064] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0065] The method of the present invention is suitable for the visual grasping training of the robot arm, and aims to improve the target grasping performance of the robot arm in the space-constrained scene in the real world. Figure 1 As shown in the figure, the RGB-D height map of the workspace is first rotated 16 times in steps of 22.5° and fed into the network. The network outputs 16 Q-value pixel maps with different rotation angles. The Q-value of each pixel represents the expected future reward obtained by grasping at that pixel. The rotation angle of the height map corresponds to the rotation angle of the end effector around the z-axis.
[0066] The entire training process of the present invention is divided into four steps: the first step is mainly to determine the environmental parameters that affect the model performance; the second step is mainly to set a reasonable distribution range for each environmental parameter, in order to ensure the diversity of the simulation environment while avoiding the influence of invalid or extreme parameter configuration; the third step is mainly to train the robot arm to grab all objects from a relatively scattered working environment, in order to initially master the grasping skills of objects of different shapes and sizes; the fourth step is mainly to train the robot arm to grab the target object in a more cluttered working environment, in order to improve the efficiency of the action and the success rate of grabbing the target.
[0067] In this embodiment, the robot arm visual grasping method based on Sim2Real migration under space-constrained conditions includes the following steps:
[0068] Step 1: Build a task scene, select the visual domain randomization parameters of the simulation environment based on certain prior knowledge and experience, and conduct a grasping simulation experiment in the task scene to determine the randomization parameters; the randomization parameters include the camera pose; the brightness, saturation and contrast of the RGB image captured by the camera; and the Gaussian noise of the image;
[0069] Step 1-1: Randomly place 7 objects of different colors and shapes in a workspace, and the grasping task is to grasp all the objects;
[0070] Step 1-2: According to the possible camera poses in the real environment, the 3D position and rotation angle of the camera are randomized to simulate different perspective changes and train the network model deployed on the robot arm to guide the robot arm to achieve the grasping task. The network model is compared with the network model without randomized camera pose training to verify the effectiveness of the randomized camera pose.
[0071] Step 1-3: According to different lighting conditions in the real environment, adjust the brightness, saturation, contrast and other visual features of the RGB image captured by the camera for image enhancement and then train the network model. Compare it with the network model without image adjustment training to verify the effectiveness of image enhancement;
[0072] Step 1-4: According to the influence of sensor noise in the real environment, add Gaussian noise to the RGB image captured by the camera during network model training to train the model, and compare it with the model trained without adding image noise to verify the effectiveness of adding random noise;
[0073] Step 2: Set a reasonable randomization range for each randomization parameter to form a variety of different simulation environments;
[0074] Step 2-1: Set the camera position distribution range and camera angle distribution range;
[0075] In order to reflect the possible changes in camera posture in the real environment as much as possible, ensure that the camera position errors that may occur in reality can be covered, and avoid the generation of invalid training data due to excessive randomization range, the camera position distribution range is set to: (±0.05, ±0.025, ±0.05) m, that is, the camera randomly selects the three-dimensional position coordinates in a cuboid with a length of 0.1 m, a width of 0.05 m, and a height of 0.1 m; the camera angle distribution range is set to: (0, ±0.1, ±0.1) rad, that is, the randomization range of the pitch angle and yaw angle is ±0.1 rad;
[0076] Step 2-2: Set the visual feature variation range of the RGB image captured by the camera;
[0077] The brightness, saturation, and contrast of the RGB image captured by the camera are set to [0.7, 1.3], that is, between 70% and 130% of the brightness, contrast, and saturation of the original image;
[0078] Step 2-3: Set the Gaussian noise added during image enhancement;
[0079] The Gaussian noise added to the RGB image captured by the camera is set to have a mean of μ = 0 and a variance of σ = 25;
[0080] Step 2-4: Each round of training is randomly sampled from the randomized parameter distribution range set above, so as to generate different parameter configurations to form simulation environments with different conditions. That is, each round of training of the network model is trained in a simulation environment with different conditions, so that the network model can be exposed to a rich variety of scenarios;
[0081] Step 3: Train the first-stage grasping model in the random simulation environment generated in step 2, so that the robot arm grasps all objects; the first-stage grasping model includes a first perception module and a grasping module; the first perception module extracts image features; the grasping module predicts the Q value of each pixel in the image;
[0082] Step 3-1: Obtain an RGB-D image of the scene from a depth camera randomly placed above the workspace;
[0083] Step 3-2: Randomly change the brightness, saturation, and contrast of the RGB-D image;
[0084] Step 3-3: Add random Gaussian noise to the changed RGB-D image;
[0085] Step 3-4: Send the RGB-D image with random Gaussian noise added to the perception module network to extract image features, and then use the grasping module to predict the Q value of each pixel, where the Q value represents the expected future reward obtained by grasping at the pixel. The perception module uses a Densenet network to extract image features. The grasping module consists of convolution, normalization and activation layers, and is connected to the perception module through upsampling and skip connection operations to predict the Q value of the pixel.
[0086] Step 3-5: Perform a grabbing action at the pixel corresponding to the maximum Q value. If the grabbing is successful, you will receive a corresponding reward.
[0087] In this embodiment, the reward function for successfully grabbing and obtaining rewards is set to:
[0088]
[0089] Among them, R g is the reward function value;
[0090] Step 3-6: If 3000 grasping actions have been performed, stop the training of the first-stage grasping model and use the trained grasping model as the initialization model for the robot arm grasping task; otherwise, go to step 3-1;
[0091] Step 4: Train the second-stage grasping model in a cluttered and crowded simulation environment, so that the robot arm can grasp the target objects in sequence from the cluttered and crowded simulation environment; the second-stage grasping model includes a second perception module and a grasping module; the second perception module uses the SE-Densenet network to extract image features; the grasping module predicts the Q value of each pixel in the image;
[0092] Step 4-1: Randomly place M objects of different colors and shapes in a workspace. The grasping task is to grasp all objects of one color. Randomly select the color of the target object to be grasped and set the number of target objects to be grasped. Initialize the number of grasped target objects to 0.
[0093] Step 4-2: Get an RGB-D image of a random scene, change its brightness, contrast, saturation, and add random Gaussian noise; send the RGB-D image to the second perception module network to extract image features, and then use the capture module to predict the Q value of each pixel;
[0094] The SE-Densenet network is obtained by connecting the DenseNet-121 network and the attention mechanism SENet network, such as Figure 2 As shown in the figure: First, the DenseNet-121 network extracts the global features of the image, and then the global average pooling operation is used to compress the global characteristics output by the DenseNet-121 network to extract features more efficiently; then the number of channels of the feature map is reduced to one-fourth of the original through the first fully connected layer and activated using Relu, and then the number of channels is restored through the second fully connected layer to model the correlation between channels; then the importance weights of each channel are normalized through the Sigmoid function to ensure that the weight value is between 0 and 1, which is convenient for subsequent weighted operations; finally, the normalized weights are applied to each channel of the feature map, giving each channel a different importance, so that the SE-Densenet network can focus more on key features when predicting.
[0095] Step 4-3: Get a binary mask image containing all target objects, and then use the connected component function to get a binary mask image containing only a single target object and its center position;
[0096] Step 4-4: Calculate the edge binary mask map and edge height map corresponding to the binary mask map of the target object, and then calculate the boundary occupancy of the target object;
[0097] The edge binary mask map, edge depth height map and boundary occupancy rate corresponding to the binary mask map of the target object are shown in the following formula:
[0098]
[0099] Among them, m represents the binary edge mask of the target object, m depth Represents the edge depth height map of the target object; m obj Represents the binary mask image of the target object, m dilation Binary mask representing the target object Figure 2 Binary mask image after value dilation operation, H depth represents the depth height map, v represents the number of pixels with a value of 1 in the edge binary mask map m, and v o Represents the edge depth height map m depth The number of pixels with non-zero median values; r is the boundary occupancy of the target object.
[0100] Step 4-5: Select the target object with the smallest boundary occupancy rate as the current target to be grasped;
[0101] Step 4-6: If the boundary occupancy rate of the current target to be grasped is lower than the set threshold γ, the target is grasped; otherwise, the first search is performed with the target as the center to find graspable obstacles within 10 cm around it;
[0102] Step 4-7: If there is an obstacle with a boundary occupancy rate lower than the set threshold γ within 10 cm, the obstacle is captured; otherwise, the obstacle corresponding to the minimum boundary occupancy rate is taken as the center, and a second search is performed to capture the obstacle with the minimum boundary occupancy rate within 10 cm around it, such as Figure 3 As shown;
[0103] Step 4-8: If the capture is successful, you will receive corresponding rewards;
[0104] In this embodiment, the reward function for successful grasping is set to:
[0105]
[0106] Among them, R g ′ is the reward function value for successful grasping;
[0107] Step 4-9: Add 1 to the number of captured target objects;
[0108] Step 4-10: If 3500 grasping actions have been performed, stop training and save the second-stage grasping model; otherwise, determine whether there is still a target object in the workspace:
[0109] Case 1: If the target object is detected, then loop steps 4-2 to 4-10;
[0110] Case 2: If the target object is not detected and the number of captured target objects is less than the total number of target objects to be captured, it means that the target object is completely blocked by other obstacles and is not detected. Then, all objects in the workspace whose height exceeds the preset threshold are regarded as potential objects to be captured, and the object corresponding to the maximum Q value is captured until the target object enters the detection range, and steps 4-2 to 4-10 are looped.
[0111] Case 3: If no target object is detected and the number of captured target objects is equal to the number of target objects to be captured, go to step 4-1 to start a new round of training;
[0112] Step 5: Use the second-stage grasping model trained in step 4 to conduct target grasping test experiments in space-constrained scenarios;
[0113] Step 5-1: Set up two different test scenarios: one is the same as the training scenario in step 4, and the other is a space-constrained work scenario;
[0114] Step 5-2: Use the grasping model trained in step 4 to perform grasping tests on the robot arm in two different test scenarios;
[0115] Step 5-3: Compare the grasping result of step 5-2 with the baseline solution of combined push and grasp to verify the advantage of the grasping model trained by the present invention in grasping the target in a space-constrained scenario;
[0116] Step 5-5: Compare the capture results of step 5-2 with the capture solution without using the Sim2Real migration method to verify the advantages of the visual domain randomization solution used in the present invention.
[0117] In order to verify the effectiveness of the method of the present invention, this example conducts a test and a comparative experiment of the grasping model in two workspaces, normal and restricted. Figure 4 As shown, the comparison results are shown in Table 1-2. The results show that the target grasping success rate and action efficiency of the method of the present invention are better than those of the baseline method, and this is particularly evident in the confined workspace. This is because in space-constrained scenarios, the "pushing" action in the baseline method is often difficult to be effective, because the complexity of the workspace is always maintained at a high level, resulting in the inability to provide sufficient grasping space for the target, resulting in poor performance of the baseline solution. The grasping scheme of the present invention effectively reduces the complexity of the workspace by using the "grabbing" action to continuously remove obstacles in the workspace, which not only provides more ample grasping space for the target object, but also significantly improves the success rate of target grasping and the overall action efficiency.
[0118] Table 1 Normal working space
[0119] method Grab target success rate Action efficiency Task completion rate BPG 79.2% 67.4% 100% Method of the present invention 86.3% 85.8% 100%
[0120] Table 2 Restricted workspace
[0121] method Grab target success rate Action efficiency Task completion rate BPG 70.5% 40.2% 96.7% Method of the present invention 83.5% 82.6% 100%
[0122] Secondly, in order to verify the effectiveness of the perception network module proposed in this invention, the network structure before and after adding the SENet network module is used for model training. As the number of training iterations increases, the model's grasping performance curve is as follows: Figure 5 As shown in the figure. The grasping success rate of each step represents the success rate of grasping actions executed within 200 steps before the current step. As can be seen from the figure, the grasping success rate of the model is increased by nearly 10% after the introduction of the attention mechanism SENet, indicating that the network structure proposed in the present invention can effectively improve the grasping performance.
[0123] Finally, in order to verify the effectiveness of the visual domain randomization method proposed in this invention, the grasping performance with and without the visual domain randomization method is compared in the normal simulation workspace, and the comparison results are shown in Table 3. The results show that after using the visual domain randomization method, the average grasping success rate is improved accordingly, and the average number of actions is reduced accordingly, indicating that the visual domain randomization method proposed in this invention can effectively improve the robustness of the model.
[0124] Table 3 Comparison of grasping performance with and without visual domain randomization in normal simulation workspace
[0125] method Average crawl success rate Average number of actions Task completion rate SE-Densenet 93.1% 7.60 100% SE-Densenet+sim2real 95.5% 7.37 100%
[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.
Claims
1. A robotic arm visual grasping method based on Sim2Real migration under space-constrained conditions, characterized by: include: Build a task scene, select the visual domain randomization parameters of the simulation environment, and conduct a grasping simulation experiment in the task scene to determine the randomization parameters; the randomization parameters include the camera pose; the brightness, saturation and contrast of the RGB image captured by the camera; and the Gaussian noise of the image; Set the randomization range for each randomization parameter to form a variety of different simulation environments; The first-stage grasping model is trained in the generated random simulation environment to enable the robot arm to grasp all objects; the first-stage grasping model includes a first perception module and a grasping module; the first perception module extracts image features; The capture module predicts the Q value of each pixel in the image; The second-stage grasping model is trained in a cluttered and crowded simulation environment, so that the robot arm grasps the target objects in sequence from the cluttered and crowded simulation environment; the second-stage grasping model includes a second perception module and a grasping module; the second perception module extracts image features; The capture module predicts the Q value of each pixel in the image; The grasping model trained in the second stage is used to conduct target grasping tests in spatially constrained scenarios.
2. According to claim 1, a robot arm visual grasping method based on Sim2Real migration under space-constrained conditions is characterized by: The step of building a task scene, selecting a visual domain randomization parameter of a simulation environment, and performing a grasping simulation experiment in the task scene to determine the randomization parameter includes: Multiple objects of different colors and shapes are randomly placed in a workspace, and the grasping task is to grasp all the objects; According to the possible camera poses in the real environment, the 3D position and rotation angle of the camera are randomized to simulate different perspective changes and train the network model deployed on the robot arm to guide the robot arm to achieve the grasping task; According to different lighting conditions in real environments, the visual features of the RGB images captured by the camera are adjusted to perform image enhancement and then the network model is trained; According to the influence of sensor noise in the real environment, Gaussian noise is added to the RGB images taken by the camera during network model training to train the network model.
3. According to claim 2, a robot arm visual grasping method based on Sim2Real migration under space-constrained conditions is characterized in that: The randomization range is set for each randomization parameter to form a variety of different simulation environments, including: Set the camera position distribution range and camera angle distribution range; Set the range of visual feature changes of the RGB image captured by the camera; Set the Gaussian noise added during image enhancement; Each round of training is randomly sampled from the randomized parameter distribution range set above, so as to generate different parameter configurations to form simulation environments with different conditions, that is, each round of training of the network model is trained in a simulation environment with different conditions.
4. According to claim 3, a robot arm visual grasping method based on Sim2Real migration under space-constrained conditions is characterized in that: The camera position distribution range is set to: (±0.05, ±0.025, ±0.05) m, that is, the camera randomly selects the three-dimensional position coordinates in a cuboid with a length of 0.1 m, a width of 0.05 m, and a height of 0.1 m; the camera angle distribution range is set to: (0, ±0.1, ±0.1) rad, that is, the randomization range of the pitch angle and the yaw angle is ±0.1 rad; The brightness, saturation, and contrast of the RGB image captured by the camera are set to a range of [0.7, 1.3], that is, ranging from 70% to 130% of the brightness, contrast, and saturation of the original image; The Gaussian noise added to the RGB image captured by the camera is set to have a mean μ=0 and a variance σ=25.
5. According to claim 4, a robot arm visual grasping method based on Sim2Real migration under space-constrained conditions is characterized in that: The first stage grasping model is trained in the generated random simulation environment to enable the robot arm to grasp all objects, including: Acquire RGB-D images of the scene from depth cameras randomly placed above the workspace; Randomly change the brightness, saturation, and contrast of RGB-D images; Add random Gaussian noise to the altered RGB-D image; The RGB-D image with random Gaussian noise added is sent to the first perception module to extract image features, and then the grasping module is used to predict the Q value of each pixel, where the Q value represents the expected future reward obtained by grasping at the pixel. The first perception module uses a Densenet network to extract image features. The grasping module consists of convolution, normalization and activation layers, and is connected to the perception module through upsampling and skip connection operations to predict the Q value of the pixel. Perform the grabbing action at the pixel point corresponding to the maximum Q value, and obtain the corresponding reward if the grabbing is successful; If the set number of grasping actions has been performed, the first-stage grasping model training is stopped, and the trained first-stage grasping model is used as the initialization model for the robotic arm grasping task; otherwise, the RGB-D image of the scene is obtained again from the depth camera randomly placed above the workspace.
6. The robot arm visual grasping method based on Sim2Real migration under space-constrained conditions according to claim 5 is characterized in that: The training of the second-stage grasping model in the cluttered and crowded simulation environment enables the robot arm to grasp the target objects in sequence from the cluttered and crowded simulation environment, including: S1: Randomly place M objects of different colors and shapes in a workspace, and the grasping task is to grasp all objects of one color; randomly select the color of the target object to be grasped and set the number of target objects to be grasped, and initialize the number of grasped target objects to 0; S2: Obtain an RGB-D image of a random scene, change its brightness, contrast, saturation, and add random Gaussian noise; send the RGB-D image to the second perception module to extract image features, and then use the capture module to predict the Q value of each pixel; the second perception module uses the SE-Densenet network to extract image features; S3: Obtain a binary mask image containing all target objects, and then use the connected component function to obtain a binary mask image containing only a single target object and its center position; S4: Calculate the edge binary mask map and edge height map corresponding to the binary mask map of the target object, and then calculate the boundary occupancy rate of the target object; S5: Select the target object with the smallest boundary occupancy rate as the current target to be grasped; S6: If the boundary occupancy rate of the current target to be grasped is lower than the set threshold γ, the target is grasped; otherwise, the first search is performed with the target as the center to find obstacles that can be grasped within the set range around it; S7: If there is an obstacle with a boundary occupancy rate lower than the set threshold value γ within the set range, the obstacle is captured; otherwise, the obstacle corresponding to the minimum boundary occupancy rate is taken as the center, and a second search is performed to capture the obstacle with the minimum boundary occupancy rate within the set range around it; S8: If the capture is successful, you will receive corresponding rewards; S9: adding 1 to the number of captured target objects; S10: If the set number of grasping actions has been performed, stop training and save the second-stage grasping model; otherwise, determine whether there is still a target object in the workspace: if the target object is detected, loop through steps S2 to S10; If the target object is not detected and the number of captured target objects is less than the total number of target objects to be captured, it means that the target object is completely blocked by other obstacles and is not detected. Then, all objects in the workspace whose height exceeds the preset threshold are regarded as potential objects to be captured, and the object corresponding to the maximum Q value is captured until the target object enters the detection range, and steps S2 to S10 are cycled. If no target object is detected and the number of captured target objects is equal to the number of target objects to be captured, step S1 is re-executed to start a new round of training.
7. The robot arm visual grasping method based on Sim2Real migration under space-constrained conditions according to claim 6 is characterized in that: The SE-Densenet network is obtained by connecting the DenseNet-121 network and the attention mechanism SENet network: first, the DenseNet-121 network extracts the global features of the image, and then uses the global average pooling operation to compress the global features output by the DenseNet-121 network; then, the number of channels of the feature map is reduced to one-fourth of the original through the first fully connected layer and activated using Relu, and then the number of channels is restored through the second fully connected layer, so as to model the correlation between channels; Then, the importance weights of each channel are normalized by the Sigmoid function to ensure that the weight value is between 0 and 1; finally, the normalized weights are applied to each channel of the feature map, giving each channel a different importance, so that the SE-Densenet network can focus on key features when predicting.
8. A robot arm visual grasping system based on Sim2Real migration under space-constrained conditions, implemented based on the method of claim 1, characterized in that: include: Simulation environment construction module: sets the randomization range for each randomization parameter to form a variety of different simulation environments; The first training module: trains the first-stage grasping model in the generated random simulation environment to enable the robot arm to grasp all objects; Second training module: training the second-stage grasping model in a cluttered and crowded simulation environment, so that the robot arm can grasp the target objects in sequence from the cluttered and crowded simulation environment; Test module: Use the grasping model trained in the second stage to conduct target grasping tests in space-constrained scenarios.
9. An electronic device, comprising: One or more processors, and a memory, wherein the memory is used to store instructions, and when the instructions are executed by the one or more processors, the one or more processors execute the robotic arm visual grasping method based on Sim2Real migration under space-constrained conditions as described in claims 1-7.
10. A computer-readable storage medium storing executable instructions, which, when executed, enable a processor to execute the robotic arm visual grasping method based on Sim2Real migration under space-constrained conditions as described in claims 1-7.
Citation Information
Patent Citations
Mitigating reality gaps by training simulated-to-real models using vision-based robotic task models
CN114585487A
Vision-language-action joint modeling-based disordered scene target object capturing method
CN115861596A
Control unit and method for controlling robot to grip object
CN116652974A
Intelligent grabbing pose estimation method based on domain randomization
CN116958252A
Active data learning selection method for robot grasp
US20220212339A1
Cited By
Data set adaptive optimization method based on artificial intelligence
CN121706872A
Robot trajectory planning method and device for grabbing and placing tasks
CN121912387A
A robot trajectory planning method and device for a pick-and-place task
CN121912387B
Robot grabbing pose generation method and device
CN122023506A
A robot grasping pose generation method and device
CN122023506B