Gripper control method and device based on artificial intelligence

Through the artificial intelligence-based gripping control method, using RGB and deep image fusion, straight edge and corner point extraction and other technologies, the grab problem of changing shapes of non-orchy substances in ore mining is solved, and the grab accuracy and stability of the manipulator are improved.

CN119973976APending Publication Date: 2025-05-13HEFEI MINGDE PHOTOELECTRIC TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411247828.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-06
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Non-orch substances mixed in during ore mining, such as wood, steel nails, rags, plastic parts, etc., have a variety of shapes and postures, which increases the difficulty of accurate grasping by the robot, resulting in low grasping accuracy.

Method used

A gripper control method based on artificial intelligence is proposed. By obtaining the RGB image and depth image of the target object, extracting depth information and fusion, straight edges and corner points are extracted, bounding boxes are generated, target geometry and center of mass are determined, parallel straight edges are searched or image feature extraction is performed, the grab success, angle and width of each pixel point is calculated, the grab strategy is dynamically adjusted, and the grab posture of the manipulator is optimized.

Benefits of technology

It improves the accuracy and stability of the manipulator to grasp irregularly shaped objects, ensures that the manipulator can adapt to different object shapes and postures, reduces mistaken grasping and misplacement, and improves the efficiency of the entire grasping process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119973976A_ABST
    Figure CN119973976A_ABST
Patent Text Reader

Abstract

The invention discloses a gripper control method and device based on artificial intelligence, and relates to the technical field of artificial intelligence. Obtaining a fusion image of the target object, performing straight edge and corner extraction to obtain a straight edge set and a corner set, and generating a plurality of bounding boxes according to the straight edge set and the corner set; for each bounding box, acquiring a central point, and connecting all the central points to determine a target centroid of the target object; searching parallel straight edges in the fused image, and if the parallel straight edges do not exist, obtaining an execution parameter corresponding to each pixel point in the fused image; obtaining a target execution parameter through a preset rectangular window and the execution parameter of each pixel point; and the grabbing posture of the manipulator is determined according to the target execution parameters. According to the shape and the mass center of the target object determined according to the bounding box, the center of the object is positioned, the grabbing success degree, the grabbing angle and the grabbing width of each pixel point are calculated, it is ensured that the mechanical arm can adapt to different object shapes and postures, and the grabbing accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence, and in particular relates to a gripper control method and device based on artificial intelligence. Background Art

[0002] During the mining process, due to the complexity of geological conditions and the limitations of mining technology, various non-ore materials, such as wood, steel nails, rags, plastic parts, and waste filling pipes, will inevitably be mixed in. The mixing of these debris not only increases the impurity content of the ore, but also causes serious interference and impact on the subsequent process of transportation, crushing, grinding, and beneficiation. In order to solve this problem, manual sorting is often used to remove these foreign objects. However, manual sorting is not only inefficient and labor-intensive, but also has serious safety and occupational health risks. Workers need to work in harsh environments for a long time, exposed to harmful factors such as dust and noise, which can easily lead to occupational diseases. At the same time, due to the limitations of human eye recognition and judgment ability, manual sorting is often difficult to be completely thorough, resulting in incomplete sorting of foreign objects, which in turn affects the effect of subsequent process flows.

[0003] In order to solve the sorting problem caused by non-ore materials mixed in the ore mining process, the company has begun to use manipulators to realize automatic sorting. This solution significantly improves the work efficiency, because the manipulator can complete the sorting task uninterruptedly and quickly, and is not affected by physical fatigue or emotional fluctuations. More importantly, through the sorting operation of the manipulator, workers can reduce the direct exposure time in the harsh environment of dust and harsh noise, thereby greatly reducing the risk of occupational diseases caused by long-term exposure to these harmful factors, and further improving the safety of production operations. However, non-ore materials in the ore, such as wood, steel nails, rags, plastic parts, etc., have a variety of shapes and postures. These irregular shapes and postures increase the difficulty of accurate grasping by the manipulator, making the grasping accuracy of the manipulator not high. Summary of the invention

[0004] The purpose of the present invention is to solve the problem that non-ore materials in ore, such as wood, steel nails, rags, plastic parts, etc., have various shapes and postures. These irregular shapes and postures increase the difficulty of accurate grasping by the manipulator, resulting in low grasping accuracy of the manipulator, and propose a gripper control method and device based on artificial intelligence.

[0005] In a first aspect of the present invention, a gripper control method based on artificial intelligence is first proposed, the method comprising:

[0006] Acquire an RGB image and a depth image of the target object, extract depth information of the depth image, and replace the G channel of the RGB image with the depth information to obtain a fused image;

[0007] Extracting straight edges and corner points from the fused image to obtain a straight edge set and a corner point set, generating multiple bounding boxes based on the straight edge set and the corner point set; obtaining a center point for each bounding box, connecting all center points to obtain a target geometric shape, and determining a target centroid based on the target geometric shape;

[0008] Searching for parallel straight edges in the fused image, if they do not exist, extracting image features from the fused image to obtain a feature image, and obtaining execution parameters corresponding to each pixel point according to the feature image; the execution parameters include grasping success degree, grasping angle, and grasping width;

[0009] The target execution parameters are obtained by presetting the rectangular window and the execution parameters of each pixel point; the target execution parameters are mapped to the manipulator coordinate system to obtain the manipulator's grasping posture, and the manipulator grasps the target object according to the grasping posture; the target execution parameters include target grasping coordinates, target grasping angle and target grasping width.

[0010] Optionally, extracting straight edges and corner points from the fused image to obtain a straight edge set and a corner point set includes:

[0011] Preprocessing the fused image to obtain an initial image, and extracting the edge of the initial image using a Canny edge detection algorithm to obtain an edge map;

[0012] A straight edge set is obtained by extracting straight lines from the edge map through Hough transform, and a corner point set is obtained by identifying corner points in the edge map through Harris corner point detection algorithm.

[0013] Optionally, after searching for parallel straight edges in the fused image, the step further includes:

[0014] Step 1: If there are parallel straight edges, determine whether the width of two straight edges among the parallel straight edges is greater than the gripper width. If there is any straight edge whose width is less than the gripper width, search for parallel straight edges again until the width of two straight edges among the parallel straight edges is greater than the gripper width, and record them as target parallel straight edges; the gripper width is the width of the instrument hand;

[0015] Step 2: Calculate the distance between two straight edges of the target parallel straight edges to obtain the target distance. If the target distance is greater than or equal to the preset grasping width, return to step 1 until the target distance is less than the preset grasping width; the preset grasping width is the maximum opening width of the instrument hand;

[0016] Step 3: respectively obtain the bounding boxes corresponding to the two straight edges on the parallel straight edges to obtain the first bounding box and the second bounding box. If the length of the first bounding box is equal to the length of the second bounding box, proceed to step 4; if the length of the first bounding box is less than the length of the second bounding box, proceed to step 5; if the length of the first bounding box is greater than the length of the second bounding box, proceed to step 6;

[0017] Step 4: taking the midpoint of the first bounding box and the second bounding box as a clamping point;

[0018] Step 5: Project the midpoint of the first bounding box onto the second bounding box to obtain a first projection point, and use the first projection point and the midpoint of the first bounding box as clamping points;

[0019] Step six: Project the midpoint of the second bounding box onto the first bounding box to obtain a second projection point, and use the second projection point and the midpoint of the second bounding box as clamping points.

[0020] Optionally, performing image feature extraction on the fused image to obtain a feature image, and obtaining an execution parameter corresponding to each pixel point according to the feature image includes:

[0021] Substituting the fused image into a backbone network, the backbone network includes RSU-7, RSU-6, RSU-5, and RSU-4F; obtaining a first image by passing the fused image through the RSU-7, and fusing the first image with the fused image to obtain a first fused image;

[0022] Substituting the fused image into the RSU-6 to obtain a second image, and fusing the second image with the fused image to obtain a second fused image;

[0023] Fusing the first fused image and the second fused image to obtain a third fused image, substituting the third fused image into the RSU-5 to obtain a third image, performing multi-scale sampling on the third image to obtain a plurality of sampling images, and obtaining an output image through the RSU-4F for each sampling image;

[0024] For each output image, the output image is activated by the Sigma id activation function to obtain a grasping success image, the grasping success image is element-wise multiplied with the output image to obtain a potential grasping feature map, and the potential grasping feature map is substituted into the angle estimation branch and the width estimation branch to obtain an angle grasping feature map and a width grasping feature map;

[0025] All grasping success images are fused to obtain a final grasping success map, all angle grasping feature maps are fused to obtain a final grasping angle map, all width grasping feature maps are fused to obtain a final grasping width map, and the final grasping success map, the final grasping angle map and the final grasping width map are mapped to the fused image to obtain the grasping success, grasping angle and grasping width corresponding to each pixel in the fused image.

[0026] Optionally, mapping the target execution parameter to a manipulator coordinate system to obtain a grasping posture of the manipulator includes:

[0027] Acquire the target grabbing coordinates, and obtain the target position and target direction according to the target grabbing coordinates; map the target position and the target direction into the manipulator coordinate system, and determine the manipulator landing area according to the target grabbing width;

[0028] The angle corresponding to the falling area of ​​the manipulator is adjusted according to the target grasping angle to obtain the grasping posture of the manipulator.

[0029] In a second aspect of the present invention, a gripper control device based on artificial intelligence is provided, comprising:

[0030] An image fusion module is used to obtain an RGB image and a depth image of a target object, extract depth information of the depth image, and replace the G channel of the RGB image with the depth information to obtain a fused image;

[0031] a bounding box determination module, configured to extract straight edges and corner points from the fused image to obtain a straight edge set and a corner point set, generate multiple bounding boxes based on the straight edge set and the corner point set; obtain a center point for each bounding box, connect all center points to obtain a target geometric shape, and determine a target centroid based on the target geometric shape;

[0032] An execution parameter determination module is used to search for parallel straight edges in the fused image. If no parallel straight edges exist, image feature extraction is performed on the fused image to obtain a feature image, and execution parameters corresponding to each pixel are obtained according to the feature image; the execution parameters include grasping success degree, grasping angle and grasping width;

[0033] A grasping posture determination module is used to obtain target execution parameters through a preset rectangular window and the execution parameters of each pixel point; the target execution parameters are mapped to the manipulator coordinate system to obtain the grasping posture of the manipulator, and the manipulator grasps the target object according to the grasping posture; the target execution parameters include target grasping coordinates, target grasping angle and target grasping width.

[0034] Optionally, the bounding box determination module includes:

[0035] An edge map determination module is used to preprocess the fused image to obtain an initial image, and extract the edge of the initial image using a Canny edge detection algorithm to obtain an edge map;

[0036] The straight edge corner point extraction module is used to extract straight lines from the edge map through Hough transform to obtain a straight edge set, and to identify corner points in the edge map through Harris corner point detection algorithm to obtain a corner point set.

[0037] Optionally, the execution parameter determination module further includes:

[0038] A parallel straight edge determination module is used to determine whether the width of two straight edges among the parallel straight edges is greater than the gripper width if there are parallel straight edges. If there is any straight edge whose width is less than the gripper width, the parallel straight edges are searched again until the width of the two straight edges among the parallel straight edges is greater than the gripper width, and the parallel straight edges are recorded as target parallel straight edges; the gripper width is the width of the instrument hand;

[0039] A target distance calculation module, used for calculating the distance between two straight edges of the target parallel straight edges to obtain a target distance, and if the target distance is greater than or equal to a preset grasping width, returning to the parallel straight edge determination module until the target distance is less than a preset grasping width; the preset grasping width is the maximum opening width of the instrument hand;

[0040] a length comparison module, used to obtain the bounding boxes corresponding to the two straight edges on the parallel straight edges to obtain the first bounding box and the second bounding box respectively, and if the lengths of the first bounding box and the second bounding box are equal, enter the first clamping point determination module; if the length of the first bounding box is less than the length of the second bounding box, enter the second clamping point determination module; if the length of the first bounding box is greater than the length of the second bounding box, enter the third clamping point determination module;

[0041] A first clamping point determination module, configured to use the midpoint of the first bounding box and the second bounding box as a clamping point;

[0042] A second clamping point determination module, configured to project the midpoint of the first bounding box onto the second bounding box to obtain a first projection point, and to use the first projection point and the midpoint of the first bounding box as clamping points;

[0043] The third clamping point determination module is used to project the midpoint of the second bounding box onto the first bounding box to obtain a second projection point, and use the second projection point and the midpoint of the second bounding box as clamping points.

[0044] Optionally, the execution parameter determination module further includes:

[0045] A first fusion module is used to substitute the fused image into a backbone network, the backbone network includes RSU-7, RSU-6, RSU-5, and RSU-4F; the fused image is passed through the RSU-7 to obtain a first image, and the first image is fused with the fused image to obtain a first fused image;

[0046] A second fusion module, used for substituting the fused image into the RSU-6 to obtain a second image, and fusing the second image with the fused image to obtain a second fused image;

[0047] A third fusion module is used to fuse the first fused image and the second fused image to obtain a third fused image, substitute the third fused image into the RSU-5 to obtain a third image, perform multi-scale sampling on the third image to obtain multiple sampling images, and obtain an output image through RSU-4F for each sampling image;

[0048] A grasping image determination module is used for activating each output image through a sigmoid activation function to obtain a grasping success image, performing element-by-element multiplication of the grasping success image and the output image to obtain a potential grasping feature map, and substituting the potential grasping feature map into an angle estimation branch and a width estimation branch to obtain an angle grasping feature map and a width grasping feature map;

[0049] The pixel point grabbing parameter determination module is used to fuse all the grabbing success images to obtain the final grabbing success map, fuse all the angle grabbing feature maps to obtain the final grabbing angle map, fuse all the width grabbing feature maps to obtain the final grabbing width map, map the final grabbing success map, the final grabbing angle map and the final grabbing width map to the fused image, and obtain the grabbing success, grabbing angle and grabbing width corresponding to each pixel point in the fused image.

[0050] Optionally, the grasping posture determination module includes:

[0051] A manipulator drop area determination module is used to obtain the target grabbing coordinates, obtain the target position and the target direction according to the target grabbing coordinates; map the target position and the target direction to the manipulator coordinate system, and determine the manipulator drop area according to the target grabbing width;

[0052] The angle adjustment module is used to adjust the angle corresponding to the falling area of ​​the manipulator according to the target grasping angle to obtain the grasping posture of the manipulator.

[0053] Beneficial effects of the present invention:

[0054] The present invention proposes an artificial intelligence-based gripper control method, which includes obtaining an RGB image and a depth image of a target object, extracting depth information of the depth image, and replacing a G channel of the RGB image with the depth information to obtain a fused image; extracting straight edges and corner points from the fused image to obtain a straight edge set and a corner point set, and generating multiple bounding boxes according to the straight edge set and the corner point set; obtaining a center point for each bounding box, connecting all center points to obtain a target geometric shape, and determining a target centroid according to the target geometric shape; searching for parallel straight edges in the fused image, and if no straight edges exist, extracting image features from the fused image to obtain a feature image, and obtaining execution parameters corresponding to each pixel point according to the feature image; the execution parameters include a grasping success degree, a grasping angle, and a grasping width; obtaining target execution parameters by presetting a rectangular window and the execution parameters of each pixel point; mapping the target execution parameters to a manipulator coordinate system to obtain a grasping posture of the manipulator, and the manipulator grasps the target object according to the grasping posture. By replacing the G channel of the RGB image with depth information, the color and three-dimensional spatial information of the object are integrated. The bounding boxes generated by straight edge and corner point extraction, as well as the target geometry and center of mass determined based on these bounding boxes, can accurately locate the center and shape of the object. This information helps to optimize the grasping point and posture of the robot, thereby improving the accuracy of grasping. By calculating the grasping success rate, grasping angle, and grasping width of each pixel point, the solution can dynamically adjust the grasping strategy to ensure that the robot can adapt to different object shapes and postures and improve the accuracy of grasping. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] The present invention will be further described below in conjunction with the accompanying drawings.

[0056] Figure 1 A flowchart of a gripper control method based on artificial intelligence is provided for an embodiment of the present invention;

[0057] Figure 2 A structural schematic diagram of a gripper control device based on artificial intelligence is provided for an embodiment of the present invention. DETAILED DESCRIPTION

[0058] Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative work shall fall within the scope of protection of the present invention.

[0059] The embodiment of the present invention provides a gripper control method based on artificial intelligence. Figure 1 , Figure 1 A flowchart of a gripper control method based on artificial intelligence provided by an embodiment of the present invention. The method comprises the following steps:

[0060] S101, obtaining an RGB image and a depth image of a target object, extracting depth information of the depth image, and replacing the G channel of the RGB image with the depth information to obtain a fused image.

[0061] S102, extract straight edges and corner points from the fused image to obtain a straight edge set and a corner point set, generate multiple bounding boxes based on the straight edge set and the corner point set; obtain a center point for each bounding box, connect all center points to obtain a target geometry, and determine the target centroid based on the target geometry.

[0062] S103, searching for parallel straight edges in the fused image, if they do not exist, extracting image features from the fused image to obtain a feature image, and obtaining execution parameters corresponding to each pixel point according to the feature image; the execution parameters include grasping success degree, grasping angle and grasping width.

[0063] S104, obtaining target execution parameters by presetting the rectangular window and the execution parameters of each pixel point; mapping the target execution parameters to the manipulator coordinate system to obtain the grasping posture of the manipulator, and the manipulator grasps the target object according to the grasping posture.

[0064] The target execution parameters include target grabbing coordinates, target grabbing angle and target grabbing width.

[0065] An artificial intelligence-based gripper control method provided by an embodiment of the present invention replaces the G channel of an RGB image with depth information, integrates the color and three-dimensional space information of the object, and accurately locates the center and shape of the object through the bounding boxes generated by straight edge and corner point extraction, as well as the target geometry and center of mass determined based on these bounding boxes. This information helps to optimize the grasping points and postures of the manipulator, thereby improving the accuracy of grasping. By calculating the grasping success rate, grasping angle, and grasping width of each pixel point, the scheme can dynamically adjust the grasping strategy to ensure that the manipulator can adapt to different object shapes and postures and improve the accuracy of grasping.

[0066] In one implementation, more comprehensive information about the target object is obtained by fusing the RGB image and the depth image. The depth image provides three-dimensional information about the target object, which helps to accurately locate the surface features of the target, thereby improving the accuracy of the robot grasping; extracting the depth information of the depth image is to obtain the grayscale value of each pixel from the image, and regard these grayscale values ​​as the corresponding depth information.

[0067] In one implementation, by extracting straight edges and corner points, the robot can better identify the geometric shape of the object and generate a more accurate bounding box, which helps to avoid misidentification or misgrasping, especially in complex backgrounds or occlusions, when the target object is debris other than ore.

[0068] In one implementation, multiple execution parameters such as grasping success, grasping angle and grasping width are considered in the grasping strategy. By extracting and analyzing feature images, the system can select the optimal grasping posture for the manipulator to improve the stability and reliability of grasping.

[0069] In one implementation, the target execution parameter is obtained by using a preset rectangular window and the execution parameters of each pixel point, that is, finding the average value. The average value of the execution parameters of all pixel points in the preset rectangular window is calculated as the target execution parameter. The calculation of the average value can smooth out abnormal values ​​or extreme values, making the grasping decision more robust, especially when the surface features of the object are uneven. The preset rectangular window is determined by the technician.

[0070] In one implementation, using fused images for grasping planning can reduce misjudgments caused by image noise or other interference. The robot can adjust its posture more quickly and perform grasping actions, improving the efficiency of the entire grasping process. When there are no parallel straight edges in the image, the system can automatically adjust and continue grasping planning through feature image extraction, allowing the robot to work in complex and irregular environments without relying on specific scene conditions.

[0071] In one implementation, each bounding box contains parameters including center coordinates, length, width and angle, where the angle is the angle relative to the positive direction of the X-axis in the robot coordinate system. The robot coordinate system is established with the position of the robot as the origin and the positive direction of the X-axis being 90 degrees left in the positive direction of the robot; the parallel straight edges in the fused image are parallel straight edges relative to the target center of mass, that is, one straight edge is on one side of the target center of mass, and the other straight edge is on the other side of the target center of mass, and the two straight edges are parallel.

[0072] In one implementation, before the robot grasps, the target grasping coordinates, grasping angle, and grasping width are determined by precise calculation, which can significantly improve the success rate of grasping, especially when facing objects with complex or irregular shapes.

[0073] In one embodiment, extracting straight edges and corner points from the fused image to obtain a straight edge set and a corner point set includes:

[0074] The fused image is preprocessed to obtain an initial image, and the edge of the initial image is extracted using the Canny edge detection algorithm to obtain an edge map;

[0075] Straight lines are extracted from the edge map through Hough transform to obtain a straight edge set, and corner points in the edge map are identified through Harris corner detection algorithm to obtain a corner point set.

[0076] In one implementation, by preprocessing the fused image and applying the Canny algorithm, the contour of the target object can be clearly identified, laying the foundation for the subsequent extraction of straight lines and corner points; combining the detection of straight lines and corner points can more comprehensively describe the geometric shape of the target object. Straight lines provide boundary information of objects, while corner points mark the turning points or key structures of objects. The combination of the two can more accurately identify and describe the overall shape of objects, especially for objects with complex shapes.

[0077] In one embodiment, searching for parallel straight edges in the fused image further includes:

[0078] Step 1: If there are parallel straight edges, determine whether the width of two straight edges in the parallel straight edges is greater than the gripper width. If the width of any straight edge is less than the gripper width, search for parallel straight edges again until the width of two straight edges in the parallel straight edges is greater than the gripper width, and record them as target parallel straight edges; the gripper width is the width of the instrument hand;

[0079] Step 2: Calculate the distance between two straight edges of the target parallel straight edges to obtain the target distance. If the target distance is greater than or equal to the preset grasping width, return to step 1 until the target distance is less than the preset grasping width; the preset grasping width is the maximum opening width of the instrument hand;

[0080] Step 3: Obtain the bounding boxes corresponding to the two straight edges on the parallel straight edges respectively to obtain the first bounding box and the second bounding box. If the lengths of the first bounding box and the second bounding box are equal, proceed to step 4; if the length of the first bounding box is less than the length of the second bounding box, proceed to step 5; if the length of the first bounding box is greater than the length of the second bounding box, proceed to step 6;

[0081] Step 4: Taking the midpoint of the first bounding box and the second bounding box as the clamping point;

[0082] Step 5: Project the midpoint of the first bounding box onto the second bounding box to obtain a first projection point, and use the first projection point and the midpoint of the first bounding box as clamping points;

[0083] Step 6: Project the midpoint of the second bounding box onto the first bounding box to obtain a second projection point, and use the second projection point and the midpoint of the second bounding box as clamping points.

[0084] In one implementation, by ensuring that the width of the parallel straight edges is greater than the width of the gripper, it is possible to avoid the gripper being unstable due to the target being too narrow when performing grasping. This condition ensures that the gripper can firmly clamp the target object during the grasping process, reducing the possibility of grasping failure.

[0085] In one implementation, if all calculated parallel edges are not satisfied, the grasping scheme is determined by calculating the grasping degree of the pixel points, and the parallel straight edges in the fused image are searched by first searching for the parallel straight edges closest to the target centroid, and then expanding outward until the distance between two straight edges in the target parallel straight edges is found to be less than the preset grasping width, or all parallel straight edges are traversed.

[0086] In one implementation, the midpoint of the first bounding box is projected onto the second bounding box to obtain a first projection point, and the midpoint of the second bounding box is projected onto the first bounding box to obtain a second projection point. Specifically, for the first bounding box and the second bounding box, the side closest to the target centroid is taken as the distance opposite side, the shortest perpendicular line to the distance opposite side is obtained to obtain the target perpendicular line, and the perpendicular bisector of the target perpendicular line is obtained. According to the perpendicular bisector, the midpoint of the small bounding box is mapped to the large bounding box to obtain the projection point.

[0087] In one implementation, by calculating the distance between the parallel straight edges and ensuring that it is less than the preset grip width, the system can determine whether the target object fits the maximum opening width of the current gripper. If the target distance is too large, the gripper may not be able to fully grip the object. By returning to step one and re-adjusting, the success rate of the gripping operation can be ensured. By projecting steps five and six, the system can accurately calculate the position of the gripping point. This calculation is based on geometric relationships and helps ensure that the gripping point is in the optimal position between the two parallel straight edges, thereby maximizing the accuracy and stability of the grip.

[0088] In one embodiment, image features are extracted from the fused image to obtain a feature image, and execution parameters corresponding to each pixel point are obtained according to the feature image, including:

[0089] Substitute the fused image into the backbone network, which includes RSU-7, RSU-6, RSU-5, and RSU-4F; the fused image is passed through RSU-7 to obtain a first image, and the first image and the fused image are fused to obtain a first fused image;

[0090] Substituting the fused image into RSU-6 to obtain a second image, and fusing the second image with the fused image to obtain a second fused image;

[0091] The first fused image and the second fused image are fused to obtain a third fused image, the third fused image is substituted into RSU-5 to obtain a third image, multi-scale sampling is performed on the third image to obtain multiple sampling images, and an output image is obtained by RSU-4F for each sampling image;

[0092] For each output image, the output image is activated by the Sigma id activation function to obtain a grasping success image, the grasping success image is element-wise multiplied with the output image to obtain a potential grasping feature map, and the potential grasping feature map is substituted into the angle estimation branch and the width estimation branch to obtain an angle grasping feature map and a width grasping feature map;

[0093] All grasping success images are fused to obtain the final grasping success map, all angle grasping feature maps are fused to obtain the final grasping angle map, and all width grasping feature maps are fused to obtain the final grasping width map. The final grasping success map, the final grasping angle map and the final grasping width map are mapped to the fused image to obtain the grasping success, grasping angle and grasping width corresponding to each pixel in the fused image.

[0094] In one implementation, by using multiple network modules such as RSU-7, RSU-6, RSU-5 and RSU-4F, the system can extract image features at different scales. Multi-scale feature extraction enables the system to capture the details and global information of the target object, enhancing the ability to recognize the target object. Each image after passing through the RSU module will be fused with the original fused image. This progressive feature fusion method can retain information at different levels, ensuring that the final generated feature image has rich multi-level features, and improving the system's ability to understand objects in complex scenes.

[0095] In one implementation, a grasping success image is generated by a Sigma ID activation function and is element-wise multiplied with the output image, which can further highlight areas with a higher grasping success probability, improve the prediction accuracy of grasping success, and reduce the risk of erroneous grasping.

[0096] In one implementation, the potential grasping feature map is passed through the angle estimation branch and the width estimation branch respectively to obtain the angle grasping feature map and the width grasping feature map, which can more accurately estimate the grasping posture of the manipulator including the grasping angle and the grasping width, and ensure the accuracy and effectiveness of the grasping action; wherein the angle estimation branch is the main task of the angle estimation branch to estimate an optimal grasping angle for each potential grasping position. The grasping angle refers to the optimal rotation angle of the manipulator relative to the target object, which usually needs to be aligned with the geometric shape of the object to ensure the stability and success rate of the grasping. The processing flow is to send the grasping feature map into the angle estimation branch. The angle estimation branch is usually composed of a series of convolutional layers, which can extract angle-related information from the input feature map and identify rotation-invariant features and edge features in the image. After being processed by multiple convolutional layers, the network outputs a prediction map, and the value of each pixel represents the optimal angle for grasping at that position; the purpose of the width estimation branch is to estimate the width of the object from the image, that is, the actual size of the object in the grasping direction. Accurate width estimation helps ensure that the gripper or tool matches the object. The width estimation branch also contains a series of convolutional layers to extract important features from the feature map, and finally predicts the width through one or more regression layers. These layers output a real value representing the width of the object.

[0097] In one embodiment, mapping the target execution parameters to the manipulator coordinate system to obtain the manipulator's grasping posture includes:

[0098] Obtain the target grasping coordinates, and obtain the target position and target direction according to the target grasping coordinates; map the target position and target direction to the manipulator coordinate system, and determine the manipulator landing area according to the target grasping width;

[0099] The angle corresponding to the manipulator's falling area is adjusted according to the target grasping angle to obtain the manipulator's grasping posture.

[0100] In one implementation, the target object can be accurately located by obtaining the target grasping coordinates and calculating the target position and target direction. Accurate positioning is the basis for efficient and reliable grasping. The target position and target direction are mapped to the manipulator coordinate system to ensure that the manipulator can accurately reach the target position and grasp in the correct direction, thereby reducing errors.

[0101] In one implementation, the falling area recognition and grasping angle of the manipulator are first obtained, and then the grasping angle of the corresponding area in the fused image is searched according to the falling area and adjusted, so as to achieve more accurate grasping.

[0102] In one implementation, the manipulator can calculate and set a suitable drop area based on the width of the target object. This helps avoid collisions or unstable grasping, ensures that the manipulator can drop and grasp the object in the correct position, adjusts the grasping posture of the manipulator based on the target grasping angle, adapts it to the specific angle of the target object, can handle objects of different postures, and improves the flexibility of the manipulator.

[0103] Based on the same inventive concept, the embodiment of the present invention also provides a gripper control device based on artificial intelligence. Figure 2 , Figure 2 A schematic structural diagram of a gripper control device based on artificial intelligence provided by an embodiment of the present invention includes:

[0104] An image fusion module is used to obtain an RGB image and a depth image of the target object, extract the depth information of the depth image, and replace the G channel of the RGB image with the depth information to obtain a fused image;

[0105] A bounding box determination module is used to extract straight edges and corner points from the fused image to obtain a straight edge set and a corner point set, and generate multiple bounding boxes based on the straight edge set and the corner point set; for each bounding box, obtain the center point, connect all the center points to obtain the target geometry, and determine the target centroid based on the target geometry;

[0106] An execution parameter determination module is used to search for parallel straight edges in the fused image. If they do not exist, the fused image is subjected to image feature extraction to obtain a feature image, and the execution parameters corresponding to each pixel are obtained according to the feature image; the execution parameters include grasping success, grasping angle, and grasping width;

[0107] The grasping posture determination module is used to obtain the target execution parameters through the preset rectangular window and the execution parameters of each pixel point; the target execution parameters are mapped to the manipulator coordinate system to obtain the grasping posture of the manipulator, and the manipulator grasps the target object according to the grasping posture; the target execution parameters include target grasping coordinates, target grasping angle and target grasping width.

[0108] An artificial intelligence-based gripper control device provided according to an embodiment of the present invention replaces the G channel of an RGB image with depth information, thereby integrating the color and three-dimensional spatial information of an object. The bounding boxes generated by straight edge and corner point extraction, as well as the target geometry and center of mass determined based on these bounding boxes, can accurately locate the center and shape of the object. This information helps to optimize the grasping points and postures of the manipulator, thereby improving the accuracy of grasping. By calculating the grasping success rate, grasping angle, and grasping width of each pixel point, the solution can dynamically adjust the grasping strategy to ensure that the manipulator can adapt to different object shapes and postures, thereby improving the accuracy of grasping.

[0109] In one embodiment, the bounding box determination module includes:

[0110] An edge map determination module is used to preprocess the fused image to obtain an initial image, and extract the edge of the initial image using the Canny edge detection algorithm to obtain an edge map;

[0111] The straight edge corner point extraction module is used to extract straight lines from the edge map through Hough transform to obtain a straight edge set, and to identify corner points in the edge map through Harris corner point detection algorithm to obtain a corner point set.

[0112] In one embodiment, the execution parameter determination module further includes:

[0113] The parallel straight edge determination module is used to determine whether the width of two straight edges in the parallel straight edges is greater than the gripper width if there are parallel straight edges. If the width of any straight edge is less than the gripper width, the parallel straight edges are searched again until the width of the two straight edges in the parallel straight edges is greater than the gripper width, and the parallel straight edges are recorded as target parallel straight edges. The gripper width is the width of the instrument hand.

[0114] The target distance calculation module is used to calculate the distance between two straight edges of the target parallel straight edges to obtain the target distance. If the target distance is greater than or equal to the preset grasping width, the module returns to the parallel straight edge determination module until the target distance is less than the preset grasping width; the preset grasping width is the maximum opening width of the instrument hand;

[0115] A length comparison module is used to obtain the corresponding bounding boxes on the two straight edges on the parallel straight edges to obtain the first bounding box and the second bounding box. If the lengths of the first bounding box and the second bounding box are equal, the first clamping point determination module is entered; if the length of the first bounding box is less than the length of the second bounding box, the second clamping point determination module is entered; if the length of the first bounding box is greater than the length of the second bounding box, the third clamping point determination module is entered;

[0116] A first clamping point determination module, used to take the midpoint of the first bounding box and the second bounding box as the clamping point;

[0117] A second clamping point determination module, used for projecting the midpoint of the first bounding box onto the second bounding box to obtain a first projection point, and taking the first projection point and the midpoint of the first bounding box as clamping points;

[0118] The third clamping point determination module is used to project the midpoint of the second bounding box onto the first bounding box to obtain a second projection point, and use the second projection point and the midpoint of the second bounding box as clamping points.

[0119] In one embodiment, the execution parameter determination module further includes:

[0120] The first fusion module is used to substitute the fused image into the backbone network, the backbone network includes RSU-7, RSU-6, RSU-5, and RSU-4F; the fused image is passed through RSU-7 to obtain a first image, and the first image is fused with the fused image to obtain a first fused image;

[0121] A second fusion module, used for substituting the fused image into RSU-6 to obtain a second image, and fusing the second image with the fused image to obtain a second fused image;

[0122] A third fusion module is used to fuse the first fused image and the second fused image to obtain a third fused image, substitute the third fused image into RSU-5 to obtain a third image, perform multi-scale sampling on the third image to obtain multiple sampling images, and obtain an output image for each sampling image through RSU-4F;

[0123] A grasping image determination module is used for activating each output image through a sigmoid activation function to obtain a grasping success image, performing element-by-element multiplication of the grasping success image and the output image to obtain a potential grasping feature map, and substituting the potential grasping feature map into an angle estimation branch and a width estimation branch to obtain an angle grasping feature map and a width grasping feature map;

[0124] The pixel grabbing parameter determination module is used to fuse all grabbing success images to obtain the final grabbing success map, fuse all angle grabbing feature maps to obtain the final grabbing angle map, fuse all width grabbing feature maps to obtain the final grabbing width map, map the final grabbing success map, the final grabbing angle map and the final grabbing width map to the fused image, and obtain the grabbing success, grabbing angle and grabbing width corresponding to each pixel in the fused image.

[0125] In one embodiment, the grasping posture determination module includes:

[0126] The manipulator falling area determination module is used to obtain the target grasping coordinates, obtain the target position and target direction according to the target grasping coordinates; map the target position and target direction to the manipulator coordinate system, and determine the manipulator falling area according to the target grasping width;

[0127] The angle adjustment module is used to adjust the angle corresponding to the manipulator's falling area according to the target grasping angle to obtain the manipulator's grasping posture.

[0128] The above is a detailed description of an embodiment of the present invention, but the content is only a preferred embodiment of the present invention and cannot be considered to limit the scope of implementation of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.

Claims

1. A gripper control method based on artificial intelligence, characterized in that: The method comprises: Acquire an RGB image and a depth image of the target object, extract depth information of the depth image, and replace the G channel of the RGB image with the depth information to obtain a fused image; Extracting straight edges and corner points from the fused image to obtain a straight edge set and a corner point set, and generating a plurality of bounding boxes according to the straight edge set and the corner point set; For each bounding box, obtain the center point, connect all the center points to obtain the target geometry, and determine the target centroid according to the target geometry; Searching for parallel straight edges in the fused image, if they do not exist, extracting image features from the fused image to obtain a feature image, and obtaining execution parameters corresponding to each pixel point according to the feature image; the execution parameters include grasping success degree, grasping angle, and grasping width; The target execution parameters are obtained by presetting the rectangular window and the execution parameters of each pixel point; the target execution parameters are mapped to the manipulator coordinate system to obtain the manipulator's grasping posture, and the manipulator grasps the target object according to the grasping posture; the target execution parameters include target grasping coordinates, target grasping angle and target grasping width.

2. The gripper control method based on artificial intelligence according to claim 1 is characterized in that: Extracting straight edges and corner points from the fused image to obtain a straight edge set and a corner point set includes: Preprocessing the fused image to obtain an initial image, and extracting the edge of the initial image using a Canny edge detection algorithm to obtain an edge map; A straight edge set is obtained by extracting straight lines from the edge map through Hough transform, and a corner point set is obtained by identifying corner points in the edge map through Harris corner point detection algorithm.

3. The gripper control method based on artificial intelligence according to claim 1 is characterized in that: After searching for parallel straight edges in the fused image, the following steps are further included: Step 1: If there are parallel straight edges, determine whether the width of two straight edges among the parallel straight edges is greater than the gripper width. If there is any straight edge whose width is less than the gripper width, search for parallel straight edges again until the width of two straight edges among the parallel straight edges is greater than the gripper width, and record them as target parallel straight edges; the gripper width is the width of the instrument hand; Step 2: Calculate the distance between two straight edges of the target parallel straight edges to obtain the target distance. If the target distance is greater than or equal to the preset grasping width, return to step 1 until the target distance is less than the preset grasping width; the preset grasping width is the maximum opening width of the instrument hand; Step 3: respectively obtain the bounding boxes corresponding to the two straight edges on the parallel straight edges to obtain the first bounding box and the second bounding box. If the length of the first bounding box is equal to the length of the second bounding box, proceed to step 4; if the length of the first bounding box is less than the length of the second bounding box, proceed to step 5; if the length of the first bounding box is greater than the length of the second bounding box, proceed to step 6; Step 4: taking the midpoint of the first bounding box and the second bounding box as a clamping point; Step 5: Project the midpoint of the first bounding box onto the second bounding box to obtain a first projection point, and use the first projection point and the midpoint of the first bounding box as clamping points; Step six: Project the midpoint of the second bounding box onto the first bounding box to obtain a second projection point, and use the second projection point and the midpoint of the second bounding box as clamping points.

4. The gripper control method based on artificial intelligence according to claim 1 is characterized in that: Extracting image features from the fused image to obtain a feature image, and obtaining execution parameters corresponding to each pixel point according to the feature image includes: Substituting the fused image into a backbone network, the backbone network includes RSU-7, RSU-6, RSU-5, and RSU-4F; obtaining a first image by passing the fused image through the RSU-7, and fusing the first image with the fused image to obtain a first fused image; Substituting the fused image into the RSU-6 to obtain a second image, and fusing the second image with the fused image to obtain a second fused image; Fusing the first fused image and the second fused image to obtain a third fused image, substituting the third fused image into the RSU-5 to obtain a third image, performing multi-scale sampling on the third image to obtain a plurality of sampling images, and obtaining an output image through the RSU-4F for each sampling image; For each output image, the output image is activated by the Sigmoid activation function to obtain a grasping success image, the grasping success image is multiplied element by element with the output image to obtain a potential grasping feature map, and the potential grasping feature map is substituted into the angle estimation branch and the width estimation branch to obtain an angle grasping feature map and a width grasping feature map; All grasping success images are fused to obtain a final grasping success map, all angle grasping feature maps are fused to obtain a final grasping angle map, all width grasping feature maps are fused to obtain a final grasping width map, and the final grasping success map, the final grasping angle map and the final grasping width map are mapped to the fused image to obtain the grasping success, grasping angle and grasping width corresponding to each pixel in the fused image.

5. The gripper control method based on artificial intelligence according to claim 1 is characterized in that: Mapping the target execution parameters to the manipulator coordinate system to obtain the manipulator's grasping posture includes: Acquire the target grabbing coordinates, and obtain the target position and target direction according to the target grabbing coordinates; map the target position and the target direction into the manipulator coordinate system, and determine the manipulator landing area according to the target grabbing width; The angle corresponding to the falling area of ​​the manipulator is adjusted according to the target grasping angle to obtain the grasping posture of the manipulator.

6. A gripper control device based on artificial intelligence, characterized in that: The device comprises: An image fusion module is used to obtain an RGB image and a depth image of a target object, extract depth information of the depth image, and replace the G channel of the RGB image with the depth information to obtain a fused image; a bounding box determination module, configured to extract straight edges and corner points from the fused image to obtain a straight edge set and a corner point set, generate multiple bounding boxes based on the straight edge set and the corner point set; obtain a center point for each bounding box, connect all center points to obtain a target geometric shape, and determine a target centroid based on the target geometric shape; An execution parameter determination module is used to search for parallel straight edges in the fused image. If no parallel straight edges exist, image feature extraction is performed on the fused image to obtain a feature image, and execution parameters corresponding to each pixel are obtained according to the feature image; the execution parameters include grasping success degree, grasping angle and grasping width; A grasping posture determination module is used to obtain target execution parameters through a preset rectangular window and the execution parameters of each pixel point; the target execution parameters are mapped to the manipulator coordinate system to obtain the grasping posture of the manipulator, and the manipulator grasps the target object according to the grasping posture; the target execution parameters include target grasping coordinates, target grasping angle and target grasping width.

7. The gripper control device based on artificial intelligence according to claim 6, characterized in that: The bounding box determination module includes: An edge map determination module is used to preprocess the fused image to obtain an initial image, and extract the edge of the initial image using a Canny edge detection algorithm to obtain an edge map; The straight edge corner point extraction module is used to extract straight lines from the edge map through Hough transform to obtain a straight edge set, and to identify corner points in the edge map through Harris corner point detection algorithm to obtain a corner point set.

8. The gripper control method based on artificial intelligence according to claim 6 is characterized in that: The execution parameter determination module also includes: A parallel straight edge determination module is used to determine whether the width of two straight edges among the parallel straight edges is greater than the gripper width if there are parallel straight edges. If there is any straight edge whose width is less than the gripper width, the parallel straight edges are searched again until the width of the two straight edges among the parallel straight edges is greater than the gripper width, and the parallel straight edges are recorded as target parallel straight edges; the gripper width is the width of the instrument hand; A target distance calculation module, used for calculating the distance between two straight edges of the target parallel straight edges to obtain a target distance, and if the target distance is greater than or equal to a preset grasping width, returning to the parallel straight edge determination module until the target distance is less than a preset grasping width; the preset grasping width is the maximum opening width of the instrument hand; a length comparison module, used to obtain the bounding boxes corresponding to the two straight edges on the parallel straight edges to obtain the first bounding box and the second bounding box respectively, and if the lengths of the first bounding box and the second bounding box are equal, enter the first clamping point determination module; if the length of the first bounding box is less than the length of the second bounding box, enter the second clamping point determination module; if the length of the first bounding box is greater than the length of the second bounding box, enter the third clamping point determination module; A first clamping point determination module, configured to use the midpoint of the first bounding box and the second bounding box as a clamping point; A second clamping point determination module, configured to project the midpoint of the first bounding box onto the second bounding box to obtain a first projection point, and to use the first projection point and the midpoint of the first bounding box as clamping points; The third clamping point determination module is used to project the midpoint of the second bounding box onto the first bounding box to obtain a second projection point, and use the second projection point and the midpoint of the second bounding box as clamping points.

9. The gripper control device based on artificial intelligence according to claim 6, characterized in that: The execution parameter determination module also includes: A first fusion module is used to substitute the fused image into a backbone network, the backbone network includes RSU-7, RSU-6, RSU-5, and RSU-4F; the fused image is passed through the RSU-7 to obtain a first image, and the first image is fused with the fused image to obtain a first fused image; A second fusion module, used for substituting the fused image into the RSU-6 to obtain a second image, and fusing the second image with the fused image to obtain a second fused image; A third fusion module is used to fuse the first fused image and the second fused image to obtain a third fused image, substitute the third fused image into the RSU-5 to obtain a third image, perform multi-scale sampling on the third image to obtain multiple sampling images, and obtain an output image through RSU-4F for each sampling image; A grasping image determination module is used to activate each output image through a Sigmoid activation function to obtain a grasping success image, multiply the grasping success image by the output image element by element to obtain a potential grasping feature map, and substitute the potential grasping feature map into the angle estimation branch and the width estimation branch to obtain an angle grasping feature map and a width grasping feature map; The pixel point grabbing parameter determination module is used to fuse all the grabbing success images to obtain the final grabbing success map, fuse all the angle grabbing feature maps to obtain the final grabbing angle map, fuse all the width grabbing feature maps to obtain the final grabbing width map, map the final grabbing success map, the final grabbing angle map and the final grabbing width map to the fused image, and obtain the grabbing success, grabbing angle and grabbing width corresponding to each pixel point in the fused image.

10. The gripper control device based on artificial intelligence according to claim 6, characterized in that: The grasping posture determination module comprises: A manipulator drop area determination module is used to obtain the target grabbing coordinates, obtain the target position and the target direction according to the target grabbing coordinates; map the target position and the target direction to the manipulator coordinate system, and determine the manipulator drop area according to the target grabbing width; The angle adjustment module is used to adjust the angle corresponding to the falling area of ​​the manipulator according to the target grasping angle to obtain the grasping posture of the manipulator.