Robot grabbing control method and device, computer equipment, readable storage medium and program product
By generating 3D point cloud information and filtering collision-free target grasping points, the adaptability and accuracy problems of traditional robot grasping control methods in complex environments are solved, and efficient grasping operations are achieved.
Patent Information
- Application Number
- CN202511848411.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-09
- Publication Date
- 2026-02-27
AI Technical Summary
Traditional robot grasping control methods are poorly adaptable to complex and unstructured environments, lack grasping accuracy, and are particularly inefficient when facing multiple targets or dynamic scenarios.
By fusing object images and depth images to generate 3D point cloud information, multiple grasping points are predicted, and non-collision target grasping points are selected based on point cloud information comparison to perform precise grasping operations.
It improves the robot's success rate and adaptability in complex scenarios, reduces the need for human intervention, and significantly enhances the accuracy and feasibility of grasp point prediction.
Smart Images

Figure CN121572307A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robot control technology, and in particular to a robot grasping control method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Technology
[0002] With the development of robot control technology, robot grasping control technology has become a core research direction in fields such as intelligent manufacturing, logistics warehousing, and service robots. Traditional robot grasping control mainly relies on preset fixed programs or offline planning to achieve grasping tasks through precise trajectory control of the robotic arm. These methods are stable in structured environments, but have poor adaptability to dynamic scenarios. Especially when facing complex stacked, multi-object, or unstructured environments, traditional methods require manual design of grasping strategies or adjustment of parameters through multiple trials and errors, resulting in low efficiency and insufficient generalization ability.
[0003] In recent years, vision-based robot grasping technology has gradually become mainstream, identifying target objects and planning grasping points through two-dimensional images or depth information. However, existing methods mostly rely on single-modal data, which makes it difficult to comprehensively describe the three-dimensional spatial characteristics of objects, resulting in limited accuracy in grasping point prediction. Therefore, traditional robot grasping control methods suffer from poor grasping accuracy. Summary of the Invention
[0004] Therefore, it is necessary to provide a robot grasping control method, device, computer equipment, computer-readable storage medium, and computer program product that can improve grasping accuracy in response to the above-mentioned technical problems.
[0005] In a first aspect, this application provides a robot grasping control method, including:
[0006] Acquire object images and depth images for the grasping area; the object images include object information for each of the grasping objects;
[0007] A three-dimensional point cloud information of the grasping area is generated based on the depth image; the three-dimensional point cloud information includes the object point cloud information of each grasping object, as well as other point cloud information of other objects in the grasping area;
[0008] For each of the aforementioned objects to be crawled, crawling point prediction is performed based on the object point cloud information and object information of the object to be crawled, thereby obtaining multiple crawling points predicted for the object to be crawled, as well as crawling point information for each of the aforementioned crawling points.
[0009] For each of the aforementioned grasping points, the grasping point information of the grasping point is compared with the other point cloud information, and the grasping points whose comparison results meet the comparison conditions are taken as target grasping points;
[0010] Perform a crawling operation based on the target crawling point information.
[0011] In one embodiment, acquiring object images and depth images for multiple objects in the grasping area includes:
[0012] Obtain the initial object image and initial depth image for the grasped area;
[0013] Object recognition is performed on the initial object image to obtain multiple initial object information;
[0014] Image analysis is performed on the initial object image based on the multiple initial object information to obtain an object image including object information of each of the multiple crawling objects; the object information is selected from the multiple initial object information; the image range of the object image is less than or equal to that of the initial object image;
[0015] Based on the size of the object image, the initial depth image is adjusted to obtain the depth image corresponding to the object image.
[0016] In one embodiment, the initial crawled object information includes object confidence; the step of performing image analysis on the initial object image based on the multiple initial crawled object information to obtain an object image including object information of each of the multiple crawled objects includes:
[0017] The initial crawling object information that does not meet the confidence condition is removed from the initial object image to obtain a candidate object image that includes the object information of each of the multiple candidate crawling objects.
[0018] Based on the central region of the object image, the candidate object image is filtered by the central region to obtain an object image that includes the object information of each of the multiple objects to be crawled.
[0019] In one embodiment, the step of filtering the candidate object image based on the central region of the object image to obtain an object image including object information of multiple crawling objects includes:
[0020] For each candidate object to be captured, the area overlap between the image range of the candidate object and the central region of the object image is determined.
[0021] The object information of candidate crawling objects whose area overlap does not meet the overlap condition is filtered out from the candidate object image to obtain an object image that includes the object information of multiple crawling objects whose area overlap meets the overlap condition.
[0022] In one embodiment, the grasping point information includes grasping point depth; the step of comparing the grasping point information with other point cloud information, and taking the grasping points whose comparison results meet the comparison conditions as target grasping points, includes:
[0023] The grab point information of the grab point is compared with the other point cloud information, and the grab points whose comparison results meet the comparison conditions are selected as candidate grab points.
[0024] From the candidate grab points, the candidate grab points whose grab point depth satisfies the optimal depth condition are determined as the target grab points.
[0025] In one embodiment, the method further includes:
[0026] Obtain the centroid position of the object to which the target grab point belongs;
[0027] Based on the centroid position of the object, the target grasping point information of the target grasping point is adjusted to obtain updated grasping point information;
[0028] The step of performing the crawling operation according to the target crawling point information includes:
[0029] Perform the crawling operation according to the updated crawling point information of the target crawling point.
[0030] Secondly, this application also provides a robot grasping control device, comprising:
[0031] The image acquisition module is used to acquire object images and depth images of the grasping area; the object images include object information of multiple grasping objects.
[0032] The point cloud information generation module is used to generate three-dimensional point cloud information of the grasping area based on the depth image; the three-dimensional point cloud information includes the object point cloud information of each of the grasping objects, as well as other point cloud information of other objects in the grasping area;
[0033] The grasping point prediction module is used to predict grasping points for each grasping object based on the object point cloud information and object information of the grasping object, so as to obtain multiple grasping points predicted for the grasping object and grasping point information of each grasping point.
[0034] The information comparison module is used to compare the grasping point information of each grasping point with the other point cloud information, and take the grasping point whose comparison result meets the comparison conditions as the target grasping point.
[0035] The capture operation execution module is used to perform capture operations according to the target capture point information of the target capture point.
[0036] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the method described above.
[0037] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the method described above.
[0038] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of the method described above.
[0039] The aforementioned robot grasping control method, apparatus, computer equipment, computer-readable storage medium, and computer program product, by fusing object images and depth images, first generate 3D point cloud information containing the grasped object and its environment, effectively integrating the spatial structure and semantic features of the object and solving the problem of insufficient single-modal data information. Second, for each grasped object, multiple candidate grasping points are independently predicted, and dynamic comparison is performed between the grasping point information and the environmental point cloud information to automatically filter out collision-free and reachable target grasping points, significantly improving the accuracy and feasibility of grasping point prediction. Finally, the grasping operation is executed based on the precise information of the target grasping points, which not only reduces the need for manual intervention but also significantly improves the robot's grasping success rate and adaptability in complex scenarios. In summary, adopting the above-mentioned robot grasping control method can improve the accuracy of robot grasping control. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 This is an application environment diagram of the robot grasping control method in one embodiment;
[0042] Figure 2 This is a flowchart illustrating a robot grasping control method in one embodiment;
[0043] Figure 3 This is a flowchart illustrating the robot grasping control method in another embodiment;
[0044] Figure 4 This is a structural block diagram of a robot grasping control device in one embodiment;
[0045] Figure 5 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0046] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0047] The robot grasping control method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, server 102 communicates with image acquisition device 104 via a network. A data storage system can store the data that server 102 needs to process. The data storage system can be integrated onto server 102, or it can be located in the cloud or on other network servers. Server 102 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. Image acquisition device 104 can include an RGB camera (red, green, and blue camera) and a depth camera. The RGB camera captures light intensity information in the visible light band (red, green, and blue channels) through a sensor to synthesize a color image. The depth camera measures the distance from objects to the camera actively or passively to generate a depth map. Specifically, during the robot grasping control process, server 102 acquires object images and depth images for the grasping area from image acquisition device 104; the object images include object information for each of the multiple grasping objects; three-dimensional point cloud information for the grasping area is generated based on the depth images; the three-dimensional point cloud information includes the object point cloud information for each grasping object, as well as other point cloud information for other objects in the grasping area; for each grasping object, grasping point prediction is performed based on the object point cloud information and object information of the grasping object, resulting in multiple predicted grasping points for the grasping object, as well as grasping point information for each grasping point; for each grasping point, the grasping point information of the grasping point is compared with other point cloud information, and the grasping point whose comparison result meets the comparison conditions is taken as the target grasping point; the grasping operation is performed according to the target grasping point information of the target grasping point.
[0048] In one exemplary embodiment, such as Figure 2 As shown, a robot grasping control method is provided, which can be applied to... Figure 1 Taking server 102 as an example, the explanation includes the following steps S202 to S210. Wherein:
[0049] Step S202: Obtain the object image and depth image captured for the grasping area.
[0050] The object image includes information about each of the multiple grasping objects. The grasping region is the spatial area within which the robot needs to perform the grasping operation, which may include multiple objects to be grasped and environmental obstacles. The object image is a color image captured by an RGB camera, containing visual information such as color and texture of all objects within the grasping region. The depth image is an image captured by a depth camera, where each pixel value represents the distance from that point to the camera, used to describe the spatial position of the object. The grasping object is the specific object that the robot needs to grasp, which may include multiple objects of the same or different types. Object information consists of the visual features of the grasping objects in the object image, such as object category, bounding box, category label, confidence score, mask, color, shape, and texture, used to distinguish different objects.
[0051] Specifically, the server synchronously acquires image data of the grasping area using an RGB camera and a depth camera. The RGB camera outputs a color image (object image) containing object information for multiple grasping objects, such as object A being a red cube with a 60% confidence level, and object B being a blue cylinder with a 70% confidence level. The depth camera outputs a depth image, recording the distance from each point on the object's surface to the camera. The two images need to be calibrated to achieve pixel-level alignment, ensuring that the object information and depth information correspond to the same spatial location. For example, the object image and depth image for the grasping area can be acquired actively or passively received.
[0052] Optionally, the server can directly acquire object images and depth images collected for the grasping area, or it can acquire initial object images and initial depth images collected for the grasping area, perform object recognition on the initial object images to obtain multiple initial object information, perform image analysis on the initial object images based on the multiple initial object information to obtain object images including object information of multiple grasping objects, the image range of the object images is less than or equal to that of the initial object images, and adjust the initial depth images based on the size of the object images to obtain the depth images corresponding to the object images.
[0053] Step S204: Generate 3D point cloud information of the grasping area based on the depth image.
[0054] The 3D point cloud information includes the individual point cloud information of each grasped object, as well as the point cloud information of other objects within the grasping area. The 3D point cloud information is a set of 3D spatial coordinates obtained by transforming a depth image; each point contains (x, y, z) coordinates and possible reflection intensity information. The object point cloud information is a subset of the 3D point cloud belonging to the grasped object, describing its surface spatial structure. Other point cloud information consists of 3D point cloud data of non-grabbed objects (such as obstacles and environmental surfaces) within the grasping area.
[0055] Specifically, when converting a depth image into a 3D point cloud, the server can backproject the depth value of each pixel into (x, y, z) coordinates in the camera coordinate system based on the camera intrinsic parameter matrix, generating an initial point cloud set. Subsequently, a point cloud segmentation algorithm divides the overall point cloud into multiple subsets. Point clouds belonging to the object to be captured are labeled as object point cloud information, while the remaining point clouds are classified as other point cloud information. The segmentation process can leverage semantic features in the object image to improve accuracy, ensuring complete extraction of the point cloud of the captured object.
[0056] Step S206: For each grasped object, grasp point prediction is performed based on the object point cloud information and object information of the grasped object to obtain multiple grasp points predicted for the grasped object, as well as the grasp point information of each grasp point.
[0057] Among them, the gripping point is a candidate position that the robot's end effector can stably grip, calculated based on information such as the object's shape and position. Gripping point information is data describing the attributes of the gripping point; for example, each gripping point is represented by a rotation matrix. Translation vector This indicates the rotation angle and gripping position of the gripping point, respectively.
[0058] Specifically, the server can use the GraspNet model to predict grab points, or it can perform geometric feature analysis on the point cloud of each grab object, calculating parameters such as its principal direction, surface normal, and curvature, while combining shape category information from the object image to predict feasible grab points. For example, for a cube, candidate points might be sampled at locations such as the center of its top face and the midpoint of its edges, while for a cylinder, candidate points distributed along its axial direction are generated on its sides. Each candidate point is associated with grab point information, and finally, the corresponding grab point is output for each grab object.
[0059] Step S208: For each grasping point, compare the grasping point information with other point cloud information, and take the grasping point whose comparison result meets the comparison conditions as the target grasping point.
[0060] The comparison criteria are preset feasibility rules, such as no obstruction point cloud around the grasping point and no occlusion in the grasping direction. In this embodiment, the comparison criteria refer to the fact that the grasping point information is compared with other point cloud information and the result is no collision. The target grasping point is a verified candidate grasping point that can be directly used to perform grasping.
[0061] Specifically, for each grab point, the server needs to verify its actual feasibility. First, the server needs to compare the grab point information with other point cloud information. Based on the comparison results between the grab point information and other point cloud information, it is determined whether the grab point will collide with other objects. If no collision occurs, the comparison result is determined to meet the comparison conditions, and the grab point can be used as the target grab point.
[0062] Optionally, the server can quickly determine spatial overlap by constructing simplified geometric models of the gripping point, obstacles, and robotic arm. Additionally, the server can directly calculate the point cloud density or nearest-point distance within the safe area surrounding the gripping point to determine if a safety threshold is exceeded.
[0063] Step S210: Perform the grabbing operation according to the target grabbing point information.
[0064] Among them, the target grab point information is the complete parameters of the finally selected grab point.
[0065] Specifically, the server can first generate a collision-free path from the current position to the gripping point through a path planning algorithm, control the robotic arm to move along the path, and after the end effector reaches the gripping point, it performs a gripping action in a preset direction and width.
[0066] The aforementioned robot grasping control method first generates 3D point cloud information containing both the grasped object and its environment by fusing object and depth images. This effectively integrates the spatial structure and semantic features of the object, solving the problem of insufficient data information from a single modality. Secondly, it independently predicts multiple candidate grasping points for each grasped object and dynamically compares the grasping point information with the environmental point cloud information. This automatically filters out collision-free and reachable target grasping points, significantly improving the accuracy and feasibility of grasping point prediction. Finally, it executes the grasping operation based on the precise information of the target grasping points, reducing the need for manual intervention and significantly improving the robot's grasping success rate and adaptability in complex scenarios. In summary, the aforementioned robot grasping control method can improve the accuracy of robot grasping control.
[0067] In an exemplary embodiment, acquiring object images and depth images for multiple grasping objects in a grasping area includes: acquiring initial object images and initial depth images for the grasping area; performing object recognition on the initial object images to obtain multiple initial object information; performing image analysis on the initial object images based on the multiple initial object information to obtain an object image including object information of each of the multiple grasping objects; selecting object information from the multiple initial object information; ensuring that the image range of the object image is less than or equal to that of the initial object image; and adjusting the initial depth image based on the size of the object image to obtain a depth image corresponding to the object image.
[0068] The initial object image is the raw image captured of the entire grasping area, containing multiple objects that may become grasping objects. The initial depth image is the corresponding raw image that records the overall object distance information within the grasping area. The initial object information is preliminary information about each object obtained from the initial object image through object recognition.
[0069] Specifically, when the server acquires object images and depth images of multiple objects to be grasped within the grasping area, it first uses appropriate image acquisition devices, such as cameras with image capture capabilities and depth sensors capable of acquiring depth data, to synchronously acquire image and depth information of the entire area where the robot plans to grasp. This yields initial object images and initial depth images. The initial object images visually present the appearance of each object within the grasping area, while the initial depth images record the distance data of each object to the acquisition device. Then, YOLO (You Only Look) is used... Once, the object detection algorithm meticulously identifies each object in the initial object image, deeply analyzing the shape, color, and other features of each object to obtain preliminary information about these objects. This information collectively constitutes the initial object information. Then, object information that meets the grasping requirements is selected from the initial object information. Based on this selected information, the initial object image is further analyzed and processed, removing non-compliant object parts from the image. Finally, an object image containing only detailed information of multiple grasping objects is obtained. The object information is strictly selected from multiple initial object information, and the image range of the object image is smaller than or equal to that of the initial object image. Finally, based on the size of the object image, the initial depth image is cropped or scaled accordingly, so that the adjusted depth image can accurately correspond to the object image, thereby accurately reflecting the depth information of each grasping object in the object image.
[0070] In this embodiment, by first acquiring the overall image and depth information, then performing object recognition and analysis, and finally adjusting the depth image, the image and depth information of the object to be grasped can be accurately obtained, providing an extremely accurate data foundation for subsequent grasping operations and effectively improving the accuracy of grasping.
[0071] In an exemplary embodiment, the initial crawled object information includes object confidence; image analysis is performed on the initial object image based on multiple initial crawled object information to obtain an object image including object information of each of the multiple crawled objects, including: removing initial crawled object information whose object confidence does not meet the confidence condition from the initial object image to obtain candidate object images including object information of each of the multiple candidate crawled objects; and performing center region filtering on the candidate object images based on the center region of the object image to obtain object images including object information of each of the multiple crawled objects.
[0072] Among them, object confidence is an indicator that measures the reliability of identifying objects in the initial object information as crawling objects.
[0073] Specifically, in the process of performing image analysis on the initial object image based on multiple initial crawled object information to obtain an object image containing the object information of each of the multiple crawled objects, a suitable object confidence threshold is first set as a confidence condition. Then, the initial crawled object information is traversed. For each initial crawled object information, its corresponding object confidence is compared with the set threshold. Objects corresponding to information with object confidence below the threshold are decisively removed from the initial object image. After this filtering, the remaining objects constitute candidate crawled objects, and their object information is combined to form a candidate object image. Next, the central region range of the object image is determined. For each candidate crawled object in the candidate object image, the positional relationship between its image range and the central region is analyzed in detail. Through filtering operations, candidate crawled objects that do not meet the requirements are removed from the candidate object image, finally obtaining an object image containing the object information of each of the multiple crawled objects.
[0074] In this embodiment, objects that obviously do not meet the requirements are first removed by confidence screening, and then further filtered by central region filtering. This can more accurately determine the objects to be grabbed, greatly reduce the interference of irrelevant objects on the grabbing operation, and improve the efficiency and accuracy of the determination of the objects to be grabbed.
[0075] In an exemplary embodiment, filtering the candidate object image based on the central region of the object image to obtain an object image including object information of multiple crawling objects includes: for each candidate crawling object, determining the area overlap between the image range of the candidate crawling object and the central region of the object image; filtering the object information of candidate crawling objects whose area overlap does not meet the overlap condition from the candidate object image to obtain an object image including object information of multiple crawling objects whose area overlap meets the overlap condition.
[0076] Candidate objects are those that have undergone preliminary screening and are likely to become the final objects to be crawled. Candidate object information includes various types of information about the candidate objects. Area overlap is the percentage of the area where the image range of a candidate object overlaps with the central region of the object image.
[0077] Specifically, when filtering the initial object image based on its central region to obtain an object image containing information about multiple crawling objects, for each candidate crawling object in the candidate object image, firstly, the area it occupies in the image is determined. Then, the area of the overlapping portion between this area and the central region of the object image is calculated. Next, depending on the specific calculation method, the overlapping area is divided by the area of the candidate crawling object image's range or the area of its central region to obtain the area overlap. Then, a suitable area overlap threshold is set as the overlap condition. The object information of candidate crawling objects whose area overlap does not meet the overlap condition is filtered out from the candidate object image. After this processing, the remaining candidate crawling objects are the final crawling objects, and their object information constitutes the object image.
[0078] In this embodiment, by calculating the area overlap and filtering, it is possible to more accurately select crawling objects located near the central area or that are sufficiently related to the central area, so that the determined crawling objects are more in line with the actual crawling needs, thereby improving the success rate of crawling.
[0079] In an exemplary embodiment, the grasp point information includes grasp point depth; comparing the grasp point information with other point cloud information, and taking the grasp points whose comparison results meet the comparison conditions as target grasp points, includes: comparing the grasp point information with other point cloud information, and taking the grasp points whose comparison results meet the comparison conditions as candidate grasp points; from the candidate grasp points, determining the candidate grasp points whose grasp point depth meets the optimal depth condition as target grasp points.
[0080] The optimal depth condition refers to the most suitable depth condition that the grab point satisfies in the depth direction. In this embodiment, there is only one target grab point. Therefore, only candidate grab points whose depth meets the optimal depth condition can be used as target grab points. For example, the optimal depth condition could be that the range of depth values is most suitable.
[0081] Specifically, when comparing the grab point information with other point cloud information to select the grab point that meets the comparison conditions as the target grab point, the server first needs to perform a comprehensive comparative analysis of the grab point information such as the position and depth with other point cloud information obtained through depth images, and set appropriate comparison conditions to select grab points that meet these conditions as candidate grab points. Then, among the candidate grab points, based on the preset optimal depth conditions, such as the depth being within a specific suitable range, the server selects the candidate grab points whose depth meets the conditions and uses them as the final target grab points.
[0082] In this embodiment, candidate grasping points are first screened by comparing with other point cloud information, and then the target grasping point is determined according to the optimal depth condition. This can select the most suitable one from multiple possible grasping points, effectively improving the stability and reliability of grasping.
[0083] In an exemplary embodiment, the robot grasping control method further includes: obtaining the centroid position of the grasped object to which the target grasping point belongs; adjusting the target grasping point information of the target grasping point based on the centroid position of the object to obtain updated grasping point information; and performing a grasping operation according to the target grasping point information of the target grasping point, including: performing a grasping operation according to the updated grasping point information of the target grasping point.
[0084] Here, the object's centroid position is the location of the center of mass of the grasped object in space. Updating the grasp point information is the new information obtained after adjusting the target grasp point information based on the object's centroid position.
[0085] Specifically, in the robot grasping control method, the server first determines the position of the center of mass of the object to be grasped in space, that is, the position of the object's centroid. Then, based on the relationship between the position of the object's centroid and the position of the target grasping point, and taking into account factors such as the mechanical properties of the object and the stability of grasping, the server makes reasonable adjustments to the position, posture, and other grasping point information of the target grasping point, so that the grasping point is more conducive to the stable grasping of the object, thereby obtaining updated grasping point information. Finally, based on the updated grasping point information, the robot precisely controls the robotic arm and other actuators to move to the corresponding position and performs the grasping operation on the object with a suitable posture.
[0086] In this embodiment, by adjusting the gripping point information based on the object's center of mass position, the gripping operation can be made more in line with the object's mechanical properties, greatly improving the stability and success rate of gripping and effectively reducing the occurrence of situations such as objects slipping during the gripping process.
[0087] In one specific embodiment, a robot grasping control method is also provided. First, the initial object image captured by the image acquisition device is used to identify objects using the YOLO (You Only Look Once) object detection algorithm. The YOLO model can quickly and accurately identify targets in the image and provide the target's location, confidence score, and mask information. To improve detection accuracy and reduce background noise interference, a central ROI (Region of Interest) filtering strategy is adopted, performing target detection only in a specific region at the center of the image. By setting the width and height of the ROI, objects located outside the ROI are filtered out, thereby ensuring the accuracy of target detection.
[0088] The core calculation formula for object detection is based on the YOLO network, assuming a given input image... :
[0089]
[0090] in: This is the detection result returned by the YOLO model, which includes the bounding box, class label, confidence score, and mask. This is the confidence threshold for the YOLO model. If the confidence of an object is lower than this value, no grab point selection will be performed. The Region of Interest (ROI) is a specific part of the image, with dimensions of [size missing]. In image width and height The positioning in is:
[0091]
[0092] By considering only the detection results within the ROI region, irrelevant targets can be effectively filtered out.
[0093] Furthermore, after object detection, the server uses the GraspNet model to predict grab points. The GraspNet model generates 3D point cloud information of the grab region based on the depth image and predicts multiple potential grab points from it. Each grab point is determined by a rotation matrix. Translation vector This indicates that the rotation and position of the grab point are respectively represented.
[0094] In this context, assuming a given 3D point cloud data ,in For the first point cloud For each point, GraspNet outputs a set of rotation matrices and translation vectors for the captured points, denoted as:
[0095]
[0096] Each grab point This corresponds to potential gripping positions on the object's surface. To avoid collisions, this invention uses a collision detection algorithm to filter gripping points. The basic method of collision detection is to compare the predicted gripping points with the spatial positions of the object's surface and other objects to calculate whether a collision exists.
[0097] Specifically, collision detection employs model-independent collision detection methods (such as model-free collision detectors). Assume a grab point is given... The corresponding gripping point posture is:
[0098]
[0099] The detection algorithm calculates whether each grasp point collides with other parts of the object based on the object's point cloud information. If a collision occurs, the grasp point is excluded.
[0100] Subsequently, to adapt to changes in a dynamic environment, this embodiment proposes an optimization method based on dynamic quality threshold adjustment. During the gripping point selection process, the system uses the depth information of the target object (such as depth values in a depth map)... Dynamically adjust the quality threshold. Adjusting the quality threshold can help the system select the most suitable grasping points for the current grasping task and avoid grasping failures caused by unstable posture or inaccurate depth information during the grasping process.
[0101] Here, the assumed depth value of the object To mitigate the impact of noise, the server dynamically adjusts the quality threshold by calculating the percentile of depth values. (Setting the depth percentile...) The server will select all points with depth values greater than that percentile as target crawling points:
[0102]
[0103] in: It is the depth value in the depth map; It is the percentile of the calculated depth value; It is the set depth percentile; by adjusting The system can respond in real time to changes in the posture of objects in a dynamic environment, thereby ensuring the accuracy of the gripping point selection.
[0104] Finally, after selecting the gripping point and performing collision detection, the server optimizes the gripping motion. To ensure the robot can grip in the most suitable posture, the system combines real-time feedback information with depth images and uses a one-click refinement algorithm to fine-tune the gripping points. Specifically, the gripping points... It will perform position adjustments to better match the position and orientation of the target object. The adjustment process optimizes the target object by calculating the position of its center of mass.
[0105] Here, we assume the pose of the current capture point is ,in For position vectors, This is a rotation matrix. The refinement process adjusts the position and rotation of the gripping point to optimally match the centroid position of the target object. .
[0106]
[0107] The refinement algorithm optimizes the position and pose of the grasping point based on the object's depth information and center of mass, ensuring the accuracy and stability of the final grasp. The optimized grasping point is then passed to the robot for execution. The robot then adjusts the target pose accordingly. It executes precise grasping actions. To ensure smooth and safe movements, the system employs a path planning algorithm to limit the robot's range of motion, preventing it from exceeding its movement boundaries.
[0108] Among them, the path executed by the robot Based on the current target pose Plan ahead to ensure the robot can successfully reach the target location and perform the grasping action:
[0109] (current_pose)
[0110] This path planning algorithm combines the robot's current pose and target pose, and ensures the smooth execution of the grasping action by limiting the range of motion.
[0111] In a specific embodiment, such as Figure 3 As shown, a robot grasping control method is also provided, including:
[0112] Step S301: Obtain the initial object image and initial depth image for the grasping area;
[0113] Step S302: Perform object recognition on the initial object image to obtain multiple initial object information;
[0114] The initial crawling information includes object confidence;
[0115] Step S303: Remove the initial crawled object information that does not meet the confidence condition from the initial object image to obtain a candidate object image that includes the object information of each of the multiple candidate crawled objects.
[0116] Step S304: For each candidate object to be captured, determine the area overlap between the image range of the candidate object and the central region of the object image.
[0117] Step S305: Filter the object information of candidate crawling objects whose area overlap does not meet the overlap condition from the candidate object image to obtain an object image that includes the object information of multiple crawling objects whose area overlap meets the overlap condition.
[0118] Among them, the object information is selected from multiple initial object information; the image range of the object image is less than or equal to the initial object image;
[0119] Step S306: Adjust the initial depth image based on the size of the object image to obtain the depth image corresponding to the object image;
[0120] The object image includes object information for each of the multiple objects being captured;
[0121] Step S307: Generate 3D point cloud information of the grasping area based on the depth image; the 3D point cloud information includes the object point cloud information of each grasping object, as well as the other point cloud information of other objects in the grasping area;
[0122] Step S308: For each grasped object, grasp point prediction is performed based on the object point cloud information and object information of the grasped object to obtain multiple grasp points predicted for the grasped object, as well as the grasp point information of each grasp point.
[0123] Step S309: For each grasping point, compare the grasping point information with other point cloud information, and take the grasping points that meet the comparison conditions as candidate grasping points.
[0124] Step S310: From the candidate grab points, determine the candidate grab points whose grab point depth satisfies the optimal depth condition as the target grab point;
[0125] Step S311: Obtain the centroid position of the object to which the target grab point belongs;
[0126] Step S312: Adjust the target grasping point information based on the object's centroid position to obtain updated grasping point information;
[0127] Step S313: Perform the crawling operation according to the updated crawling point information of the target crawling point.
[0128] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0129] Based on the same inventive concept, this application also provides a robot grasping control device for implementing the robot grasping control method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more robot grasping control device embodiments provided below can be found in the limitations of the robot grasping control method described above, and will not be repeated here.
[0130] In one exemplary embodiment, such as Figure 4 As shown, a robot grasping control device 400 is provided, including: an image acquisition module 402, a point cloud information generation module 404, a grasping point prediction module 406, an information comparison module 408, and a grasping operation execution module 410, wherein:
[0131] Image acquisition module 402 is used to acquire object images and depth images for the grasping area; the object image includes object information of multiple grasping objects;
[0132] The point cloud information generation module 404 is used to generate three-dimensional point cloud information of the grasping area based on the depth image; the three-dimensional point cloud information includes the object point cloud information of each grasping object, as well as the other point cloud information of other objects in the grasping area;
[0133] The grasp point prediction module 406 is used to predict grasp points for each grasp object based on the object point cloud information and object information of the grasp object, so as to obtain multiple grasp points predicted for the grasp object and grasp point information of each grasp point.
[0134] The information comparison module 408 is used to compare the capture point information of each capture point with other point cloud information, and to take the capture point whose comparison result meets the comparison conditions as the target capture point.
[0135] The capture operation execution module 410 is used to perform capture operations according to the target capture point information.
[0136] In one exemplary embodiment, the image acquisition module 402 includes:
[0137] The image acquisition unit is used to acquire the initial object image and the initial depth image for the grasping area;
[0138] An object recognition unit is used to perform object recognition on the initial object image to obtain multiple initial object information;
[0139] An image analysis unit is configured to perform image analysis on the initial object image based on the plurality of initial object information to obtain an object image including object information of each of the plurality of grabbing objects; the object information is selected from the plurality of initial object information; the image range of the object image is less than or equal to that of the initial object image;
[0140] An image adjustment unit is used to adjust the initial depth image based on the size of the object image to obtain a depth image corresponding to the object image.
[0141] In one exemplary embodiment, the initial crawling information includes object confidence. In this embodiment, the image analysis unit includes:
[0142] The elimination component is used to remove the initial crawled object information that does not meet the confidence condition from the initial object image, so as to obtain a candidate object image that includes the object information of each of the multiple candidate crawled objects.
[0143] A filtering component is used to perform center region filtering on the candidate object image based on the center region of the object image, so as to obtain an object image including the object information of each of the multiple objects to be captured.
[0144] In one exemplary embodiment, the filtering component is specifically used for:
[0145] For each candidate object to be captured, the area overlap between the image range of the candidate object and the central region of the object image is determined.
[0146] The object information of candidate crawling objects whose area overlap does not meet the overlap condition is filtered out from the candidate object image to obtain an object image that includes the object information of multiple crawling objects whose area overlap meets the overlap condition.
[0147] In one exemplary embodiment, the gripping point information includes gripping point depth. In this embodiment, the information comparison module 408 is specifically used for:
[0148] The grab point information of the grab point is compared with the other point cloud information, and the grab points whose comparison results meet the comparison conditions are selected as candidate grab points.
[0149] From the candidate grab points, the candidate grab points whose grab point depth satisfies the optimal depth condition are determined as the target grab points.
[0150] In an exemplary embodiment, the robot grasping control device 400 further includes a centroid position acquisition module, specifically used for:
[0151] Obtain the centroid position of the object to which the target grab point belongs;
[0152] Based on the centroid position of the object, the target grasping point information of the target grasping point is adjusted to obtain updated grasping point information;
[0153] The capture operation execution module 410 is specifically used for:
[0154] Perform the crawling operation according to the updated crawling point information of the target crawling point.
[0155] Each module in the aforementioned robot grasping control device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0156] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 5As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a robot grasping control method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0157] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0158] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method described above.
[0159] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.
[0160] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the method described above.
[0161] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0162] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0163] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0164] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A robot grasping control method, characterized in that, The method includes: Acquire object images and depth images for the grasping area; the object images include object information for each of the grasping objects; A three-dimensional point cloud information of the grasping area is generated based on the depth image; the three-dimensional point cloud information includes the object point cloud information of each grasping object, as well as other point cloud information of other objects in the grasping area; For each of the aforementioned objects to be crawled, crawling point prediction is performed based on the object point cloud information and object information of the object to be crawled, thereby obtaining multiple crawling points predicted for the object to be crawled, as well as crawling point information for each of the aforementioned crawling points. For each of the aforementioned grasping points, the grasping point information of the grasping point is compared with the other point cloud information, and the grasping points whose comparison results meet the comparison conditions are taken as target grasping points; Perform a crawling operation based on the target crawling point information.
2. The method according to claim 1, characterized in that, The acquisition of object images and depth images for multiple objects in the grasping area includes: Obtain the initial object image and initial depth image for the grasped area; Object recognition is performed on the initial object image to obtain multiple initial object information; Image analysis is performed on the initial object image based on the multiple initial object information to obtain an object image including object information of each of the multiple crawling objects; the object information is selected from the multiple initial object information; the image range of the object image is less than or equal to that of the initial object image; Based on the size of the object image, the initial depth image is adjusted to obtain the depth image corresponding to the object image.
3. The method according to claim 2, characterized in that, The initial crawling information includes object confidence; the step of performing image analysis on the initial object image based on the multiple initial crawling object information to obtain an object image including object information of each of the multiple crawling objects includes: The initial crawling object information that does not meet the confidence condition is removed from the initial object image to obtain a candidate object image that includes the object information of each of the multiple candidate crawling objects. Based on the central region of the object image, the candidate object image is filtered by the central region to obtain an object image that includes the object information of each of the multiple objects to be crawled.
4. The method according to claim 3, characterized in that, The process of filtering the candidate object image based on the central region of the object image to obtain an object image including object information of multiple crawling objects includes: For each candidate object to be captured, the area overlap between the image range of the candidate object and the central region of the object image is determined. The object information of candidate crawling objects whose area overlap does not meet the overlap condition is filtered out from the candidate object image to obtain an object image that includes the object information of multiple crawling objects whose area overlap meets the overlap condition.
5. The method according to claim 1, characterized in that, The grasping point information includes grasping point depth; the step of comparing the grasping point information with other point cloud information, and taking the grasping points that meet the comparison conditions as target grasping points, includes: The grab point information of the grab point is compared with the other point cloud information, and the grab points whose comparison results meet the comparison conditions are selected as candidate grab points. From the candidate grab points, the candidate grab points whose grab point depth satisfies the optimal depth condition are determined as the target grab points.
6. The method according to claim 1, characterized in that, The method further includes: Obtain the centroid position of the object to which the target grab point belongs; Based on the centroid position of the object, the target grasping point information of the target grasping point is adjusted to obtain updated grasping point information; The step of performing the crawling operation according to the target crawling point information includes: Perform the crawling operation according to the updated crawling point information of the target crawling point.
7. A robot grasping control device, characterized in that, The device includes: The image acquisition module is used to acquire object images and depth images of the grasping area; the object images include object information of multiple grasping objects. The point cloud information generation module is used to generate three-dimensional point cloud information of the grasping area based on the depth image; the three-dimensional point cloud information includes the object point cloud information of each of the grasping objects, as well as other point cloud information of other objects in the grasping area; The grasping point prediction module is used to predict grasping points for each grasping object based on the object point cloud information and object information of the grasping object, so as to obtain multiple grasping points predicted for the grasping object and grasping point information of each grasping point. The information comparison module is used to compare the capture point information of each capture point with the other point cloud information, and to take the capture point whose comparison result meets the comparison conditions as the target capture point. The capture operation execution module is used to perform capture operations according to the target capture point information of the target capture point.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.