Method, apparatus, device and humanoid robot for generating target grasping point data
By aligning and deep learning segmentation of the color images and depth images collected by the depth camera, the center of mass and three-dimensional coordinates of the target object are extracted, and the problem of low accuracy of the grab point data in complex dynamic scenarios is solved, and more accurate target point data generation is achieved.
Patent Information
- Application Number
- CN202510329600.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-20
AI Technical Summary
In complex dynamic scenarios such as changes in lighting conditions or partial target objects are blocked, the accuracy of the captured point data detected in the prior art is low.
By obtaining the color images and depth images taken by the depth camera in real time, performing alignment correction processing to obtain a pseudo-color depth map, using pre-trained deep learning segmentation model for detection and segmentation processing, extracting the center of mass coordinates of the target object, and obtaining the three-dimensional coordinates of the target through the depth correction processing, and determining it as the target grab point data.
The accuracy of the grab point data obtained in complex dynamic scenarios is improved, and the problems of light changes, blurred target boundaries, and occlusion are solved, making the geometric feature extraction of the target area more accurate.
Smart Images

Figure CN119850739B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robotics, specifically to the field of target detection technology, and particularly to a method, device, equipment and humanoid robot for generating target grasping point data. Background Art
[0002] With the development of robotics and computer vision, target detection based on depth cameras and generation of grasping point coordinates have been widely used in fields such as industrial automation, service robots, and drones. Among them, the grasping function of a robot for a target object is one of the essential functions of a robot. The key to the grasping action of a target object is to obtain the data of the target grasping point.
[0003] In the prior art, the method for determining the grasping point data is to use the depth image captured by a depth camera combined with target segmentation technology, and use the specific position data of the detected target object as the grasping point data.
[0004] The inventors found that the prior art has at least the following technical problems: in complex dynamic scenarios such as scenes with changing light conditions or partial occlusion of target objects, there is a problem of low accuracy of the detected grasping point data. Summary of the Invention
[0005] Embodiments of this application provide a method, device, equipment and humanoid robot for generating target grasping point data, so as to achieve the effect of improving the accuracy of the grasping point data obtained in complex dynamic scenarios.
[0006] In a first aspect, embodiments of this application provide a method for generating target grasping point data, including:
[0007] Obtaining in real time a color image and a depth image of a target object captured by a photographing device;
[0008] Performing alignment and correction processing on the color image and the depth image to obtain a pseudo-color depth map;
[0009] Inputting the pseudo-color depth map into a pre-trained deep learning segmentation model for detection and segmentation processing to obtain a target segmentation mask map of the target object;
[0010] Performing centroid extraction processing on the target segmentation mask map to determine the centroid coordinates of the target object;
[0011] Performing depth correction processing on the centroid coordinates of the target object, the pseudo-color depth map and a preset internal parameter matrix to obtain target three-dimensional coordinates, and determining the target three-dimensional coordinates as the target grasping point data.
[0012] In a possible implementation, the alignment and correction processing based on the color image and the depth image to obtain a pseudo-color depth map includes: performing alignment processing on the color image and the depth image through a preset depth camera technology library to obtain data to be corrected, where the data to be corrected is color information and depth information at the same pixel position; performing correction processing on the data to be corrected and a preset depth scale factor to obtain a pseudo-color depth map.
[0013] In a possible implementation, the data to be corrected includes the original depth of each pixel point; correspondingly, the calculation formula used when performing correction processing on the data to be corrected and a preset depth scale factor is:
[0014]
[0015] In the formula, is the depth value of the pixel point (x, y) after correction, Z is the original depth of the pixel point (x, y), and S is a preset depth calibration factor.
[0016] In a possible implementation, the centroid extraction processing based on the target segmentation mask map to obtain the centroid coordinates of the target object includes: determining the area of the mask region, the first-order geometric moment about the x-axis in the mask map, and the first-order geometric moment about the y-axis in the mask map according to the target segmentation mask map; determining the centroid x-axis coordinate value according to the area of the mask region and the first-order geometric moment about the x-axis in the mask map; determining the centroid y-axis coordinate value according to the area of the mask region and the first-order geometric moment about the y-axis in the mask map; determining the centroid coordinates of the target object according to the centroid x-axis coordinate value and the centroid y-axis coordinate value.
[0017] In a possible implementation, the depth correction processing based on the centroid coordinates of the target object, the pseudo-color depth map, and a preset internal parameter matrix to obtain the target three-dimensional coordinates includes: determining the centroid depth value of the target according to the centroid coordinates of the target object and the pseudo-color depth map; performing depth correction processing on the centroid coordinates of the target object, the centroid depth value of the target, and the preset internal parameter matrix to obtain the target three-dimensional coordinates.
[0018] In a possible implementation, the calculation formula used when performing depth correction processing on the centroid coordinates of the target object, the centroid depth value of the target, and the preset internal parameter matrix to obtain the target three-dimensional coordinates is:
[0019]
[0020] In the formula, is the target three-dimensional coordinate, is the target centroid depth value, K is the preset internal parameter matrix, cX is the x-axis coordinate value in the centroid coordinates of the target object, and cY is the y-axis coordinate value in the centroid coordinates of the target object.
[0021] In a possible implementation, it further includes: performing coordinate conversion processing on the target grasping point data to obtain the robot grasping coordinates in the robot coordinate system; generating a target grasping path according to the robot grasping coordinates; generating a grasping instruction according to the target grasping path; synthesizing the robot grasping coordinates into the real-time image captured by the imaging device according to the grasping instruction and controlling the actuator to complete the grasping of the target object along the target grasping path.
[0022] In a second aspect, an embodiment of the present application further provides a target grasping point data generation device, including:
[0023] An image acquisition module that acquires a color image and a depth image of a target object captured by an imaging device in real time;
[0024] An image alignment module for performing alignment and correction processing according to the color image and the depth image to obtain a pseudo-color depth map;
[0025] A detection and segmentation module for inputting the pseudo-color depth map into a pre-trained deep learning segmentation model for detection and segmentation processing to obtain a target segmentation mask map of the target object;
[0026] An operation module for performing centroid extraction processing according to the target segmentation mask map to determine the centroid coordinates of the target object;
[0027] A depth correction module for performing depth correction processing according to the centroid coordinates of the target object, the pseudo-color depth map, and the preset internal parameter matrix to obtain target three-dimensional coordinates, and determining the target three-dimensional coordinates as the target grasping point data.
[0028] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor;
[0029] The memory stores computer execution instructions;
[0030] The processor executes the computer execution instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementation manners of the first aspect.
[0031] In a fourth aspect, an embodiment of the present application provides a humanoid robot, including a robot body, an imaging device provided on the robot body, and the electronic device as described in the above third aspect provided in the robot body.
[0032] The embodiments of the present application provide a method, apparatus, device, and humanoid robot for generating target grasping point data. By aligning and correcting the color image and depth image collected by a depth camera, a pseudo-color image is obtained, and then a pre-trained deep learning segmentation model is used to achieve precise detection and segmentation, resulting in a target segmentation mask image. Then, by performing centroid extraction processing on the target segmentation mask image, the problem of low segmentation accuracy in dynamic and complex scenarios such as light changes, blurred target boundaries, and occlusions is solved, making the extraction of geometric features of the target area more accurate in dynamic and complex scenarios, thereby improving the accuracy of the centroid coordinates of the target object. Finally, depth correction processing is performed based on the centroid coordinates of the target object, the pseudo-color depth map, and a preset internal parameter matrix to obtain the target three-dimensional coordinates, and the target three-dimensional coordinates are determined as the target grasping point data, thereby improving the accuracy of the target grasping point data. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.
[0034] Figure 1 It is a schematic diagram of the application scenario of the method for generating target grasping point data provided by the embodiments of the present application;
[0035] Figure 2 It is a schematic flow chart of the method for generating target grasping point data provided by the embodiments of the present application;
[0036] Figure 3 It is a schematic application flow chart of the method for generating target grasping point data provided by the embodiments of the present application;
[0037] Figure 4 It is a schematic structural diagram of the device for generating target grasping point data provided by the embodiments of the present application Figure 1 ;
[0038] Figure 5 It is a schematic structural diagram of the device for generating target grasping point data provided by this embodiment; Figure 2 ;
[0039] Figure 6 It is a schematic structural diagram of the electronic device provided by the embodiments of the present application.
[0040] Through the above accompanying drawings, the clear embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and text descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0042] When the inventors used the methods in the prior art to obtain the grasping point data, they found that the depth map of the target object is generally first captured by a depth camera, which lacks support for accurate segmentation of the target object, resulting in blurry or deviation of the target boundary, thereby reducing the accuracy of the grasping point data obtained. In addition, in complex dynamic scenes such as changing lighting conditions or partial occlusion of the target object, the grasping point data obtained is even more inaccurate.
[0043] Therefore, the inventors proposed the following inventive concept to solve the above-mentioned technical problems: while accurately segmenting the target object, the problem of low target object detection and segmentation accuracy in complex dynamic scenes is solved through depth information correction, centroid calculation of the target mask, and alignment of the depth image and color image, thereby improving the accuracy of the final grasping point data.
[0044] Figure 1 A schematic diagram of an application scenario of the target grab point data generation method provided in an embodiment of the present application, such as Figure 1 As shown, the specific application scenario of the present application includes a humanoid robot 101 and a target object 102 .
[0045] Among them, a depth camera and a chip are installed on the humanoid robot 101, and the chip is connected to the depth camera for communication. The depth camera is used to shoot the target object 102 in real time to obtain a depth map and a color map. Then the chip receives the depth map and color map transmitted by the depth camera, and executes the relevant steps of the target grasping point data generation method for the depth map and the color map, and finally generates the target grasping point data of the humanoid robot 101 for the target object 102.
[0046] In an optional embodiment of the present application, the processor of the depth camera of the humanoid robot can also execute the relevant steps of the target grasping point data generation method, and finally generate the target grasping point data of the humanoid robot 101 for the target object 102, and transmit it to the robot control chip inside the humanoid robot 101.
[0047] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0048] Figure 2 This is a schematic flowchart of the method for generating target grasping point data provided by the embodiment of the present application. The execution subject of the method for generating target grasping point data can be a chip or a processor in the humanoid robot 101 as shown in Figure 1 Figure, or a device related to a computer wirelessly connected to the humanoid robot 101. This embodiment does not limit this.
[0049] As shown in Figure 2 Figure, the method for generating target grasping point data includes:
[0050] S201: Real-time obtain the color image and depth image of the target object captured by the imaging device.
[0051] In this embodiment, the imaging device can be a depth camera. The color image refers to an RBG image, and the depth image refers to an image containing the depth information of each pixel point. The depth information can be the distance between the target object and the camera.
[0052] S202: Perform alignment and correction processing on the color image and the depth image to obtain a pseudo-color depth map.
[0053] In this embodiment, the alignment and correction processing means that all pixel points in the color image and all pixel points in the depth image are in a one-to-one correspondence relationship, ensuring that the color information and depth information at the same pixel position are accurately corresponding.
[0054] Specifically, in an optional embodiment of the present application, step S202 includes:
[0055] S202a: Perform alignment processing on the color image and the depth image through a preset depth camera technology library to obtain data to be corrected, where the data to be corrected is the color information and depth information at the same pixel position.
[0056] In this embodiment, the preset depth camera technology library can be the pyrealsense2 library. The pyrealsense2 library can perform alignment of the depth image and the color image to ensure that the color information and depth information at the same pixel position are accurately corresponding.
[0057] S202b: Perform correction processing on the data to be corrected and a preset depth scale factor to obtain a pseudo-color depth map.
[0058] In this embodiment, the depth value of each corrected pixel point can be obtained by using the data to be corrected, the preset depth scale factor, and the corresponding correction formula, and then the depth value of each corrected pixel point can be corresponding to the corresponding pixel point on the color map to obtain a pseudo-color depth map.
[0059] In an optional embodiment of the present application, the data to be corrected includes the original depth of each pixel point. Correspondingly, the calculation formula used for correction processing according to the data to be corrected and the preset depth ratio factor in step S202b is:
[0060]
[0061] In the formula, is the depth value of the pixel point (x, y) after correction, Z is the original depth of the pixel point (x, y), and S is the preset depth calibration factor.
[0062] In this embodiment, it is assumed that the pixel point in the color image is (x, y), and the original depth of this pixel point in the corresponding depth image is Z. At this time, the corrected depth value can be calculated according to the calculation formula . At this time, the corrected depth value is fused into the position of the pixel point (x, y) in the color image. After all pixel points have gone through the above process, the finally obtained image is a pseudo-color depth map.
[0063] In this embodiment, using the corrected depth value to fuse into the color image can provide a more reliable data basis for subsequent steps to achieve the effect of improving the accuracy of the final target grasping point data.
[0064] S203: Input the pseudo-color depth map into a pre-trained deep learning segmentation model for detection and segmentation processing to obtain a target segmentation mask map of the target object.
[0065] In this embodiment, the pre-trained deep learning segmentation model can be a target segmentation model that is trained using training data for complex dynamic scenarios and then goes through deep learning. For example, the YOLOv8 segmentation model that supports high-precision and target diversity processing.
[0066] In an optional embodiment of the present application, it further includes the training and optimization process of the pre-trained deep learning segmentation model as follows:
[0067] S203a: Obtain a training data set, where the training data set includes pseudo-color depth maps under light change conditions, pseudo-color depth maps under occlusion conditions, and pseudo-color depth maps under complex background conditions.
[0068] S203b: Perform model training on the initial segmentation model according to the training data set to obtain a pre-trained deep learning segmentation model.
[0069] In this embodiment, the pre-trained deep learning segmentation model can perform more accurate detection and segmentation on pseudo-color depth maps under light change conditions, pseudo-color depth maps under occlusion conditions, and pseudo-color depth maps under complex background conditions, obtain the target type, bounding box, and mask, and the generated mask is represented as a binary map.
[0070] In this embodiment, through the efficient inference of the pre-trained deep learning segmentation model optimized and trained on the pseudo-color depth map, more accurate target segmentation can be achieved in a complex dynamic environment to cope with complex dynamic environments such as light changes, occlusions, and backgrounds, and the robustness of target detection can be improved.
[0071] S204: Perform centroid extraction processing according to the target segmentation mask map to determine the centroid coordinates of the target object.
[0072] In this embodiment, the centroid coordinates of the target object can be calculated by performing morphological analysis on each target segmentation mask map and using the image moment method.
[0073] Based on the above embodiment, in an optional embodiment of the present application, step S204 includes:
[0074] S204a: According to the target segmentation mask map, determine the area of the mask region, the first-order geometric moment about the x-axis in the mask map, and the first-order geometric moment about the y-axis in the mask map.
[0075] S204b: Determine the x-axis coordinate value of the centroid according to the area of the mask region and the first-order geometric moment about the x-axis in the mask map.
[0076] S204c: Determine the y-axis coordinate value of the centroid according to the area of the mask region and the first-order geometric moment about the y-axis in the mask map.
[0077] S204d: Determine the centroid coordinates of the target object according to the x-axis coordinate value of the centroid and the y-axis coordinate value of the centroid.
[0078] In this implementation, all pixel points within the image region of the target segmentation mask map are known. Then, the sum of the x coordinates of all pixel points within the image region of the target segmentation mask map can be obtained through image processing software, and the sum of the x coordinates of all pixel points within the image region of the target segmentation mask map is determined as the first-order geometric moment about the x-axis in the mask map. Similarly, the sum of the y coordinates of all pixel points within the image region of the target segmentation mask map can also be obtained, and the sum of the y coordinates of all pixel points within the image region of the target segmentation mask map is determined as the first-order geometric moment about the y-axis in the mask map.
[0079] In this embodiment, the calculation formulas used in steps S204b and S204c are:
[0080]
[0081] In the formula, is the area of the mask region, is the first-order geometric moment about the x-axis in the mask map, $m_{10}$ is the first-order geometric moment about the y-axis in the mask image, $cX$ is the x-axis coordinate value of the centroid, and $cY$ is the y-axis coordinate value of the centroid.
[0082] In this embodiment, the above calculation process can be combined with the optimized performance of OpenCV, enabling the centroid extraction process to be completed within milliseconds and fully meeting the requirements of dynamic scenarios.
[0083] S205: Perform depth correction processing based on the centroid coordinates of the target object, the pseudo-color depth map, and the preset internal parameter matrix to obtain the target three-dimensional coordinates, and determine the target three-dimensional coordinates as the target grasping point data.
[0084] In this embodiment, the depth correction processing can be a process of calculating the target three-dimensional coordinates through the depth correction and back-projection formula, and finally obtaining accurate target three-dimensional coordinates. This process relies on high-precision coordinate mapping and real-time correction functions.
[0085] Specifically, in an optional embodiment of the present application, in step S205, performing depth correction processing based on the centroid coordinates of the target object, the pseudo-color depth map, and the preset internal parameter matrix to obtain the target three-dimensional coordinates includes:
[0086] S205a: Determine the target centroid depth value according to the centroid coordinates of the target object and the pseudo-color depth map.
[0087] In this embodiment, the centroid coordinates of the target object correspond to a fixed pixel point in the pseudo-color depth map. Combining the pseudo-color depth map and the centroid coordinates of the target object can match the depth value of the corresponding pixel point in the pseudo-color depth map, and this depth value is determined as the target centroid depth value. Assume that the centroid coordinates of the target object are (cX, cY), and the depth value corresponding to the centroid coordinates of the target object in the pseudo-color depth map is recorded as D(cX, cY).
[0088] S205b: Perform depth correction processing according to the centroid coordinates of the target object, the target centroid depth value, and the preset internal parameter matrix to obtain the target three-dimensional coordinates.
[0089] Specifically, in an optional embodiment of the present application, when performing depth correction processing in step S205 according to the centroid coordinates of the target object, the target centroid depth value, and the preset internal parameter matrix to obtain the target three-dimensional coordinates, the calculation formula used is:
[0090]
[0091] In the formula, is the target three-dimensional coordinate, is the target centroid depth value, K is the preset internal parameter matrix, cX is the x-axis coordinate value in the centroid coordinates of the target object, and cY is the y-axis coordinate value in the centroid coordinates of the target object.
[0092] Based on the above embodiments, as an alternative embodiment of the present application, the method for generating grasping point data further includes:
[0093] Step A: Perform coordinate transformation processing on the target grasping point data to obtain the robot grasping coordinates in the robot coordinate system.
[0094] Step B: Generate a target grasping path according to the robot grasping coordinates.
[0095] Step C: Generate a grasping instruction according to the target grasping path.
[0096] Step D: Synthesize the robot grasping coordinates into the real-time image captured by the imaging device according to the grasping instruction and control the actuator to complete the grasping of the target object along the target grasping path.
[0097] In this embodiment, Steps A to D refer to generating operation instructions for the grasping points according to the target grasping point data, i.e., the target three-dimensional coordinates, in combination with the working range of the robot end effector and the coordinate transformation rules, for controlling the actuator to complete the grasping of the target object along the target grasping path, where the target grasping path is the movement route of the center coordinate of the actuator to the target three-dimensional coordinates. Finally, the target grasping point data is displayed on the color image and depth image collected by the imaging device for real-time monitoring and debugging. For example: in the form of video synthesis or picture pixel point synthesis, the robot grasping point coordinates are synthesized into the real-time video image captured by the imaging device.
[0098] Figure 3 It is a schematic application flow diagram of the method for generating target grasping point data provided by the embodiments of the present application.
[0099] As Figure 3 shown, this method takes the start step as the first step and successively includes the following steps: First is the initialization process of the hardware and software, that is, setting the RealSense configuration of the imaging device, starting the pipeline for the image transmission of the target object captured in real time, and loading the pre-trained deep learning segmentation model (hereinafter replaced by the YOLOv8 model used in this embodiment). To avoid the data generated during the automatic parameter adjustment stage of the imaging device from affecting the generation efficiency and accuracy of the target grasping point data, the first 30 frames of depth images and color images (hereinafter represented by frame data such as depth and color frames) are skipped. Then, by obtaining the frame data from the imaging device in real time, aligning the depth and color frames, and then converting them into arrays to generate a pseudo-color depth map. Then enter the detection mask step, that is, detecting whether a mask is detected by using the YOLOv8 model to detect the pseudo-color depth map corresponding to each frame of data.
[0100] When a mask is detected, the process enters the step of processing the mask, corresponding to calculating the centroid coordinates of the target object, correcting the depth information corresponding to the centroid, converting it into the three-dimensional coordinates of the target object in the above embodiments, and then generating target grasping point data. And generate a grasping instruction based on the target grasping point data, and then draw the grasping point in the image of the current frame. Then enter the step of displaying the frame, and the displayed frame includes an annotated image and a depth map. Then enter the key detection step to determine whether a user or a humanoid robot automatically presses the corresponding stop key, such as whether the q key is pressed.
[0101] If so, stop the pipeline to cut off the image transmission channel. Then close the window to end the execution of the target grasping point data generation method. On the contrary, if not, return to the step of obtaining frame data.
[0102] When the mask is not detected, directly jump to the step of displaying the frame.
[0103] In summary, the target grasping point data generation method provided by the embodiments of the present application performs alignment and correction processing on the color image and the depth image collected by the depth camera to obtain a pseudo-color image, and then uses a pre-trained deep learning segmentation model to achieve accurate detection and segmentation to obtain a target segmentation mask map. Then, by performing centroid extraction processing on the target segmentation mask map, the problem of low segmentation accuracy in dynamic complex scenarios such as illumination changes, blurred target boundaries, and occlusions is solved, making the geometric feature extraction of the target area more accurate in dynamic complex scenarios, thereby improving the accuracy of the centroid coordinates of the target object. Finally, depth correction processing is performed according to the centroid coordinates of the target object, the pseudo-color depth map, and the preset internal parameter matrix to obtain the target three-dimensional coordinates, and the target three-dimensional coordinates are determined as the target grasping point data, thereby improving the accuracy of the target grasping point data.
[0104] At the same time, the accuracy of the segmentation operation in the dynamic scene is also improved through the pre-trained deep learning segmentation model, and the generation efficiency of the target three-dimensional coordinates is realized more efficiently, so that the method can meet the application requirements of multi-target and complex scenarios.
[0105] At the same time, by performing centroid extraction processing and subsequent steps on the target segmentation mask map, it is ensured that the geometric feature extraction of the target area is more accurate, thereby improving the accuracy of the target grasping point data.
[0106] At the same time, the internal parameter correction and the conversion from pixels to three-dimensional coordinates further improve the calculation accuracy of the target three-dimensional coordinates, and can maintain high robustness even in the case of discontinuous depth value-related data or large noise interference.
[0107] Meanwhile, through efficient alignment processing and centroid extraction processing of the target segmentation mask image, the computational overhead is reduced. Meanwhile, combining the built-in hardware acceleration characteristics of the depth camera and the inference ability of the segmentation model, high-real-time object detection and segmentation are achieved, meeting the real-time requirements of scenarios such as industrial robots and service robots.
[0108] Meanwhile, it can also segment multi-object objects and generate target grasping point data according to the embodiments of the present method, and generate independent target grasping point results based on the depth and target three-dimensional coordinates of each target object. In complex scenarios, it can effectively distinguish target objects and generate corresponding grasping instructions for each target object, providing technical support for the autonomous operation of the robot in multi-object scenarios.
[0109] Figure 4 The structural schematic of the target grasping point data generation device provided by the embodiment of the present application Figure 1 is as Figure 4 shown. The target grasping point data generation device provided in this embodiment includes: an image acquisition module 41, an image alignment module 42, a detection and segmentation module 43, an operation module 44, and a depth correction module 45.
[0110] The image acquisition module 41 acquires a color image and a depth image of the target object captured by the shooting device in real time.
[0111] The image alignment module 42 is used to perform alignment and correction processing on the color image and the depth image to obtain a pseudo-color depth image.
[0112] The detection and segmentation module 43 is used to input the pseudo-color depth image into a pre-trained deep learning segmentation model for detection and segmentation processing to obtain a target segmentation mask image of the target object.
[0113] The operation module 44 is used to perform centroid extraction processing on the target segmentation mask image to determine the centroid coordinates of the target object.
[0114] The depth correction module 45 is used to perform depth correction processing according to the centroid coordinates of the target object, the pseudo-color depth image, and a preset internal parameter matrix to obtain the target three-dimensional coordinates, and determine the target three-dimensional coordinates as the target grasping point data.
[0115] In a possible implementation manner, the image alignment module 42 is specifically used to: perform alignment processing on the color image and the depth image through a preset depth camera technology library to obtain data to be corrected, where the data to be corrected is the color information and depth information at the same pixel position; perform correction processing on the data to be corrected and a preset depth ratio factor to obtain a pseudo-color depth image.
[0116] In a possible implementation, the data to be corrected includes the original depth of each pixel point; correspondingly, the image alignment module 42 is specifically configured to: The calculation formula used for correction processing according to the data to be corrected and the preset depth scaling factor is:
[0117]
[0118] In the formula, is the depth value of the pixel point (x, y) after correction, Z is the original depth of the pixel point (x, y), and S is the preset depth calibration factor.
[0119] In a possible implementation, the operation module 44 is specifically configured to: determine the mask area, the first-order geometric moment about the x-axis in the mask map, and the first-order geometric moment about the y-axis in the mask map according to the target segmentation mask map; determine the x-axis coordinate value of the centroid according to the mask area and the first-order geometric moment about the x-axis in the mask map; determine the y-axis coordinate value of the centroid according to the mask area and the first-order geometric moment about the y-axis in the mask map; determine the centroid coordinates of the target object according to the x-axis coordinate value of the centroid and the y-axis coordinate value of the centroid.
[0120] In an alternative embodiment of the present application, the depth correction module 45 is specifically configured to: determine the centroid depth value of the target according to the centroid coordinates of the target object and the pseudo-color depth map; perform depth correction processing according to the centroid coordinates of the target object, the centroid depth value of the target, and the preset internal parameter matrix to obtain the target three-dimensional coordinates.
[0121] In an alternative embodiment of the present application, the depth correction module 45 is specifically configured to: The calculation formula used for performing depth correction processing according to the centroid coordinates of the target object, the centroid depth value of the target, and the preset internal parameter matrix to obtain the target three-dimensional coordinates is:
[0122]
[0123] In the formula, is the target three-dimensional coordinate, is the centroid depth value of the target, K is the preset internal parameter matrix, cX is the x-axis coordinate value in the centroid coordinates of the target object, and cY is the y-axis coordinate value in the centroid coordinates of the target object.
[0124] Figure 5 This is the structural schematic of the target grasping point data generation device provided in this embodiment Figure 2 .
[0125] Such as Figure 5As shown, the target grasping point generation device further includes a display execution module 46. In an optional embodiment of the present application, the display execution module 46 is specifically configured to: perform coordinate conversion processing on the target grasping point data to obtain the robot grasping coordinates in the robot coordinate system; generate a target grasping path according to the robot grasping coordinates; generate a grasping instruction according to the target grasping path; synthesize the robot grasping coordinates into the real-time image captured by the imaging device according to the grasping instruction and control the actuator to complete the grasping of the target object along the target grasping path.
[0126] The target grasping point data generation device provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.
[0127] Figure 6 It is a schematic structural diagram of an electronic device provided in an embodiment of the present application. As Figure 6 shown, the electronic device provided in this embodiment includes: at least one processor 601 and a memory 602. Optionally, the device further includes a communication component 603. Among them, the processor 601, the memory 602, and the communication component 603 are connected through a bus 604.
[0128] In the specific implementation process, at least one processor 601 executes the computer execution instructions stored in the memory 602, so that at least one processor 601 executes the above method.
[0129] The specific implementation process of the processor 601 can refer to the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.
[0130] An embodiment of the present application further provides a humanoid robot, including a robot body, an imaging device provided on the vehicle body, and an electronic device as described in the above embodiment provided in the robot body.
[0131] In this embodiment, both being provided on the robot body and being provided within the robot body can be achieved through a hardware connection structure such as bolt connection, screw connection, etc.
[0132] An embodiment of the present application further provides a computer-readable storage medium, in which computer execution instructions are stored. When the processor executes the computer execution instructions, the above target grasping point data generation method is implemented.
[0133] An embodiment of the present application further provides a computer program product, including a computer program, which implements the above target grasping point data generation method when executed by the processor.
[0134] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU for short), or may also be other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly implemented by the execution of the hardware processor, or can be implemented by the combination of the hardware and software modules in the processor.
[0135] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.
[0136] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.
[0137] This application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.
[0138] This application also provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the processor executes the computer-executable instructions, the above method is implemented.
[0139] The above-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk or an optical disk. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.
[0140] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in a device.
[0141] The division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Additionally, the couplings or direct couplings or communication connections shown or discussed between each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0142] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0143] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0144] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0145] Those of ordinary skill in the art will understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical disks that can store program codes.
[0146] Finally, it should be noted that: after considering the specification and practicing the invention disclosed herein, those skilled in the art will readily think of other embodiments of the present invention. The present invention is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include known common knowledge or conventional technical means in the technical field not disclosed by the present invention. It is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.
Claims
1. A method for generating target grab point data, characterized in that: include: Acquire the color image and depth image of the target object taken by the shooting device in real time; Performing alignment processing on the color image and the depth image through a preset depth camera technology library to obtain data to be corrected, wherein the data to be corrected is color information and depth information at the same pixel position; Performing correction processing according to the data to be corrected and a preset depth scale factor to obtain a pseudo-color depth map; Inputting the pseudo-color depth map into a pre-trained deep learning segmentation model for detection and segmentation processing to obtain a target segmentation mask map of the target object; Performing centroid extraction processing according to the target segmentation mask image to determine the centroid coordinates of the target object; Determine the target object centroid depth value according to the target object centroid coordinates and the pseudo-color depth map; Depth correction processing is performed according to the target object's centroid coordinates, the target centroid depth value and a preset internal parameter matrix to obtain the target three-dimensional coordinates, and the target three-dimensional coordinates are determined as the target grabbing point data.
2. The method according to claim 1, characterized in that The data to be corrected includes the original depth of each pixel; Accordingly, the calculation formula used when performing the correction process according to the data to be corrected and the preset depth scale factor is: In the formula, is the corrected depth value of the pixel point (x, y), Z is the original depth of the pixel point (x, y), and S is the preset depth calibration factor.
3. The method according to claim 1, characterized in that The centroid extraction process is performed according to the target segmentation mask image to obtain the centroid coordinates of the target object, including: According to the target segmentation mask image, determine the area of the mask region, the first-order geometric distance about the x-axis in the mask image, and the first-order geometric distance about the y-axis in the mask image; Determine the x-axis coordinate value of the centroid according to the area of the mask region and the first-order geometric distance about the x-axis in the mask image; Determine the y-axis coordinate value of the centroid according to the area of the mask region and the first-order geometric distance about the y-axis in the mask image; The center of mass coordinates of the target object are determined based on the center of mass x-axis coordinate value and the center of mass y-axis coordinate value.
4. The method according to claim 1, characterized in that: The calculation formula used for performing depth correction processing according to the target object centroid coordinates, the target centroid depth value and the preset internal parameter matrix to obtain the target three-dimensional coordinates is: In the formula, is the three-dimensional coordinate of the target, is the target centroid depth value, K is the preset internal parameter matrix, cX is the x-axis coordinate value in the target object centroid coordinates, and cY is the y-axis coordinate value in the target object centroid coordinates.
5. The method according to any one of claims 1 to 4, characterized in that Also includes: Performing coordinate conversion processing on the target grasping point data to obtain the robot grasping coordinates in the robot coordinate system; Generate a target grasping path according to the robot grasping coordinates; Generate a grabbing instruction according to the target grabbing path; According to the grasping instruction, the robot grasping coordinates are synthesized into the real-time picture taken by the shooting device and the actuator is controlled to complete the grasping of the target object according to the target grasping path.
6. A target grab point data generating device, characterized in that: include: An image acquisition module acquires the color image and depth image of the target object captured by the shooting device in real time; An image alignment module is used to perform alignment processing according to the color image and the depth image through a preset depth camera technology library to obtain data to be corrected, wherein the data to be corrected is color information and depth information at the same pixel position; perform correction processing according to the data to be corrected and a preset depth scale factor to obtain a pseudo-color depth map; A detection and segmentation module is used to input the pseudo-color depth map into a pre-trained deep learning segmentation model for detection and segmentation processing to obtain a target segmentation mask map of the target object; A calculation module, used for performing centroid extraction processing according to the target segmentation mask image to determine the centroid coordinates of the target object; The depth correction module is used to determine the target center of mass depth value according to the target object center of mass coordinates and the pseudo-color depth map; perform depth correction processing according to the target object center of mass coordinates, the target center of mass depth value and a preset internal parameter matrix to obtain the target three-dimensional coordinates, and determine the target three-dimensional coordinates as the target grab point data.
7. An electronic device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 5.
8. A humanoid robot, characterized in that: The invention comprises a robot body, a photographing device arranged on the robot body, and the electronic device as claimed in claim 7 arranged in the robot body.
Citation Information
Patent Citations
Pose determination method and device, equipment, storage medium and product
CN119048601A
Battery replacement robot navigation pose measurement method based on three-dimensional visual features
CN119594979A
Cited By
Industrial material package automatic feeding system and method based on depth map key point 3D positioning
CN121717145A