Humanoid robot control method based on image segmentation

By employing instance segmentation and point cloud adjustments based on depth values, the method improves the accuracy of target object position calculation and grasping success rates in complex scenes for human-like robots.

CN120307300APending Publication Date: 2025-07-15人形机器人(上海)有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510697352.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In complex scenarios, when a humanoid robot grabs an object in the prior art, the boundary segmentation accuracy is poor, resulting in low accuracy in computing the position information of the target object and the grab task may fail.

Method used

The image information of the target object is segmented based on the instance segmentation algorithm, point cloud information of the target area is obtained, direction and position information of the target object are determined through principal component analysis, grab control information is generated, and the terminal executor is controlled to capture the target object.

Benefits of technology

It improves the calculation accuracy of target object position information in complex scenarios, increases the crawling success rate, and reduces the amount of data processed by point cloud information, and improves data processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120307300A_ABST
    Figure CN120307300A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a humanoid robot control method based on image segmentation. Belongs to the technical field of robots. The method comprises the following steps: acquiring image information about a target object; segmenting the image information of the target object based on an instance segmentation algorithm to obtain a target area corresponding to the target object; obtaining point cloud information corresponding to the target area according to the image information; according to coordinate information in the point cloud information, position information of the target object is obtained through calculation; performing principal component analysis on the point cloud information to obtain direction information of the target object; determining grabbing control information of the end effector according to the position information and the direction information of the target object; and controlling the end effector to grab the target object according to the grabbing control information. According to the method, the calculation accuracy of the position information of the target object in a complex scene can be improved, so that the capture success rate in the scene can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robot technology, and particularly to a control method for a humanoid robot based on image segmentation. Background Art

[0002] In the process of the evolution of robot technology towards intelligence and anthropomorphism, humanoid robots have become one of the core tracks for industry innovation and development due to their unique advantages of being highly adaptable to human living and working scenarios.

[0003] Currently, in the application scenario of a humanoid robot grasping an object, by extracting and analyzing the feature of the image data collected by a vision sensor, the position information of the target object in space is identified, and then the object is grasped. However, this method has the problem of poor boundary segmentation accuracy. Especially in complex scenarios, such as scenarios where the object is partially occluded, there are shadows, or texture is missing, the occluded part may not be recognized, resulting in poor calculation accuracy of the position information of the target object, and the grasping task may fail to execute. Summary of the Invention

[0004] An embodiment of this application provides a control method for a humanoid robot based on image segmentation, which can improve the calculation accuracy of the position information of the target object in a complex scenario, thereby facilitating the improvement of the grasping success rate in this scenario.

[0005] In a first aspect, an embodiment of this application provides a control method for a humanoid robot based on image segmentation. The humanoid robot includes an end effector. The method includes: obtaining image information about a target object; segmenting the image information of the target object based on an instance segmentation algorithm to obtain a target area corresponding to the target object; obtaining point cloud information corresponding to the target area according to the image information corresponding to the target area; calculating the position information of the target object according to the coordinate information in the point cloud information; performing principal component analysis on the point cloud information to obtain the direction information of the target object; determining the grasping control information of the end effector according to the position information and direction information of the target object; and controlling the end effector to grasp the target object according to the grasping control information.

[0006] In a possible implementation manner, segmenting the image information of the target object based on an instance segmentation algorithm to obtain a target area corresponding to the target object includes: associating the target object in different frame images based on a temporal tracking algorithm; and segmenting the associated target object between different frames based on the instance segmentation algorithm to obtain a target area corresponding to the target object.

[0007] In a possible implementation, based on the coordinate information in the point cloud information, the position information of the target object is calculated, including: obtaining the offset between the instance mask output by the instance segmentation algorithm and the true boundary of the target object; in the case where the offset is greater than a preset distance threshold, determining the point coordinate adjustment weight of the corresponding area of the point cloud according to the pixel depth value of the instance mask; wherein, the point coordinate adjustment weight is negatively correlated with the pixel depth value of the instance mask; adjusting the coordinates in the point cloud information according to the point coordinate adjustment weight of the corresponding area of the point cloud to obtain the adjusted point cloud coordinate information; and calculating the position information of the target object according to the adjusted point cloud coordinate information.

[0008] In a possible implementation, based on the position information and orientation information of the target object, the grasping control information of the end effector is determined, including: determining the grasping contact area according to the position information of the target object; determining the mask area corresponding to the grasping contact area according to the instance mask output by the instance segmentation algorithm; determining the clamping force of the end effector according to the mask area corresponding to the grasping contact area, and the grasping control information includes the clamping force; wherein, the clamping force is positively correlated with the mask area corresponding to the grasping contact area.

[0009] In a possible implementation, based on the coordinate information in the point cloud information, the position information of the target object is calculated, including: obtaining the weight distribution information corresponding to the target object; the weight distribution information is uniform weight distribution or non-uniform weight distribution; if the weight distribution information corresponding to the target object is uniform weight distribution, then based on the point cloud information corresponding to all parts included in the target object, calculating the position information of the target object; if the weight distribution information corresponding to the target object is non-uniform weight distribution, then determining the weight concentration area corresponding to the target object; and calculating the position information of the target object according to the coordinates of the points located in the weight concentration area.

[0010] In a possible implementation, determining the weight concentration area corresponding to the target object; calculating the position information of the target object according to the coordinates of the points located in the weight concentration area, including: obtaining the material distribution condition of the target object; determining the weight concentration area corresponding to the target object according to the material distribution condition of the target object; the weight concentration area is the largest continuous area corresponding to a single material inside the target object; extracting the point cloud corresponding to the weight concentration area from the point cloud information; and taking the average value of the coordinates of all points in the point cloud corresponding to the weight concentration area as the position information of the target object.

[0011] In a possible implementation manner, based on the coordinate information in the point cloud information, the position information of the target object is calculated, including: judging, according to the image information, whether the target object is composed of at least two non-overlapping parts, and at least a part of the non-overlapping parts does not have a grasping position; if the target object is composed of at least two non-overlapping parts, and at least a part of the non-overlapping parts does not have a grasping position, then extracting the point cloud information including the area corresponding to the grasping position from the point cloud information corresponding to the target area; and calculating the position information of the target object according to the coordinate information in the point cloud information including the area corresponding to the grasping position.

[0012] In a possible implementation manner, based on the point cloud information including the area corresponding to the grasping position, the position information of the target object is calculated, including: taking the average value of all the point coordinates in the point cloud of the area corresponding to the grasping position as the position information of the target object.

[0013] In a possible implementation manner, the position information includes a center point; according to the position information and the orientation information of the target object, the grasping control information of the end effector is determined, including: determining the grasping contact area of the end effector according to the position information of the target object; the grasping contact area can wrap the center point; determining the grasping posture of the end effector according to the orientation information of the target object; and generating the grasping control information according to the grasping contact area and the grasping posture.

[0014] In a possible implementation manner, the instance segmentation algorithm is the YOLACT algorithm.

[0015] In a second aspect, an embodiment of the present application provides an electronic device, including: a memory and a processor;

[0016] The memory stores computer-executable instructions;

[0017] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementation manners of the first aspect.

[0018] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the above first aspect and / or various possible implementation manners of the first aspect.

[0019] In a fourth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the above first aspect and / or various possible implementation manners of the first aspect.

[0020] A humanoid robot control method based on image segmentation provided by an embodiment of the present application segments the image information of a target object based on an instance segmentation algorithm to obtain a target region of the target object, which can improve the calculation accuracy of the position information of the target object in a complex scene, thereby facilitating the improvement of the grasping success rate in the complex scene; and since only the point cloud information corresponding to the target region needs to be processed, the data volume of the point cloud information to be processed is reduced, and the data processing efficiency is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0022] Figure 1 It is a schematic diagram of the scenario of the humanoid robot control method based on image segmentation provided by the present application;

[0023] Figure 2 It is a schematic flow diagram of the humanoid robot control method based on image segmentation provided by an embodiment of the present application;

[0024] Figure 3 It is a schematic flow diagram of the target region segmentation provided by an embodiment of the present application;

[0025] Figure 4 It is a schematic flow diagram of the position information determination provided by an embodiment of the present application;

[0026] Figure 5 It is a schematic flow diagram of the generation of grasping control information provided by an embodiment of the present application;

[0027] Figure 6 It is a schematic flow diagram of the point cloud interception provided by an embodiment of the present application;

[0028] Figure 7 It is a schematic diagram of the structure of the electronic device provided by an embodiment of the present application.

[0029] Through the above-mentioned accompanying drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0031] In the field of humanoid robots, the target detection algorithm is the mainstream technology for object localization. By analyzing the image data of the vision sensor, the position of the object is obtained, and then the object is grasped. However, the current object localization process requires a large amount of data to be processed, consumes a large amount of computing resources, and has a low data processing efficiency.

[0032] In view of the above technical problems, the inventors propose the following technical concept: obtaining the target point cloud information corresponding to the target object, performing instance segmentation on the target point cloud information to obtain the target area corresponding to the target object, using the point cloud information in the target area to determine the position information and orientation information of the target object, determining the grasping control information according to the position information and orientation information, and using the grasping control information to control the end effector to grasp the target object.

[0033] Figure 1 It is a schematic diagram of the scenario of the humanoid robot control method based on image segmentation provided by the present application. As Figure 1 , in this scenario, it includes: a humanoid robot 10 and a target object 20 to be grasped.

[0034] Among them, the humanoid robot 10 includes an image acquisition device 101, at least one end effector 102, and a control unit 103. The image acquisition device 101 may include a depth camera, a vision sensor, etc. The end effector 102 may be a parallel gripper, a vertical gripper, a finger gripper, a dexterous hand, etc. The target object 20 includes an object to be grasped.

[0035] The control unit 103 is configured to obtain the image information of the current scene collected by the image acquisition device 101. The image information obtained by the image acquisition device may include a two-dimensional RGB image and a three-dimensional depth map. The depth map is the distance information of each pixel point. Then, the depth map can be directly converted into point cloud information based on the built-in SDK or hardware processing module of the image acquisition device. Determine the position and orientation of the target object 20 according to the point cloud information, and control the end effector 102 to grasp the target object 20 according to the position and orientation.

[0036] It can be understood that the scenarios illustrated in the embodiments of the present application do not constitute specific limitations on the method for controlling a humanoid robot based on image segmentation. In other feasible embodiments of the present application, the above scenarios may include more or fewer components than those shown in the figures, or combine certain components, or split certain components, or have different component arrangements, which can be specifically determined according to the actual application scenarios and are not limited herein. Figure 1 The illustrated scenarios can be implemented by hardware, software, or a combination of software and hardware.

[0037] The following uses specific embodiments to elaborate in detail on the technical solutions of the present application and how the technical solutions of the present application solve the above technical problems. These several specific embodiments below can be combined with each other, and for the same or similar concepts or processes, they may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0038] Figure 2 It is a flowchart of the method for controlling a humanoid robot based on image segmentation provided by the embodiments of the present application. The execution subject of the embodiments of the present application can be Figure 1 the entire humanoid robot or the control unit in the humanoid robot. As Figure 2 shown, the method includes: step S201 to step S207.

[0039] S201: Obtain image information about the target object.

[0040] In this step, an image acquisition device can be used to periodically or continuously capture the target object to obtain image information.

[0041] Among them, the image information can include one image or N images, or a new image obtained by combining N images, where N is a positive integer.

[0042] S202: Segment the image information of the target object based on the instance segmentation algorithm to obtain the target region corresponding to the target object.

[0043] In this step, the instance segmentation algorithm can include using the Mask R-CNN model, the SOLO (Single-stage instance segmentation) model, and the YOLOCAT algorithm. By segmenting the image information of the target object through the instance segmentation algorithm, each independent object in the point cloud can be distinguished, and thus the category of the target object in the image information and the target region corresponding to the target object can be obtained. The target region corresponding to the target object is usually calibrated in the form of extracting the precise boundary of the target object.

[0044] S203: Obtain the point cloud information corresponding to the target area according to the image information corresponding to the target area.

[0045] In this step, the point cloud information corresponding to the target area in the image information can be directly read.

[0046] Among them, the point cloud is a data set composed of a large number of discrete points in three-dimensional space. The point cloud information includes multiple points and the coordinate information corresponding to each point, and may also include other information (such as color, intensity, normal vector, etc.).

[0047] S204: Calculate the position information of the target object according to the coordinate information in the point cloud information corresponding to the target area.

[0048] In this step, it can be directly calculating the average value of the coordinates of all points in the point cloud information corresponding to the target area as the center point of the target object, that is, as the position information of the target object.

[0049] S205: Perform principal component analysis on the point cloud information to obtain the direction information of the target object.

[0050] In this step, the process of performing principal component analysis on the point cloud information may include:

[0051] First, take the average value of the coordinates of all points in the target point cloud information as the coordinates of the centroid, and the centroid coordinates are expressed as , represents the x-axis coordinate of the centroid, represents the y-axis coordinate corresponding to the centroid, represents the z-axis coordinate corresponding to the centroid. That is, the centroid of the target point cloud is: , and the centroid coordinates are used as the position information of the target object, where N is the total number of points in the target point cloud information, is the three-dimensional coordinate of the i-th point. Among them, represents the x-axis coordinate corresponding to the i-th point, represents the y-axis coordinate corresponding to the i-th point, represents the z-axis coordinate corresponding to the i-th point. Then calculate the deviation of the coordinates of each point in the point cloud relative to the centroid . The deviation means subtracting the coordinate values corresponding to each direction in the three directions respectively. And construct the covariance matrix C of the point cloud data:

[0052]

[0053] Among them, represents the matrix formed by converting the coordinates corresponding to the i-th point. This matrix is one column with three rows, that is, the size is 3×1. For example, if If it is (1, 2, 3), then the corresponding transformed matrix is . represents the matrix formed by the conversion of the centroid coordinates . represents the corresponding matrix transpose. C is a 3×3 covariance matrix. The eigenvalue decomposition of the covariance matrix C gives the eigenvalues and the corresponding eigenvectors . and are in one-to-one correspondence. The larger the eigenvalue, the more important the direction corresponding to the eigenvector. Select the first eigenvector corresponding to the largest eigenvalue among the three eigenvalues as the main direction (major axis) of the object, and the second eigenvector corresponding to the second largest eigenvalue as the secondary direction (minor axis). Then, based on the right-hand rule of vectors, the vector product of the first eigenvector and the second eigenvector gives the third eigenvector , that is , represents the first eigenvector, represents the second eigenvector. Among them, is one of, is one of.

[0054] S206: Determine the grasping control information of the end effector according to the position information and orientation information of the target object.

[0055] In this step, the coordinates and orientation information (pitch angle, azimuth angle) of the target object are written into the grasping control information template to obtain the grasping control information. It also includes using the position information and orientation information to determine the jaw opening degree, grasping point position, posture, moving distance, moving direction, moving trajectory, rotation angle, etc. of the end effector, and writing these parameters into the grasping control information template to obtain the grasping control information.

[0056] The center point of the above target object is located inside the area enclosed by the grasping posture of the end effector.

[0057] Specifically, after obtaining the position information and orientation information of the target object, the grasping strategy and the generation of grasping points can be determined. Among them, the grasping strategy can be determined based on the task requirements and the object attributes. The finger opening amplitude and grasping force can be determined based on the object size, weight, material, etc. Regarding the grasping point, the grasping point can be generated based on the offline generated grasping template, or the grasping point pose can be calculated in real time based on the neural network or the optimization algorithm.

[0058] S207: Control the end effector to grasp the target object according to the grasping control information.

[0059] In this step, it includes sending the grasping control information to the end effector to enable the end effector to grasp the target object, or sending the grasping control information to the controller of the end effector to enable the controller of the end effector to control the end effector to grasp the target object.

[0060] From the description of the above embodiments, it can be seen that the embodiments of the present disclosure segment the image information of the target object based on the instance segmentation algorithm to obtain the target area of the target object, which can improve the calculation accuracy of the position information of the target object in complex scenarios, thereby facilitating the improvement of the grasping success rate in such complex scenarios; and since only the point cloud information corresponding to the target area needs to be processed, the data volume of the point cloud information to be processed is reduced, and the data processing efficiency is improved.

[0061] Figure 3 It is a schematic diagram of the target area segmentation process provided by the embodiment of the present application. As Figure 3 shown, in some optional embodiments, on the basis of any embodiment, step S202 includes: step S2021 and step S2022.

[0062] S2021, associate the target object for different frame images based on the temporal tracking algorithm.

[0063] S2022, segment the associated target object between different frames based on the instance segmentation algorithm to obtain the target area corresponding to the target object.

[0064] Exemplarily, the above temporal tracking algorithm can be DeepSORT, for example. This embodiment can realize the association of objects between different images, avoid the generation of segmentation jitter, realize temporal consistency, improve the accuracy of the position information of the finally determined target object, and thus facilitate the improvement of the success rate of the humanoid robot grasping an object.

[0065] In some optional embodiments, on the basis of any embodiment, step S202 includes:

[0066] Detect the connectivity of the instance mask output by the instance segmentation algorithm. If a non-connected situation is found, control the side-view camera of the humanoid robot to collect the second image of the target object. Fuse the instance mask results obtained by segmenting the two images as the final output result of the instance segmentation algorithm.

[0067] This embodiment can solve the problems of inaccurate mask caused by missing areas, inaccurate calculation results of the position information of the target object, and grasping failure, and ensure the grasping success rate and grasping stability.

[0068] Figure 4Schematic diagram of the location information determination process provided by the embodiments of the present application. As Figure 4 shown, in some alternative embodiments, based on any one of the embodiments, step S204 includes: step S2041 to step S2044.

[0069] S2041, obtain the offset between the instance mask output by the instance segmentation algorithm and the true boundary of the target object.

[0070] S2042, when the offset is greater than the preset distance threshold, determine the point coordinate adjustment weight of the corresponding area of the point cloud according to the pixel depth value of the instance mask. Among them, the point coordinate adjustment weight is negatively correlated with the pixel depth value of the instance mask. That is, the larger the pixel depth value of the instance mask, the smaller the point coordinate adjustment weight. The smaller the pixel depth value of the instance mask, the larger the point coordinate adjustment weight.

[0071] S2043, adjust the coordinates in the point cloud information according to the point coordinate adjustment weight of the corresponding area of the point cloud to obtain the adjusted point cloud coordinate information. That is, multiply the coordinates in the point cloud information by the point coordinate adjustment weight to obtain the adjusted point cloud coordinate information.

[0072] S2044, calculate the location information of the target object according to the adjusted point cloud coordinate information.

[0073] In the above step S2041, the comparison can be based on the coordinate offset between the mask edge and the true contour, or on the offset between the center point of the instance mask and the center point of the true boundary of the target object. Based on the above embodiments, when the depth value is large, that is, when the object surface is far from the camera, the influence on the centroid calculation is reduced after weighting; when the depth value is small, that is, when the object surface is close to the camera, it is possible that part of the object surface segmentation is missed, and the missing area can be partially compensated after weighting. The boundary offset problem is solved, and the problem that the grasping point deviates from the object centroid and causes slipping can be avoided, ensuring the grasping success rate and grasping stability.

[0074] Figure 5 Schematic diagram of the grasping control information generation process provided by the embodiments of the present application. In some alternative embodiments, based on any one of the embodiments, step S206 includes: step S2061 to step S2063.

[0075] S2061, determine the grasping contact area according to the location information of the target object. Specifically, the location information of the target object can be used as the input parameter of the neural network model, and several grasping points are calculated. The continuous area on the object surface corresponding to the arc surface formed by connecting these grasping points is the above-mentioned grasping contact area.

[0076] S2062. Determine the mask area corresponding to the grasping contact area according to the instance mask output by the instance segmentation algorithm. That is, the area of the mask region corresponding to the above-mentioned grasping contact area in the instance mask is the mask area.

[0077] S2063. Determine the clamping force of the end effector according to the mask area corresponding to the grasping contact area. The grasping control information includes the clamping force. Among them, the clamping force is positively correlated with the mask area corresponding to the grasping contact area. That is, the smaller the mask area, the smaller the clamping force. The larger the mask area, the larger the clamping force.

[0078] In this embodiment, for example, for a smaller mask area, a compliant grasping mode can be adopted, and the corresponding clamping force is smaller. For a larger mask area, a high-stiffness grasping mode can be adopted, and the corresponding clamping force is larger. This embodiment can solve the problem that the contact surface of the gripper is insufficient due to the missing of some areas during the segmentation process. The clamping force can be adjusted according to the mask area, so that the grasping force matches the contact area, ensuring the success rate and stability of the humanoid robot in grasping objects.

[0079] In a possible implementation manner, in step S204, according to the coordinate information in the point cloud information, the position information of the target object is calculated, including: steps S204B1 to S204B3.

[0080] S204B1: Obtain the weight distribution information corresponding to the target object. The weight distribution information is evenly distributed or unevenly distributed.

[0081] In this step, it may include determining the weight distribution information corresponding to the target object according to the density of points in the point cloud information. It may also include dividing the space where the point cloud is located into multiple voxels (pixels in three-dimensional space), counting the number of points in each voxel as the density of points in the voxel, calculating the mean and variance of the density of points, and determining whether the weight distribution is uniform or not according to the mean and variance. It is also possible to determine whether the above weight distribution is uniform according to the ray device or image information.

[0082] For example, if the density of some points in the point cloud information is concentrated and the density of some points is dispersed, it is determined that the weight distribution is uneven. Another example is that if the variance of the density is less than the preset variance threshold, it is determined that the weight distribution is uniform. If the variance of the density is greater than or equal to the preset variance threshold, it is determined that the weight distribution is uneven.

[0083] S204B2: If the weight distribution information corresponding to the target object is evenly distributed, calculate the position information of the target object based on the point cloud information corresponding to all parts included in the target object.

[0084] In this step, under the condition of uniform weight distribution, the average value is calculated by computing the coordinate values of all the point cloud information of the target object, and the position information of the target object is obtained.

[0085] S204B3: If the weight distribution information corresponding to the target object is non-uniform weight distribution, then determine the weight concentration area corresponding to the target object. Calculate the position information of the target object according to the coordinates of the points located within the above weight concentration area. In this embodiment, the weight concentration area is the continuous area with the largest weight corresponding to a single substance inside the target object.

[0086] In this step, it includes determining the voxels with the number of points inside the voxel greater than the average density as the target voxels, combining the target voxels to obtain the weight concentration area, and taking the average value of the coordinates of the points located within the weight concentration area as the position information of the target object.

[0087] From the description of the above embodiments, on the one hand, the embodiments of the present disclosure determine whether the weight distribution of the target object is uniform, determine the weight concentration area in the case of non-uniformity, and use the point cloud cluster in the weight concentration area to determine the position information of the target object, so as to locate the position information to a position closer to the weight concentration area. Then the corresponding grasping points generated subsequently are closer to the position of the weight concentration area, which can make the humanoid robot more stable during the task execution and grasping process. For example, when grasping a mineral water bottle, if the grasping position is in a non-weight concentration area, that is, the internal hollow area, it may cause the bottle to shake left and right in the air after being grasped, that is, the bottle is in an unstable state, then it may cause the bottle to fall and the task execution to fail. Especially for the end effector being a parallel gripper or a two-point gripper, the possibility of execution failure is greater.

[0088] On the other hand, since calculating the centroid of the target object based on all the point cloud coordinates corresponding to all parts included in the target object may cause the centroid to deviate significantly from the true grasping position, and then the grasping position calculated based on the position of the target object may not conform to the true grasping position, resulting in the failure of the grasping task. Therefore, in the case of non-uniform weight distribution, this embodiment only uses the average value of the coordinates of the discrete points located within the weight concentration area as the position information of the target object, rather than calculating the center point coordinates based on the coordinates of all the discrete points in the target object and using this to determine the grasping point. This can make the finally obtained grasping point coordinates more in line with the requirements of the humanoid robot grasping scenario, that is, the grasping point is more in line with the actual operation requirements, making the calculation result of the grasping point coordinates more accurate for this scenario requirement. Thus, while improving the success rate of the humanoid robot grasping the target object, it can improve the calculation efficiency of the position information of the target object, and further improve the task operation efficiency.

[0089] Figure 6Schematic diagram of the point cloud extraction process provided by the embodiments of the present application. As Figure 6 shown, in some embodiments, the above step S204B3 includes:

[0090] S204B31, if the weight distribution information corresponding to the target object is uneven in weight distribution, obtain the substance distribution of the target object.

[0091] S204B32, determine the weight concentration area corresponding to the target object according to the substance distribution of the target object. The weight concentration area is the continuous area with the largest weight corresponding to a single substance inside the target object.

[0092] S204B33, extract the point cloud corresponding to the weight concentration area from the point cloud information.

[0093] And S204B34, take the average value of the coordinates of all points in the point cloud corresponding to the weight concentration area as the position information of the target object.

[0094] When specifically implementing this embodiment, the method of dividing the space where the target object is located into voxels in the above step S204B1 can be used to determine the substance distribution of the target object. It is also possible to determine the above substance distribution according to the ray device or image information. Through the ray device or image information, the distribution of various types of substances in each area of the target object can be identified (including the substance category, the distribution areas of various substances, and the volume of each distribution area). According to the density and volume of different substances, the weight of each type of substance can be calculated, and the weight concentration area can be determined. Then, point cloud extraction can be performed according to the point coordinate range corresponding to the weight concentration area.

[0095] From the description of the above embodiments, it can be seen that the embodiments of the present disclosure calculate the target point coordinates only according to the point cloud information of the weight concentration area for an object with uneven weight distribution. The target point coordinates can be used as the coordinates of the grasping point, rather than calculating the center point coordinates according to the coordinates of all discrete points in the target object. On the one hand, it can make the finally obtained target point coordinates more in line with the grasping scenario of this special type of object, improve the calculation accuracy of the position information of this type of object, and thus facilitate improving the success rate of the humanoid robot in grasping this type of object.

[0096] On the other hand, it makes the located target point closer to the area where the substantial object of the target object is located, so that the grasping point is closer to the area where the substantial object is located, which can make the humanoid robot more stable during the task execution and grasping process, and avoid the grasping point of the humanoid robot being located in the area where the non-substantial object is located, resulting in unstable center of gravity of the object after grasping and causing the phenomenon of left and right shaking, which in turn leads to the failure of task execution.

[0097] In a possible implementation, in the above step S203, according to the image information, obtaining the point cloud information corresponding to the target area includes: steps S2031 to S2033.

[0098] S2031: Obtain the point cloud information of the image information.

[0099] In this step, it includes reading the point cloud information from the image information.

[0100] S2032: According to the preset camera parameters, project each point of the point cloud information of the image information onto a two-dimensional plane to obtain the pixel coordinates of each point.

[0101] In this step, the camera parameters such as the focal length , , the optical center coordinates , , and the internal parameter matrix constructed by using the above camera parameters is as follows:

[0102]

[0103] The internal parameter matrix is a parameter matrix used to describe the camera. It includes the above information such as the focal length and optical center coordinates of the camera, and can convert the image coordinates collected by the camera into the real-world coordinates in the camera coordinate system. In the embodiments of the present application, the above camera is a depth camera. Using the image information and depth information obtained by the depth camera, the point cloud information in the scene can be obtained through a three-dimensional point cloud reconstruction algorithm.

[0104] Using the internal parameter matrix, each point of the point cloud information can be projected onto a two-dimensional plane to obtain the pixel coordinates of each point.

[0105] S2033: If the pixel coordinates of the target point are not within the target area, remove the target point to obtain the point cloud information corresponding to the target area.

[0106] In this step, if the abscissa of the pixel coordinates is greater than the maximum abscissa of the target area and less than the minimum abscissa of the target area, remove the target point; if the ordinate of the pixel coordinates is greater than the maximum ordinate of the target area and less than the minimum ordinate of the target area, remove the target point, and only retain the discrete points located inside the target area. The remaining points are the point cloud information corresponding to the target area.

[0107] From the description of the above embodiments, it can be seen that the embodiments of the present disclosure project the point cloud information onto a two-dimensional plane by combining camera parameters, and remove the points that do not belong to the target object to obtain the point cloud information of the target area.

[0108] In some alternative embodiments, on the basis of the above Figure 2 corresponding embodiments, step S204 includes:

[0109] S2045. Determine whether the target object consists of at least two non - overlapping parts, and at least one of the non - overlapping parts does not have a grasping position according to the image information.

[0110] In this step, determine the target object category according to the instance segmentation algorithm or any object recognition model, and then combine with the preset object category feature library to determine whether it contains a dedicated grasping position according to the target object category. If it contains a dedicated grasping position, use a grasp detection model (such as GraspNet, GG - CNN) to identify whether each part of the target object contains a grasping position. The above - mentioned preset object category feature library records the characteristics of whether each category of object has a grasping position.

[0111] For example, a glass water cup consists of a cup body and a cup handle. The orthographic projections of the two do not overlap. Among them, the cup handle has a grasping position, and the cup body does not have a grasping position.

[0112] S2046. If the target object consists of at least two non - overlapping parts, and at least one of the non - overlapping parts does not have a grasping position, extract the point cloud information containing the continuous region corresponding to the grasping position from the point cloud information corresponding to the target area.

[0113] In this step, it may include determining the region corresponding to the grasping position in the image information, removing the points outside the region corresponding to the grasping position in the point cloud information, and obtaining the point cloud information corresponding to the grasping position.

[0114] S2047. Calculate the position information of the target object according to the coordinate information in the point cloud information containing the continuous region corresponding to the grasping position.

[0115] In this step, take the average value of the coordinates of all points in the point cloud of the region corresponding to the grasping position as the position information of the target object.

[0116] As can be seen from the description of the above embodiments, the embodiments of the present disclosure are directed to a specific type of target object, that is, it is composed of at least two non-overlapping parts, and at least a part of the non-overlapping parts does not have a grasping position. For this type of target object, if the centroid of the target object is calculated according to all the point clouds corresponding to all the parts included in the target object, it may cause the centroid to deviate significantly from the true grasping position, and then the grasping position calculated according to the position of the target object may not conform to the true grasping position, resulting in the failure of the grasping task. In this embodiment, when it is determined through the image information that the target object meets the above type conditions, the position information of the target object is determined only according to the point cloud information of the area corresponding to the grasping position, rather than calculating the center point coordinates according to the coordinates of all the discrete points in the target object, which is beneficial to improving the accuracy of the calculation of the grasping position for this special type of target object, and further beneficial to improving the success rate of the humanoid robot in grasping this special type of target object; on the other hand, only the position information of the target object needs to be calculated according to part of the point cloud, which can improve the calculation efficiency, and thus is beneficial to improving the operation efficiency of the humanoid robot.

[0117] In a possible implementation manner, the image information is collected periodically. The above step S201 is replaced with: obtaining at least three sets of consecutive image information about the target object. Step S204 includes: step S240 and step S241.

[0118] S240: Determine whether the target object is moving according to at least three sets of image information.

[0119] In this step, the method of the above steps S202 to S204 can be adopted to determine three position information according to three sets of image information. If the difference between the three position information is greater than a preset difference threshold, it is determined whether the target object is a dynamically moving object.

[0120] Among them, a set of images can be composed of multiple images, or one image can be used as a set, or multiple images can be combined to obtain a set of images. In this step, the determination of whether the target object is a dynamically moving object can be performed by the above processing unit as the execution subject. For example, if all the images taken from the same shooting angle are the same, it means that the target object has not moved, that is, it is a static object. Otherwise, it means that it is moving, that is, it is a dynamic object.

[0121] S241: If the target object is moving, then predict the position information corresponding to the next period according to at least three sets of image information, and determine the position information corresponding to the next period as the position information of the target object.

[0122] If the target object does not move, continue to execute steps S202 to S207.

[0123] In this step, it may include determining the position information by using each group of image information, calculating the moving speed of the target object based on the acquisition period of the image information and the position information, and predicting the position information of the target object at the next moment according to the position information and the moving speed. It is also possible to calculate the acceleration of the target object on this basis, and combine the acceleration, position information, and moving speed to determine the speed between the current moment and the next moment of image information acquisition, and combine the position information and the time interval of image information acquisition to determine the position information at the next moment of image information acquisition.

[0124] Regarding the direction information of the dynamic object, for each group of images, respective corresponding multiple groups of direction information can be calculated based on step S205, and then information such as the angular velocity and angular acceleration during the movement can be calculated based on the above multiple groups of direction information, and then the direction information of the target object at the next moment can be predicted. Alternatively, the above multiple groups of direction information are respectively converted into rotation matrices, the translation vector is solved through the position change, the angular velocity is estimated through the change of the rotation matrix, so as to deduce the three-dimensional linear velocity and angular velocity of the target object at the next moment.

[0125] As can be seen from the description of the above embodiments, the embodiments of the present disclosure determine whether the target object is actually a dynamically moving object by combining at least three groups of images, so that it can accurately determine whether the target object moves. Different calculation methods for position information and direction information are adopted for static objects and dynamic objects respectively. For the position information of dynamic objects, parameters such as its speed and acceleration can be calculated based on multiple groups of image information, so that the determined new position information is more accurate, and the success rate of the humanoid robot grasping the target object is higher.

[0126] In a possible implementation manner, after executing step S240, if the target object is in a stationary state, steps S202 to S206 are continued to be executed.

[0127] In a possible implementation manner, the position information includes the center point. In the above step S206, according to the position information and direction information of the target object, the grasping control information of the end effector is determined, including: steps S2064 to S2066.

[0128] S2064: Determine the grasping contact area of the end effector according to the position information of the target object. The grasping contact area can wrap the center point.

[0129] Specifically, the position information of the target object can be used as an input parameter of the neural network model to calculate a number of grasping points. The continuous area on the object surface corresponding to the arc surface formed by connecting these grasping points is the above-mentioned grasping contact area. The fact that the above-mentioned grasping contact area can wrap the center point means that: based on the arc surface where the above-mentioned grasping contact area is located, a corresponding virtual closed shape can be formed, that is, the arc surface where the above-mentioned grasping contact area is located is used as a part of a virtual closed shape such as a circle or an ellipse, and the above-mentioned virtual closed shape can wrap the above-mentioned center point.

[0130] S2065: Determine the grasping posture of the end effector according to the direction information of the target object.

[0131] In this step, it includes using the grasping angle corresponding to the direction information, the rotation direction of the end effector, etc. The grasping posture of the end effector matches the direction of the target object.

[0132] S2066: Generate grasping control information according to the grasping contact area and the grasping posture.

[0133] In this step, the grasping contact area and the grasping posture can be input into the grasping control information template to obtain the grasping control information. It is also possible to select a suitable type of end effector according to the grasping contact area and generate the grasping control information corresponding to the type of end effector and the grasping posture.

[0134] From the description of the above embodiments, it can be seen that the embodiments of the present disclosure determine the grasping contact area by position information, use direction information to determine the grasping posture, and generate the grasping control information corresponding to the grasping contact area and the grasping posture, enabling the humanoid robot to stably grasp the target object.

[0135] In a possible implementation manner, the instance segmentation algorithm is the YOLACT algorithm. The idea of the YOLACT algorithm is to decompose the instance segmentation task into two parallel subtasks:

[0136] Generate Prototype Masks: Predict a set of basic masks (similar to semantic segmentation) through a fully convolutional network.

[0137] Predict Mask Coefficients: Generate a set of linear combination coefficients for each detected instance to combine the prototype masks into the final instance masks.

[0138] Through the ingenious design of Prototype Masks + coefficient combination, YOLACT realizes real-time instance segmentation, balances speed and accuracy, and is especially suitable for real-time computing scenarios such as humanoid robots performing grasping tasks.

[0139] By adopting the YOLACT algorithm for instance segmentation, high-precision instance segmentation can be achieved. Moreover, since the YOLACT algorithm can adapt to different types of objects and complex backgrounds, the accuracy of instance segmentation is increased. It should be noted that all the above embodiments disclosed in this application can be freely combined, and the technical solutions obtained after combination are also within the protection scope of this application.

[0140] To implement the above embodiments, an electronic device is further provided in an embodiment of this application.

[0141] Reference Figure 7 , which shows a schematic structural diagram of an electronic device 700 suitable for implementing the embodiments of this application. The electronic device 700 can be the humanoid robot in any of the above embodiments. Figure 7 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of this application.

[0142] As Figure 7 shown, the electronic device 700 can include a processor (such as a central processing unit, a graphics processing unit, etc.) 701, and a memory 702 communicatively connected to the processor. It can perform various appropriate actions and processes according to the programs, computer-executable instructions stored in the memory 702, or the programs loaded from the storage device 708 into the random access memory (Random Access Memory, abbreviated as RAM) 703, and implement the humanoid robot control method based on image segmentation in any of the above embodiments. The memory can be a read-only memory (Read Only Memory, abbreviated as ROM). In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The processing device 701, the memory 702, and the RAM 703 are connected to each other through a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.

[0143] Generally, the following devices can be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a speaker; a storage device 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 can allow the electronic device 700 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 7 shows the electronic device 700 having various devices, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.

[0144] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product that includes a computer program carried on a computer-readable storage medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 709, or installed from the storage device 708, or installed from the memory 702. When the computer program is executed by the processing device 701, the above functions defined in the method of the embodiment of the present application are executed.

[0145] It should be noted that the above computer-readable storage medium in the present application can be a computer-readable signal medium, a computer storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable storage medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0146] The above computer-readable storage medium can be included in the above electronic device; or it can exist separately and not be assembled into the electronic device.

[0147] The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above embodiment.

[0148] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a Local Area Network (LAN) or a Wide Area Network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0149] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and this module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutively represented blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0150] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), System on a Chip (SOC), Complex Programmable Logic Devices (CPLD), and so on.

[0151] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When the processor executes the computer-executable instructions, the technical solution of the humanoid robot control method based on image segmentation in any of the above embodiments is implemented. The implementation principle and beneficial effects are similar to those of the humanoid robot control method based on image segmentation. For details, refer to the implementation principle and beneficial effects of the humanoid robot control method based on image segmentation, which will not be elaborated here.

[0152] In the context of the present application, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0153] The present application also provides a computer program product, including a computer program, which when executed by a processor, implements the technical solution of the humanoid robot control method based on image segmentation in any of the above embodiments. The implementation principle and beneficial effects are similar to those of the humanoid robot control method based on image segmentation. For details, refer to the implementation principle and beneficial effects of the humanoid robot control method based on image segmentation, which will not be elaborated here.

[0154] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principle. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solution formed by the specific combination of the above technical features. It should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, a technical solution formed by mutually replacing the above features with (but not limited to) technical features having similar functions disclosed in the present application.

[0155] Those of ordinary skill in the art will understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments; and the aforementioned storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0156] Finally, it should be noted that: After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily think of other embodiments of the present invention. The present invention is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include the common general knowledge or conventional technical means in the technical field not disclosed in the present invention. It is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A humanoid robot control method based on image segmentation, characterized in that, The humanoid robot includes an end effector, and the method includes: Obtaining image information about a target object; Segmenting the image information of the target object based on an instance segmentation algorithm to obtain a target region corresponding to the target object; Obtaining point cloud information corresponding to the target region according to the image information corresponding to the target region; Calculating the position information of the target object according to the coordinate information in the point cloud information; Performing principal component analysis on the point cloud information to obtain the orientation information of the target object; Determining the grasping control information of the end effector according to the position information and orientation information of the target object; Controlling the end effector to grasp the target object according to the grasping control information.

2. The method according to claim 1, wherein The segmenting the image information of the target object based on an instance segmentation algorithm to obtain a target region corresponding to the target object includes: Associating target objects in different frame images based on a temporal tracking algorithm; Segmenting the associated target objects between different frames based on an instance segmentation algorithm to obtain a target region corresponding to the target object.

3. The method according to claim 1, wherein The calculating the position information of the target object according to the coordinate information in the point cloud information includes: Obtaining an offset between the instance mask output by the instance segmentation algorithm and the true boundary of the target object; When the offset is greater than a preset distance threshold, determining a point coordinate adjustment weight for the corresponding region of the point cloud according to the pixel depth value of the instance mask; wherein, the point coordinate adjustment weight is negatively correlated with the pixel depth value of the instance mask; Adjusting the coordinates in the point cloud information according to the point coordinate adjustment weight of the corresponding region of the point cloud to obtain adjusted point cloud coordinate information; Calculating the position information of the target object according to the adjusted point cloud coordinate information.

4. The method according to claim 1, wherein The determining the grasping control information of the end effector according to the position information and orientation information of the target object includes: Determining a grasping contact area according to the position information of the target object; Determining a mask area corresponding to the grasping contact area according to the instance mask output by the instance segmentation algorithm; Determining the clamping force of the end effector according to the mask area corresponding to the grasping contact area, and the grasping control information includes the clamping force; wherein, the clamping force is positively correlated with the mask area corresponding to the grasping contact area.

5. The method according to claim 1, wherein The calculating the position information of the target object according to the coordinate information in the point cloud information includes: Obtaining the weight distribution information corresponding to the target object; the weight distribution information is evenly distributed or unevenly distributed; If the weight distribution information corresponding to the target object is evenly distributed, calculating the position information of the target object based on the point cloud information corresponding to all parts included in the target object; If the weight distribution information corresponding to the target object is unevenly distributed, determining the weight concentration area corresponding to the target object; calculating the position information of the target object according to the coordinates of the points located in the weight concentration area.

6. The method according to claim 5, wherein Determining the weight concentration area corresponding to the target object; Calculate the position information of the target object according to the coordinates of the points located within the weight concentration area, including: Obtain the substance distribution of the target object; Determine the weight concentration area corresponding to the target object according to the substance distribution of the target object; the weight concentration area is the continuous area with the largest weight corresponding to a single substance inside the target object; Extract the point cloud corresponding to the weight concentration area from the point cloud information; Take the average value of the coordinates of all points in the point cloud corresponding to the weight concentration area as the position information of the target object.

7. The method according to claim 1, characterized in that, The calculation of the position information of the target object according to the coordinate information in the point cloud information includes: Judge whether the target object is composed of at least two non-overlapping parts according to the image information, and at least a part of the non-overlapping parts does not have a grasping position; If the target object is composed of at least two non-overlapping parts, and at least a part of the non-overlapping parts does not have a grasping position, then extract the point cloud information including the area corresponding to the grasping position from the point cloud information corresponding to the target area; Calculate the position information of the target object according to the coordinate information in the point cloud information including the area corresponding to the grasping position.

8. The method according to claim 7, wherein The calculation of the position information of the target object according to the point cloud information including the area corresponding to the grasping position includes: Take the average value of the coordinates of all points in the point cloud of the area corresponding to the grasping position as the position information of the target object.

9. The method according to any one of claims 1 to 8, characterized in that The position information includes a center point; The determination of the grasping control information of the end effector according to the position information and the direction information of the target object includes: Determine the grasping contact area of the end effector according to the position information of the target object; the grasping contact area can wrap the center point; Determine the grasping posture of the end effector according to the direction information of the target object; Generate the grasping control information according to the grasping contact area and the grasping posture.

10. The method according to any one of claims 1 to 8, characterized in that, The instance segmentation algorithm is the YOLACT algorithm.

Citation Information

Cited By

  • Vision-based stable grabbing control method and device for humanoid robot

    CN121018572A