Humanoid Robot Control Method Based on Point Cloud Clustering

By acquiring continuous image information and performing point cloud clustering processing, the static or dynamic state of the target object of the humanoid robot is determined, its position and direction information is calculated, and the grab control information is generated, which solves the problem of low success rate of humanoid robot crawling and improves the accuracy of position and direction.

CN120206542BActive Publication Date: 2025-07-29人形机器人(上海)有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510696161.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-07-29
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

In the prior art, humanoid robots have low accuracy in positioning objects, resulting in a low success rate of grabbing objects.

Method used

By acquiring at least two consecutive sets of image information, the target clustering algorithm is used to cluster the target point cloud information, determine the static or dynamic state of the target object, and calculate the position and direction information according to different states to generate the grab control information of the end effector.

Benefits of technology

The accuracy of the position and direction of the target object is improved, and the success rate of humanoid robot crawling is enhanced. In particular, different calculation methods are adopted for static and dynamic objects, further improving the accuracy of position calculation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120206542B_ABST
    Figure CN120206542B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a humanoid robot control method based on point cloud clustering, belonging to the technical field of robots. The method includes: obtaining at least two sets of continuous image information about a target object; obtaining target point cloud information corresponding to the image information, and performing clustering processing on the target point cloud information based on a target clustering algorithm to obtain first point cloud information corresponding to the target object; determining the state of the target object according to the image information, where the state is static or dynamic; determining position information and orientation information of the target object according to the first point cloud information and the state corresponding to the target object; determining grasping control information of an end effector according to the position information and the orientation information of the target object; and controlling the end effector to grasp the target object according to the grasping control information. This method solves the problem that the success rate of current humanoid robots in grasping objects is relatively low.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robot technology, and particularly to a control method for a humanoid robot based on point cloud clustering. Background Art

[0002] Humanoid robots are one of the important development directions in the field of robots. During the operation of humanoid robots, there are various tasks that require object grasping.

[0003] Currently, in the related art, when a humanoid robot grasps an object, a target detection algorithm is usually used to locate the object, and then the object is grasped.

[0004] However, the inventors found that the related art has at least the following technical problems: because the current humanoid robot has a low accuracy in object pose positioning, the success rate of grasping the object is low. Summary of the Invention

[0005] The embodiments of this application provide a control method for a humanoid robot based on point cloud clustering to solve the problem of the low success rate of the current humanoid robot in grasping objects.

[0006] In a first aspect, the embodiments of this application provide a control method for a humanoid robot based on point cloud clustering. The humanoid robot includes an end effector. The method includes: obtaining at least two sets of consecutive image information about a target object; obtaining the target point cloud information corresponding to the image information, and performing clustering processing on the target point cloud information based on a target clustering algorithm to obtain the first point cloud information corresponding to the target object; determining the state of the target object according to the image information; the state is static or dynamic; determining the position information and orientation information of the target object according to the first point cloud information and the state corresponding to the target object; the calculation methods of the position information corresponding to the static target object and the dynamic target object are different; determining the grasping control information of the end effector according to the position information and orientation information of the target object; controlling the end effector to grasp the target object according to the grasping control information.

[0007] In a possible implementation manner, determining the position information and orientation information of the target object according to the first point cloud information and the state corresponding to the target object includes: if the target object is dynamic, associating the target object based on a target tracking algorithm; modeling the motion process of the target object using a uniform motion model; predicting the position information of the target object at the next moment according to the state at the current moment using a dynamic prediction algorithm.

[0008] In a possible implementation, the method includes: obtaining at least three sets of consecutive image information about a target object; determining whether the target object is a dynamically moving object according to the at least three sets of consecutive image information; and in the case that the target object is a dynamically moving object, predicting the position information corresponding to the target object at the next moment according to the at least three sets of consecutive image information, and using the predicted position information as the position information of the target object.

[0009] In a possible implementation, after controlling the end effector to grasp the target object according to the grasping control information, the method further includes: if the grasping fails, obtaining at least two sets of consecutive new image information about the target object; using the new image information to determine whether the state of the target object changes from static to dynamic; and if the state of the target object changes from static to dynamic, determining the new position information of the target object according to the new image information.

[0010] In a possible implementation, the robot includes at least two types of end effectors; after obtaining at least two sets of consecutive new image information about the target object when the grasping fails, the method further includes: determining the environmental information where the target object is located according to the new image information; generating the target grasping control information of the target type of end effector according to the environmental information and the new position information; and using the target grasping control information to control the target type of end effector to grasp the target object.

[0011] In a possible implementation, determining the grasping control information of the end effector according to the position information and the direction information of the target object includes: obtaining the current task information of the humanoid robot; determining the candidate grasping points of the end effector according to the position information and the direction information of the target object; screening out the target grasping points from the candidate grasping points according to the current task information; the grasping control information includes the target grasping points; and determining the grasping control information of the end effector according to the target grasping points.

[0012] In a possible implementation, screening out the target grasping points from the candidate grasping points according to the current task information includes: obtaining the density information and the volume information of the target object; and screening out the target grasping points from the candidate grasping points according to the current task information, the density information, and the volume information.

[0013] In a possible implementation, screening out the target grasping points from the candidate grasping points according to the current task information, the density information, and the volume information includes: obtaining the joint limit information of the humanoid robot; and screening out the target grasping points from the candidate grasping points according to the current task information, the density information, the volume information, and the joint limit information of the humanoid robot.

[0014] In a possible implementation, the position information and orientation information of the target object are determined according to the first point cloud information and status corresponding to the target object, including: determining the position information of the target object according to the first point cloud information and status corresponding to the target object; performing principal component analysis on the first point cloud information corresponding to the target object to obtain the first sub-orientation information and the second sub-orientation information of the target object; calculating the third sub-orientation information according to the first sub-orientation information and the second sub-orientation information; and generating the orientation information according to the first sub-orientation information, the second sub-orientation information, and the third sub-orientation information.

[0015] In a possible implementation, the target point cloud information corresponding to the image information is obtained, and the target point cloud information is clustered based on the target clustering algorithm to obtain the first point cloud information corresponding to the target object, including: obtaining the initial point cloud information in the image information; identifying the bounding box of the target object in the image information; cropping the initial point cloud information with the bounding box to obtain the valid point cloud information; determining the normal vector of each point in the valid point cloud information; if the included angle between the normal vector of the target point in the valid point cloud information and the preset reference normal is less than the preset angle threshold, and the ordinate of the target point is within the preset ordinate range, then removing the target point from the valid point cloud information; downsampling the remaining valid point cloud information to obtain the target point cloud information; and clustering the target point cloud information to obtain the first point cloud information corresponding to the target object.

[0016] In a second aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor;

[0017] The memory stores computer-executable instructions;

[0018] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementation manners of the first aspect.

[0019] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the above first aspect and / or various possible implementation manners of the first aspect.

[0020] In a fourth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the above first aspect and / or various possible implementation manners of the first aspect.

[0021] The humanoid robot control method based on point cloud clustering provided by the embodiments of the present application obtains at least two sets of consecutive image information, clusters the point cloud information corresponding to the image information to obtain the first point cloud information corresponding to the target object, obtains the position information and orientation information of the target object according to the first point cloud information, and determines the grasping control information of the end effector from the position information and orientation information. Since clustering the point cloud information makes the obtained clustering clusters correspond to the size of the target object itself, it is beneficial to further improve the accuracy of the position and orientation. And different position information calculation methods are adopted for static objects and dynamic objects respectively, which can further improve the accuracy of position calculation. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Other features, objects, and advantages of the present application will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings.

[0023] Figure 1 It is a schematic diagram of the scenario of the humanoid robot control method based on point cloud clustering provided by the present application;

[0024] Figure 2 It is a schematic flowchart of the humanoid robot control method based on point cloud clustering provided by the embodiments of the present application;

[0025] Figure 3 It is a schematic flowchart of the generation process of the grasping control information provided by the embodiments of the present application;

[0026] Figure 4 It is a schematic flowchart of the target grasping point screening process provided by the embodiments of the present application;

[0027] Figure 5 It is a schematic diagram of the candidate grasping points provided by the embodiments of the present application;

[0028] Figure 6 It is a schematic flowchart of the point cloud information clustering process provided by the embodiments of the present application;

[0029] Figure 7 It is a schematic flowchart of the object grasping process provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] The following uses specific specific examples to illustrate the implementation manners of the present application. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in the present application. The present application can also be implemented or applied through other different specific implementation manners. The details in the present application can also be variously modified or changed according to different viewpoints and application systems without departing from the spirit of the present application. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0031] Hereinafter, with reference to the accompanying drawings, embodiments of the present application will be described in detail so that those skilled in the art to which the present application pertains can easily implement it. The present application can be embodied in various different forms and is not limited to the embodiments described herein.

[0032] In the description of the present application, the reference to the expression of terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics represented in connection with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics represented can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples represented in the present application and the features of different embodiments or examples.

[0033] In addition, the terms "first" and "second" are only used for the purpose of indication and cannot be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of such features. In the description of the present application, the meaning of "a plurality" is two or more unless otherwise specifically defined.

[0034] To clearly illustrate the present application, devices irrelevant to the description are omitted, and the same or similar components throughout the specification are given the same reference numerals.

[0035] The technical terms used herein are only for referring to specific embodiments and are not intended to limit the present application. The singular forms used herein also include the plural forms as long as the statements do not explicitly indicate the contrary meaning. The meaning of "including" used in the specification is to embody specific characteristics, regions, integers, steps, operations, elements, and / or components, and does not exclude the existence or addition of other characteristics, regions, integers, steps, operations, elements, and / or components.

[0036] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application will be thorough and complete, and the concept of the example embodiments will be fully conveyed to those skilled in the art. Like reference numerals in the figures denote like or similar structures, and thus their repetitive description will be omitted.

[0037] Figure 1 It is a schematic diagram of the scenario of the humanoid robot control method based on point cloud clustering provided for the present application. As Figure 1 , in this scenario, it includes: a humanoid robot 100 and a target object 200 to be grasped.

[0038] Among them, the humanoid robot 100 includes a depth camera 1001, at least one end effector 1002, and a processor 1003. The end effector can be a parallel gripper, a rotary gripper, a dexterous hand, etc. The target object 200 includes an object to be grasped.

[0039] The image information obtained by the depth camera may include a two-dimensional RGB image and a three-dimensional depth map. The depth map is the distance information of each pixel point. Then, the depth map can be converted into point cloud information based on the SDK built in the depth camera. The processor 1003 determines the position and orientation of the target object 200 according to the point cloud information, and controls the end effector 1002 to grasp the target object 200 according to the position and orientation.

[0040] Figure 2 It is a schematic flowchart of a humanoid robot control method based on point cloud clustering provided by an embodiment of the present application. The execution subject of the embodiment of the present application can be Figure 1 the entire humanoid robot in Figure 2 As shown, the method includes:

[0041] S201: Obtain at least two sets of continuous image information about the target object.

[0042] In this step, the depth camera can be used to periodically or continuously photograph the target object to obtain at least two sets of continuous image information.

[0043] Among them, the image information may include a two-dimensional black and white or color image.

[0044] S202: Obtain the target point cloud information corresponding to the image information, and perform clustering processing on the target point cloud information based on the target clustering algorithm to obtain the first point cloud information corresponding to the target object.

[0045] In this step, it may include reading the target point cloud information from the image information in the manner described above. The target clustering algorithm may include the K-means clustering algorithm, the point cloud region growing clustering algorithm, the DBSCAN clustering algorithm, etc.

[0046] In some optional embodiments, before performing clustering processing on the target point cloud information based on the target clustering algorithm in the above step S202 to obtain the first point cloud information corresponding to the target object, a tabletop removal operation may be performed on the first point cloud information based on the relevant tabletop removal algorithm. The tabletop removal operation can be implemented based on the existing related technologies.

[0047] S203: Determine the state of the target object according to the above image information. The above state is static or dynamic.

[0048] For example, if all the images captured by a humanoid robot at the same position and with the same shooting angle are the same, the target object is a static object; otherwise, it is a dynamic object.

[0049] Alternatively, the following steps can be used to determine whether the target object is static or dynamic: First, identify the target object in the image through a target detection algorithm, then achieve cross-frame association of the target object according to a target tracking algorithm, and then perform motion analysis, such as calculating displacement, velocity, or optical flow changes. Exemplarily, the inter-frame Euclidean distance of the center point coordinates of the object can be calculated. If the displacement exceeds a threshold (such as 5 pixels), it is determined to be dynamic; otherwise, it is determined to be a static object.

[0050] S204: Determine the position information and orientation information of the target object according to the first point cloud information and status corresponding to the target object.

[0051] In this step, a calculation method corresponding to the status of the target object is adopted to determine the position information of the target object. The calculation methods for the position information of static objects and dynamic objects are different. The orientation information may include the posture and rotation direction of the target object. The calculation methods for the orientation information corresponding to static target objects and dynamic target objects are also different.

[0052] Among them, the calculation method for the position information of a static object can be: calculate the average value of the coordinates of each point in the first point cloud information as the central coordinate or centroid position corresponding to the first point cloud information, which is used as the position information of the target object.

[0053] The calculation method for the orientation information of a static object can be:

[0054] Input the above first point cloud information into a trained neural network model such as an RNN (Recurrent Neural Network), and the neural network model can automatically extract the features in the point cloud and directly calculate the orientation information of the target object. The training process of the neural network model includes: annotating the three-dimensional coordinates and orientation of the target object on the three-dimensional point cloud data as training samples, and training the neural network model.

[0055] The calculation method for the position information of a dynamic object can be:

[0056] First, identify the target object in the image through a target detection algorithm, then achieve cross-frame association of the same object according to a target tracking algorithm (such as Optical Flow), and then perform motion modeling, such as using a constant velocity model (CV) to predict the position of the target object at the next moment, or using a dynamic prediction algorithm such as a neural network or Kalman filter to predict the position at the next moment.

[0057] The calculation method of the orientation information of the dynamic object can be as follows:

[0058] In a time series of three frames or more, calculate the orientation information corresponding to each image respectively, convert the orientation information corresponding to each image into a rotation matrix, solve the translation vector through the position change, estimate the angular velocity through the change of the rotation matrix, so as to deduce the three-dimensional linear velocity and angular velocity of the target object at the next moment.

[0059] Alternatively, use the principal component analysis method to calculate the principal direction corresponding to each image respectively. Specifically, that is, establish a point cloud covariance matrix according to the first point cloud information, then perform eigenvalue decomposition on the point cloud covariance matrix to obtain each eigenvalue and the eigenvector corresponding to each eigenvalue, and use the eigenvector with the largest eigenvalue as the principal direction. Then represent the above principal direction as a rotation matrix, solve the translation vector through the position change, estimate the angular velocity through the change of the rotation matrix, so as to deduce the three-dimensional linear velocity and angular velocity of the target object at the next moment, and determine the orientation information at the next moment according to the orientation information, three-dimensional linear velocity, and angular velocity of the last or the last N images.

[0060] S205: Determine the grasping control information of the end effector according to the position information and orientation information of the target object.

[0061] In this step, it may include inputting the position information and orientation information of the target object into the grasping control information template to obtain the grasping control information. It may also include converting the position information and orientation information to the base coordinate system of the robot, and then generating the corresponding grasping control information. It may also include determining the moving distance and direction of the end effector from the position information and orientation information of the target object, writing the moving distance and direction into the control information template to obtain the grasping control information.

[0062] Among them, the grasping control information may include the moving distance, speed, acceleration, jaw opening degree, grasping point position and posture, etc. of the end effector. The centroid position of the above target point cloud is located inside the area enclosed by the grasping posture of the end effector.

[0063] Specifically, in the case of obtaining the position information and orientation information of the target object, the object size and shape can be known based on point cloud clustering. Then the object position can be converted to the base coordinate system of the robotic arm through hand-eye calibration to ensure the alignment of the coordinate systems of the robot vision system (camera) and the robotic arm.

[0064] Then, grasping planning can be carried out. The planning specifically includes: determining the grasping strategy, generating grasping points, and obstacle avoidance verification. Regarding the grasping strategy, the grasping strategy can be determined based on the task requirements and object attributes. For example, for a cube, a parallel gripping strategy can be adopted; for a cylinder, a wrapping grasping strategy can be adopted; for an irregular object, a multi-point adsorption strategy (such as a dexterous hand with suction cups) can be adopted. The finger opening amplitude and grasping force can be determined based on the object size and weight respectively. For example, for tasks that require precise operation (such as picking up an egg), a compliant control strategy can be adopted; for tasks that require strong grasping (such as carrying a box), a high stiffness mode can be adopted. Regarding the grasping points, the grasping points can be generated based on the offline generated grasping template, or the pose of the grasping points can be calculated in real time based on a neural network or an optimization algorithm. Then, the motion planning algorithm can be used to check whether the motion path of the robotic arm collides.

[0065] This step may further include: determining the control parameters of the end effector. For example, the pose of the grasping point (position + direction) is converted into the joint angles of the robotic arm based on the inverse kinematics solution method. The robotic arm moves to the pre-grasping position according to the joint angles of the robotic arm, and the end effector closes until it contacts the cup handle, and the grasping force is maintained based on the force control module.

[0066] S206: Control the end effector to grasp the target object according to the grasping control information.

[0067] In this step, the generated grasping control information is transmitted to the drive system of the end effector in the form of an electrical signal or a digital signal, so that the drive system performs corresponding actions, thereby enabling the end effector to grasp the target object.

[0068] As can be seen from the description of the above embodiments, the embodiments of the present disclosure obtain at least two sets of continuous image information, cluster the point cloud information corresponding to the image information to obtain the first point cloud information corresponding to the target object, obtain the position information and orientation information of the target object according to the first point cloud information, and determine the grasping control information of the end effector from the position information and orientation information. Since clustering the point cloud information makes the obtained clustering clusters correspond to the size of the target object body, it is beneficial to further improve the accuracy of the position and orientation; and different position information calculation methods are adopted for static objects and dynamic objects respectively, which is further beneficial to improving the accuracy of position calculation.

[0069] In some alternative embodiments, for a static target object, in the above step S201, the shooting angles of at least two captured images are different. In the above step S204, based on the images with different angles, the respective corresponding first point cloud information is extracted, and then the intersection of all the first point cloud information is taken as the target point cloud. The average value of the coordinates of all discrete points in the target point cloud is calculated as the position information of the target object. In this way, since different-angle images of the target object to be grasped by the humanoid robot are comprehensively considered, the accuracy of the position information of the target object can be further improved, which is beneficial to improving the grasping success rate of the humanoid robot.

[0070] In some alternative embodiments, in the above step S202, at least two clustering algorithms are used to cluster the target point cloud information, and at least two first point cloud information are obtained. Among them, the clustering algorithms may include at least two of the DBSCAN clustering algorithm, the K-means clustering algorithm, the point cloud region growing clustering algorithm, etc.

[0071] The above step S204 includes: step S204A1 to step S204A3.

[0072] S204A1: Merge the respective first point cloud information corresponding to the target object to obtain the merged point cloud information corresponding to the target object.

[0073] In this step, it includes taking the union of the respective first point cloud information to obtain the merged point cloud information corresponding to the target object.

[0074] S204A2: Determine the position information and orientation information of the target object according to the merged point cloud information and state corresponding to the target object.

[0075] In this step, for a static object, it includes adding the three-dimensional coordinates of each point in the merged point cloud information and then dividing by the number of all points in the merged point cloud information to obtain the position information of the static target object, and performing principal component analysis on the three-dimensional coordinates of each point in the merged point cloud information to obtain the orientation information.

[0076] As can be seen from the description of the above embodiments, the embodiments of the present disclosure achieve avoiding the errors caused by a single clustering algorithm and increasing the accuracy of the position information by using multiple clustering algorithms for clustering and merging the respective first point cloud information, thereby increasing the accuracy of grasping the target object.

[0077] In some alternative embodiments, obtaining the target point cloud information corresponding to the image information in the above step S202 includes: step S202A1 to step S202A3.

[0078] S202A1: Obtain the initial point cloud information corresponding to the image information.

[0079] In this step, it includes reading point cloud information from the image information to obtain the initial point cloud information.

[0080] S202A2: Perform object detection on the image information according to the object detection algorithm to obtain the target region corresponding to the target object.

[0081] In this step, the object detection algorithm may include YOLO, R-CNN (Region-based Convolutional Neural Network) algorithm, etc. The target region may include the region framed by the coordinates of the recognition box of the target object, and the diagonal coordinates of the recognition box can be used to represent it.

[0082] S202A3: Crop the target point cloud information from the initial point cloud information according to the target region.

[0083] In this step, it includes using the point cloud information within the target region in the initial point cloud information as the target point cloud information.

[0084] As can be seen from the description of the above embodiments, the embodiments of the present disclosure perform object detection on the initial point cloud information in the image information, and crop the initial point cloud information according to the object detection result to obtain the target point cloud information, realizing the deletion of the point cloud information, reducing the amount of data processing, increasing the data processing speed; and removing interference information, making the calculation result more accurate.

[0085] In some optional embodiments, in the above step S202, clustering the target point cloud information based on the target clustering algorithm to obtain the first point cloud information corresponding to the target object includes: step S202B1 and step S202B2.

[0086] S202B1: Cluster the target point cloud information based on the target clustering algorithm to obtain the second point cloud information corresponding to multiple objects.

[0087] This step is similar to the above step S202, and clusters to obtain the second point cloud information corresponding to one or more objects.

[0088] For example, the current image information corresponds to 3 objects A, B, and C. After clustering the target point cloud information, the second point cloud information corresponding to objects A, B, and C is obtained respectively.

[0089] S202B2: Classify the second point cloud information corresponding to multiple objects based on the classifier to obtain the first point cloud information corresponding to the target object.

[0090] In this step, the second point cloud information can be identified and classified based on the trained classifier. For example, the second point cloud information corresponding to objects A, B, and C can be identified, the target object can be selected from them, and the second point cloud information corresponding to the target object (such as object A) can be determined as the first point cloud information.

[0091] As can be seen from the description of the above embodiments, the embodiments of the present disclosure obtain the first point cloud information by classifying the second point cloud information after clustering, so as to accurately find the object to be grasped in the case of multiple objects.

[0092] In some alternative embodiments, the classifier in the above step S202B2 can be replaced by a convolutional neural network, where the convolutional neural network includes a feature extraction network and a classifier.

[0093] Figure 3 It is a schematic diagram of the process for generating the grasping control information provided by the embodiments of the present application. As Figure 3 shown, in the above step S205, according to the position information and orientation information of the target object, the grasping control information of the end effector is determined, including: step S205A1 to step S205A4.

[0094] S205A1: Obtain the current task information of the humanoid robot.

[0095] In this step, the current task information is, for example, grasping the target object, pushing the target object, dragging the target object, etc.

[0096] S205A2: Determine the candidate grasping points of the end effector according to the position information and orientation information of the target object.

[0097] In this step, it may include selecting candidate grasping points from the graspable positions of the target object according to the position information and orientation information of the target object. The above candidate grasping points can be determined based on the offline generated grasping templates, or can be calculated in real time using a neural network (such as 6-DoF GraspNet) or an optimization algorithm. There are multiple above candidate grasping points.

[0098] For example, if the target object is a box, candidate grasping points can be selected on the surface of the box according to the position information and orientation information. The target object in the embodiments of the present application can also be an object such as a cardboard box or a barrel, which will not be elaborated here.

[0099] S205A3: Screen out the target grasping points from the candidate grasping points according to the current task information. The grasping control information includes the target grasping points. The target grasping point is one of the candidate grasping points.

[0100] In this step, for example, if the current task information is to carry or push a box, different candidate grasping points can be selected as the target grasping point. When pushing a box, a point near the center of the side of the box body can be selected as the target grasping point, which is more labor-saving and can save the energy consumption of the humanoid robot. When carrying a box, a candidate grasping point located on the edge of the box body can be selected as the target grasping point, which can ensure the success rate of the humanoid robot in grasping the box body.

[0101] S205A4: Determine the grasping control information of the end effector according to the target grasping point.

[0102] In this step, the appropriate grasping posture, grasping position, and jaw opening degree can be selected according to the target grasping point to generate the grasping control information of the end effector. The size of the target object can also be determined by combining the target grasping point and the point cloud of the target object. The appropriate grasping strategy, jaw opening degree, and posture can be selected according to the size of the object, and the grasping control information can be generated by using the grasping strategy, jaw opening degree, and posture.

[0103] As can be seen from the description of the above embodiments, the embodiments of the present disclosure determine the candidate grasping points through the position information, and combine the current task information of the humanoid robot to select the target grasping point from the candidate grasping points, so that the humanoid robot has lower energy consumption when performing task operations.

[0104] Figure 4 It is a schematic diagram of the target grasping point screening process provided by the embodiment of the present application. As Figure 4 shown, in the above step S205A3, screening the target grasping point from the candidate grasping points according to the current task information includes: step S531 and step S532.

[0105] S531: Obtain the density information and volume information of the target object.

[0106] In this step, the spatial size occupied by the clustering cluster obtained by clustering can be used as the volume information, and the density information can be obtained by dividing the number of points in the clustering cluster by the volume information. The material and volume information of the target object can also be identified according to the image information, and the corresponding relationship between the material and the density can be found according to the material of the target object to obtain the density information of the target object. In the case where the target object includes multiple materials, the volume occupied by various materials of the target object in the target object can be used as the object density information and volume information. The target object can also be scanned by a ray device to obtain various internal materials and corresponding volumes. The density can be determined according to the material, and then combined with the volume occupied by the substances corresponding to various materials, and thus the weights corresponding to various materials (i.e., the product of density and corresponding volume) can be determined.

[0107] For example, the target object is a box filled with water and the water level is lower than the maximum height inside the box, that is, the box is not full of water. In this case, find the density of water and the density of the box, and combine the volumes of water and the box to determine the overall density of the target object. The volume information can be obtained based on the length, width, height, and thickness of the box, as well as the length, width, and height occupied by the water inside the box, to get the volumes of water and the box.

[0108] S532: Filter out the target grasping points from the candidate grasping points according to the current task information, density information, and volume information.

[0109] In this step, it may include determining the weight concentration area according to the density information and volume information, and selecting the target grasping point from the candidate grasping points in the weight concentration area. Based on the above steps, the weights corresponding to various materials inside the target object can be obtained, and the area where the material with the largest weight is located is determined as the weight concentration area. That is, the above weight concentration area is the continuous area with the largest weight corresponding to a single substance inside the target object. Among them, the above candidate grasping points and target grasping points are both located in the weight concentration area. For example, if the water in the above water tank is not full, then for the scenario where the target object is the water tank, the weight concentration area is the area where the water is located. Or for the scenario where the target object is a mineral water bottle, the weight concentration area is also the area where the water is located.

[0110] In this step, select the point that matches the current task information and is located in the weight concentration area from the candidate grasping points as the target grasping point. Specifically, the coordinate average value of all discrete points included inside the above weight concentration area can be calculated, and according to the coordinates of each candidate grasping point and the above coordinate average value, calculate the candidate grasping point with the shortest distance to the point corresponding to the above coordinate average value as the target grasping point.

[0111] From the description of the above embodiments, it can be seen that the embodiments of the present disclosure determine the target grasping point uniquely from the candidate grasping points according to the current task information and the weight concentration area by obtaining the object density information and volume information in the target object, and combining the current task information, density information, and volume information, rather than calculating the center point coordinates based on the coordinates of all discrete points in the target object and using this to determine the target grasping point. This can make the coordinates of the finally obtained target grasping point more in line with the requirements of the humanoid robot grasping scenario, that is, make the target grasping point more in line with the actual operation requirements, make the calculation result of the target grasping point coordinates more accurate for the scenario requirements, thereby improving the success rate of the humanoid robot grasping the target object and improving the task operation efficiency at the same time.

[0112] On the other hand, since the target grasping point matches the weight distribution of the object, the target grasping point determined based on the weight distribution of the target object can complete the task more labor - saving, which is beneficial to saving the energy consumption of the humanoid robot and improving its endurance. Moreover, it can avoid the situation that the finally calculated grasping point deviates too much from the center of gravity of the target object, resulting in shaking of the target object after grasping, and improve the stability of the object after grasping.

[0113] Figure 5 This is a schematic diagram of candidate grasping points provided by an embodiment of the present application. As Figure 5 shown, the target object is a water - filled water tank, and the current task information is to push the water tank. Then, the points on the side of the water tank below the water surface inside the tank can all be candidate grasping points, which are simplified to three points: upper, middle, and lower in Figure 5 . Determine the density information and volume information according to the point cloud information, and determine the weight - distribution concentration area as the part of the water in the box body according to the density information and volume information. If the current task information is to push the box body of the water tank, since the water tank is half - filled with water, considering the magnitude of the force required and the structural stability of the box body, it is more appropriate to push the middle candidate grasping point (i.e., point middle) or the lower candidate grasping point (i.e., point lower) of the water tank. Taking the middle candidate grasping point or the lower candidate grasping point of the water tank as the target grasping point will be more labor - saving compared to selecting the candidate grasping point located above, which is beneficial to saving the energy consumption of the robot and improving its endurance.

[0114] In some optional embodiments, the above - mentioned S532 screens out the target grasping point from the candidate grasping points according to the current task information, object density information, and volume information, including: step S5321 and step S5322.

[0115] S5321: Obtain the joint limit information of the humanoid robot.

[0116] In this step, the joint limit information may include the rotation - angle limit of the end - effector, the rotation - angle limit of the knee joint, the rotation - angle limit of the knee joint, etc., and may also include the ground - clearance limit of the joint, etc.

[0117] For example, the rotation angle of the knee joint cannot make the angle between the thigh and the calf less than the first preset angle, and the rotation angle of the end - effector cannot be greater than the second preset angle, etc. The ground - clearance limit of the joint, for example, the end - effector cannot be lower than the first preset distance from the ground, and the ground - clearance limit of the hip joint, for example, the distance between the hip joint and the ground cannot be lower than the second preset distance.

[0118] Among them, each preset angle and preset distance can be pre - set by the staff.

[0119] S5322: Select a target grasping point from the candidate grasping points according to the current task information, density information, volume information, and joint limit information of the humanoid robot. Specifically, in implementation, this step can be achieved through a neural network model such as the 6-DoFGraspNet model.

[0120] In this step, continue to refer to Figure 5 , determine the area where the weight distribution in the target object is concentrated from the density information and volume information. According to the area where the weight distribution is concentrated (concentrated in the middle and lower parts of the box, including Figure 5 the points "middle" and "lower" in it) and the current task information (pushing the box), two candidate grasping points, the middle and the lower, can be selected. Considering from the perspective of saving effort and energy consumption, the lower candidate grasping point (i.e., point lower) is more labor-saving than the middle candidate grasping point (i.e., point middle). However, considering the joint limit information of the humanoid robot, such as the ground clearance limit of the joint, the joints and end effectors of the humanoid robot may not be able to reach the lower candidate grasping point, which may lead to the failure of task execution. Therefore, in this embodiment, when the robot cannot reach the lower candidate grasping point, the "middle" candidate grasping point is used as the target grasping point. This can take into account both labor-saving and the success rate of task execution. While ensuring the success rate of task execution, it can save effort, is beneficial to saving the energy consumption of the robot, and improves its endurance.

[0121] From the description of the above embodiments, it can be seen that the embodiments of the present disclosure screen a suitable point as the target grasping point by obtaining the joint limit information of the humanoid robot and combining the current task information, density information, volume information, and joint limit information of the humanoid robot. While reducing the energy consumption of the humanoid robot, it can improve the accuracy of the finally determined target grasping point, thereby facilitating the improvement of the probability of successful task execution and avoiding damage or failure of the humanoid robot during task execution.

[0122] In some alternative embodiments, in the above step S204, according to the first point cloud information and state corresponding to the target object, determining the position information and orientation information of the target object includes: steps S204B1 to S204B4.

[0123] S204B1: Determine the position information of the target object according to the first point cloud information and state corresponding to the target object.

[0124] In this step, the determination method of the position information is similar to the above step S204 and will not be elaborated here.

[0125] S204B2: Perform principal component analysis on the first point cloud information corresponding to the target object to obtain the first sub-orientation information and the second sub-orientation information of the target object.

[0126] Specifically, a point cloud covariance matrix is established based on the first point cloud information, and then the point cloud covariance matrix is eigen-decomposed to obtain each eigenvalue and the eigenvector corresponding to each eigenvalue. The eigenvector corresponding to the largest eigenvalue is used as the first sub-direction, and the eigenvector corresponding to the second largest eigenvalue (i.e., when the eigenvalues are sorted from largest to smallest and located at the second position in the sequence) is used as the second sub-direction.

[0127] S204B3: Calculate the third sub-direction information based on the first sub-direction information and the second sub-direction information.

[0128] In this step, the third sub-direction information is perpendicular to the first sub-direction information and the second sub-direction information and conforms to the right-hand rule. Therefore, the third sub-direction information can be determined from the first sub-direction information, the second sub-direction information, and the vector right-hand rule. Specifically, the vector corresponding to the third sub-direction is obtained by performing a vector product on the vector corresponding to the first sub-direction and the vector corresponding to the second sub-direction based on the vector right-hand rule.

[0129] For example, the first sub-direction information is used as the X-axis direction of the target object, and the second sub-direction information is used as the Y-axis direction of the target object.

[0130] S204B4: Generate direction information based on the first sub-direction information, the second sub-direction information, and the third sub-direction information.

[0131] In this step, it includes combining the first sub-direction information, the second sub-direction information, and the third sub-direction information to obtain the direction information. A local coordinate system including three coordinate axes is constructed based on the vectors corresponding to the three sub-directions, which can represent the corresponding direction information.

[0132] As can be seen from the description of the above embodiments, the embodiments of the present disclosure perform principal component analysis on the point cloud information corresponding to the target object to obtain two directions, and determine the third direction from these two directions, thereby obtaining three directions of the target object.

[0133] In some alternative embodiments, the above step S201 is replaced with the step of:

[0134] Obtain at least three sets of consecutive image information about the target object.

[0135] Step S204 includes: step S210 and step S211.

[0136] S210: Determine whether the target object is a dynamically moving object based on at least three sets of consecutive image information.

[0137] In this step, the methods of the above steps S202 and S204 can be adopted to determine three position information based on three sets of image information. If the difference between the three position information is greater than a preset difference threshold, it is determined whether the target object is a dynamic moving object.

[0138] Among them, a set of images can be composed of multiple images, or a single image can be used as a set, or multiple images can be combined to obtain a set of images. In this step, the above processor can be used as the execution entity to determine whether the target object is a dynamic moving object.

[0139] S211: When the target object is a dynamic moving object, based on at least three sets of continuous image information, predict the position information corresponding to the target object at the next moment as the position information of the target object.

[0140] In this step, it can include using each set of image information to determine the position information, calculating the moving speed of the target object from the acquisition period of the image information and the position information, and predicting the position information of the target object at the next moment based on the position information and the moving speed. It can also calculate the acceleration of the target object on this basis, combine the acceleration, position information and moving speed to determine the speed between the current moment and the next image information acquisition moment, and combine the position information and the time interval of the image information acquisition to determine the position information at the next image information acquisition moment.

[0141] In a possible implementation manner, when the target object is a dynamic moving object in the above step S211, predicting the position information corresponding to the target object at the next moment based on at least three sets of continuous image information as the position information of the target object further includes: steps S2111 to S2114.

[0142] S2111: When the target object is a dynamic moving object, determine the position information of the target object in at least three sets of continuous image information.

[0143] In this step, for each set of image information, adopt steps S202 to S204 to obtain the position information of the target object.

[0144] S2112: Determine the object speed of the target object according to the position information of the target object.

[0145] In this step, it includes calculating the position information of every two groups, calculating the moving distance, dividing the moving distance by the preset acquisition period of the image information to obtain the moving speed of the target object during the acquisition process of every two sets of continuous image information, and calculating the average value of N moving speeds to obtain the object speed.

[0146] S2113: Calculate the object acceleration of the target object according to the object speed of the target object.

[0147] In this step, it includes dividing the speeds of two adjacent objects obtained by calculation by a preset acquisition period of image information to obtain the object acceleration. For example, the Kalman filtering method or the particle filtering method can be used to predict motion trends such as speed and acceleration.

[0148] S2114: Calculate the position information corresponding to the target object at the next moment according to the position information, object speed, and object acceleration of the target object.

[0149] In this step, it includes inputting the position information, object speed, and object acceleration into a pre-trained position information determination model to obtain the position information corresponding to the target object at the next moment. It may also include inputting the position information, object speed, and object acceleration into a preset equation to obtain the position information corresponding to the target object at the next moment. The Kalman filtering method can also be used to determine the position information corresponding to the target object at the next moment.

[0150] As can be seen from the description of the above embodiments, the embodiments of the present disclosure determine the position information of the target object in each image information, determine the object speed using the position information, calculate the object acceleration from the object speed, and calculate the position information of the target object at the next moment by combining the position information, object speed, and acceleration, making the calculated position information more accurate.

[0151] In some optional embodiments, after the above step S210, it further includes: if the target object is a static object, continue to execute steps S204 to S205.

[0152] As can be seen from the description of the above embodiments, the embodiments of the present disclosure determine whether the target object is actually a dynamically moving object by combining at least three sets of consecutive images, thereby more accurately determining whether the target object moves and calculating more accurate position information of the target object, making the success rate of the humanoid robot grasping the target object higher, that is, improving the probability of successful task execution.

[0153] In some optional embodiments, the target clustering algorithm is a point cloud region growing clustering algorithm.

[0154] Among them, the point cloud region growing clustering algorithm forms a growing region by selecting seed points, judging which points meet the criteria of the growing region (color and distance meet the preset requirements) according to the attributes of the seed points, starting from the seed points, and gradually adding adjacent points to the current region according to the growing criteria until all points are divided into a certain region, completing the clustering of the point cloud.

[0155] As can be seen from the description of the above embodiments, the embodiments of the present disclosure use a point cloud region growing clustering algorithm, which considers the spatial relationship between points and the information of multiple adjacent points for clustering, rather than relying only on the features of a single point. Therefore, it can suppress the influence of noise to a certain extent and can process objects with complex shapes.

[0156] Figure 6 It is a schematic diagram of the point cloud information clustering process provided by the embodiments of the present application. As Figure 6 shown, in the above step S202, the target point cloud information corresponding to the image information is obtained, and the target point cloud information is clustered based on the target clustering algorithm to obtain the first point cloud information corresponding to the target object, including: step S202C1 to step S202C7.

[0157] S202C1: Obtain the initial point cloud information in the image information.

[0158] In this step, the initial point cloud information in the image information can be directly read.

[0159] S202C2: Identify the bounding box of the target object in the image information.

[0160] In this step, the YOLO algorithm or other object recognition algorithms can be used to obtain the bounding box of the target object.

[0161] S202C3: Crop the initial point cloud information using the bounding box to obtain the effective point cloud information.

[0162] In this step, the points with coordinates within the bounding box in the initial point cloud information are retained to obtain the effective point cloud information.

[0163] S202C4: Determine the normal vector of each point in the effective point cloud information.

[0164] In this step, surface fitting can be performed based on the local neighborhood, and the normal of the surface is used as the normal vector of the point. Principal component analysis can also be used to determine the normal vector of each point.

[0165] S202C5: If the included angle between the normal vector of the target point in the effective point cloud information and the preset reference normal is less than the preset angle threshold, and the ordinate of the target point is within the preset ordinate interval range, then the target point is removed from the effective point cloud information.

[0166] In this step, the preset angle threshold and the ordinate interval range can be preset by the staff according to experimental data or empirical parameters.

[0167] Among them, the target point can be a point belonging to the plane where the target object is placed.

[0168] S202C6: Downsample the remaining valid point cloud information to obtain target point cloud information.

[0169] In this step, downsampling the remaining valid point cloud information may include dividing the point cloud into three-dimensional grids (voxels) of uniform size using a voxel filtering method, and replacing all points in the voxel with the centroid or center point within each voxel to obtain target point cloud information.

[0170] S202C7: Perform clustering processing on the target point cloud information to obtain the first point cloud information corresponding to the target object.

[0171] In this step, the process of this step is similar to the clustering process in the above step S202 and will not be elaborated here.

[0172] As can be seen from the description of the above embodiments, the embodiments of the present disclosure identify the bounding box of the target object, use the bounding box to crop the initial point cloud information, retain the valid point cloud information, remove the points in the valid point cloud information where the included angle between the normal vector and the preset reference normal is less than the preset angle threshold, thereby removing the points on the plane where the target object is placed, then downsample the valid point cloud information to reduce the point cloud density, perform clustering on the target point cloud information to obtain the first point cloud information, achieving a reduction in the number of points in the point cloud, reducing the duration and energy consumption of data processing, and increasing the data processing efficiency.

[0173] In some alternative embodiments, after controlling the end effector to grasp the target object according to the grasping control information in the above step S206, the following steps are further included: steps S240 to S242.

[0174] S240: If the grasping fails, obtain at least two sets of consecutive new image information about the target object.

[0175] In this step, it can be determined whether the grasping of the target object fails according to the feedback of the force / torque sensor. It may also include determining whether the grasping fails by taking image information of the grasping process.

[0176] S241: Use the new image information to determine whether the state of the target object changes from static to dynamic.

[0177] In this step, if it is determined according to the image information obtained in step S201 that the state of the target object is static, and it is determined according to the new image information obtained in step S240 that the state of the target object is dynamic, then it is determined that the state of the target object changes from static to dynamic. The method for determining whether the state of the target object is dynamic is similar to the above step S210 and will not be elaborated here.

[0178] S242: If the state of the target object changes from static to dynamic, determine the new position information of the target object according to the new image information.

[0179] In this step, the position information can be determined in the same way as in steps S202 to S204. Alternatively, the position information of the target object in each image information can be determined separately, and then based on the position information of the target object in each image information, the speed and acceleration of the target object can be calculated. Based on the position information, speed, and acceleration of the target object in each image information, the new position information can be calculated.

[0180] As can be seen from the description of the above embodiments, in the embodiments of the present disclosure, after a grasping failure and when it is detected that the target object changes from static to dynamic, different position information calculation methods are adopted for static objects and dynamic objects respectively, and parameters such as the speed and acceleration of the dynamic object are calculated based on the current new image information, so that the determined new position information is more accurate, which is conducive to improving the success rate of the next grasping.

[0181] In some alternative embodiments, the robot includes at least two types of end effectors.

[0182] Figure 7 This is a schematic diagram of the object grasping process provided by the embodiments of the present application. As Figure 7 shown, after the grasping fails in step S240 above and at least two sets of consecutive new image information about the target object are obtained, the following steps are further included: steps S250 to S252.

[0183] S250: Determine the environmental information where the target object is located according to the new image information.

[0184] In this step, image recognition is performed according to the new image information to obtain the environmental information where the target object is located.

[0185] Among them, the environmental information is, for example, whether the target object is in a corner or whether the target object is in an environment with obstacles.

[0186] S251: Generate the target grasping control information of the end effector of the target type according to the environmental information and the new position information.

[0187] In this step, it may include selecting a suitable end effector as the end effector of the target type according to the environmental information and the new position information. Write the environmental information into the grasping control information template of the end effector of the target type to obtain the target grasping control information.

[0188] For example, if the target object moves to a narrow position, the target grasping control information of the parallel gripper can be generated. Another example is that if the environmental information is a position with many obstacles, the target grasping control information of the dexterous hand can be generated.

[0189] S252: Use the target grasping control information to control the end effector of the target type to grasp the target object.

[0190] This step can be similar to the above step S206 and will not be elaborated here.

[0191] As can be seen from the description of the above embodiments, in the embodiments of the present disclosure, after the target object moves, a suitable end effector is used to grasp the target object according to the new environmental information, so that the success rate of grasping is higher.

[0192] In some alternative embodiments, there are at least two end effectors. In the above step S205, according to the position information and orientation information of the target object, determining the grasping control information of the end effector includes: steps S205B1 to S205B5.

[0193] S205B1: Obtain the relative coordinates of each end effector, where the relative coordinates are the relative coordinates of the end effector relative to the main body of the humanoid robot.

[0194] In this step, the relative coordinates of the end effector can be calculated by reading the angles of the joints of the humanoid robot and the preset arm lengths. It is also possible to obtain the relative coordinates of the end effector after the completion of the execution of the historical task by reading the execution records of the historical tasks.

[0195] S205B2: Determine the target end effector closest to the target object from at least two end effectors according to the position information of the target object and the relative coordinates.

[0196] In this step, it includes calculating the relative distance between each end effector and the target object according to the position information of the target object and the relative coordinates, and using the end effector with a closer relative distance as the target end effector.

[0197] S205B3: Obtain the relative angle of the target end effector with respect to the preset angle.

[0198] In this step, the preset angle can be the angle of the end effector in the initial state and can be preset by the staff. The relative angle can be obtained by subtracting the current angle from the preset angle.

[0199] S205B4: Determine the rotation angle according to the relative angle and the orientation information.

[0200] In this step, it includes subtracting the relative angle from the orientation information and then adding the preset angle to obtain the rotation angle.

[0201] S205B5: Use the position information, relative coordinates and rotation angle to determine the grasping control information of the target end effector.

[0202] In this step, it includes inputting the position information, relative coordinates, and relative angle into the control information template to obtain the grasping control information. It may also include generating a moving distance using the position information and relative coordinates, and generating the grasping control information using the moving distance and rotation angle.

[0203] As can be seen from the description of the above embodiments, the embodiments of the present disclosure obtain the relative coordinates of the end effector, combine the relative coordinates with the position information of the target object, select the end effector closest to the target object to perform the grasping action, and combine the current angle of the end effector to control the rotation angle of the end effector during grasping, making the success rate of grasping higher.

[0204] In some alternative embodiments, based on the above Figure 2 corresponding embodiments, between step S202 and step S204, there is also a step:

[0205] S220: According to the image information, determine whether the target object meets the preset conditions. The preset conditions are that the target object has multiple structural components, the projection areas of all structural components are mutually exclusive, and the grasping area of the target object is only located on one of the structural components.

[0206] S221: If the target object meets the above preset conditions, then crop from the first point cloud information to obtain the third point cloud information corresponding to the structural component containing the above grasping area.

[0207] Step S204 is correspondingly replaced with:

[0208] Determine the position information and orientation information of the target object according to the third point cloud information and status of the target object.

[0209] In the above step S220 of this embodiment, it can be judged by matching the target object category template or the preset database, or can be realized by a neural network or object recognition model obtained through pre-training. The matching situation between the structural components of each target object and the above preset conditions is recorded in the above target object category template or preset database.

[0210] For example, if the target object is a frying pan, the frying pan consists of a pan body and a pan handle. The pan handle has a grasping area, and the pan body does not have a grasping area. If the grasping point coordinates are calculated based on the coordinates of all discrete points in the point cloud corresponding to the frying pan, the calculation result will not be located in the grasping area, that is, on the pan handle, but in the pan body, resulting in a grasping failure. After determining the grasping area according to the image information in this embodiment, only the grasping point coordinates are calculated based on the point cloud corresponding to the grasping area, which is beneficial to improving the success rate of the humanoid robot grasping the target object, and it is more stable after grasping, especially suitable for static objects.

[0211] Among them, the above steps S220 and S221 can be located between S202 and S203, or between S203 and S204. This application does not limit this.

[0212] In some alternative embodiments, based on the corresponding embodiments above, Figure 2 between the above steps S202 and S204, there are further steps:

[0213] S223: Obtain the density information and volume information of the target object.

[0214] S224: According to the density information and volume information, determine the weight concentration area in the target object, and crop the fourth point cloud information corresponding to the weight concentration area from the first point cloud information.

[0215] Step S204 is correspondingly replaced with:

[0216] Determine the position information and orientation information of the target object according to the fourth point cloud information and state of the target object.

[0217] Among them, the above weight concentration area is the continuous area with the largest weight corresponding to a single substance inside the target object. For the determination process of the weight concentration area, reference can be made to the descriptions of the corresponding embodiments of the above steps S531 and S532.

[0218] In this step, for a static object, first obtain the point cloud information corresponding to the weight concentration part, and then only calculate the average value of the coordinates of all points within this part of the point cloud as the position information of the target object. For a dynamic object, through the first point cloud information corresponding to the image information, determine the position information in the way of a static object, determine the speed and acceleration by using the position information corresponding to each first point cloud information, and calculate the position information of the target object at the next moment by using the position information, speed, and acceleration.

[0219] For example, if the water in a mineral water bottle is not full, that is, there is only a part of the water, then determine the position information according to the part of the water. Another example is that the target object is a segmented object, and different segments are connected by a flexible material, and the heavier parts are in different segments, then use the heavier segment to determine the position information of the target object.

[0220] Based on this, this embodiment can improve the success rate and stability of a humanoid robot in grasping an object, especially an object with uneven weight distribution, and is especially suitable for static objects. On the other hand, it can reduce the calculation amount and improve the calculation efficiency.

[0221] Among them, the above steps S223 and S224 can be located between S202 and S203, or between S203 and S204. This application does not limit this.

[0222] It should be noted that all the above embodiments disclosed in this application can be freely combined, and the technical solutions obtained after combination are also within the protection scope of this application.

[0223] The above content is a further detailed description of this application in combination with specific preferred embodiments. It cannot be determined that the specific implementation of this application is only limited to these descriptions. For those of ordinary skill in the technical field to which this application belongs, without departing from the concept of this application, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of this application.

Claims

1. A humanoid robot control method based on point cloud clustering, characterized in that, The humanoid robot includes an end effector, and the method includes: Obtain at least two sets of consecutive image information about a target object; Obtain the target point cloud information corresponding to the image information, and perform clustering processing on the target point cloud information based on a target clustering algorithm to obtain the first point cloud information corresponding to the target object; Determine the state of the target object according to the image information; the state is static or dynamic; Determine the position information and orientation information of the target object according to the first point cloud information corresponding to the target object and the state; the calculation methods of the position information corresponding to the static target object and the dynamic target object are different; Determine the grasping control information of the end effector according to the position information and orientation information of the target object; Control the end effector to grasp the target object according to the grasping control information.

2. The method according to claim 1, characterized in that, The determining the position information and orientation information of the target object according to the first point cloud information corresponding to the target object and the state includes: If the target object is dynamic, associate the target object based on a target tracking algorithm; Model the motion process of the target object using a uniform motion model; Use a dynamic prediction algorithm to predict the position information of the target object at the next moment according to the state at the current moment.

3. The method according to claim 1, wherein The method includes: Obtain at least three sets of consecutive image information about a target object; Judge whether the target object is a dynamic moving object according to the at least three sets of consecutive image information; In the case where the target object is a dynamic moving object, predict the position information corresponding to the target object at the next moment according to the at least three sets of consecutive image information, and use it as the position information of the target object.

4. The method according to claim 1, characterized in that After controlling the end effector to grasp the target object according to the grasping control information, it further includes: If the grasping fails, obtain at least two sets of consecutive new image information about the target object; Use the new image information to determine whether the state of the target object has changed from static to dynamic; If the state of the target object has changed from static to dynamic, determine the new position information of the target object according to the new image information.

5. The method according to claim 4, wherein The robot includes at least two types of end effectors; After the step of if the grasping fails, obtain at least two sets of consecutive new image information about the target object, it further includes: Determine the environmental information where the target object is located according to the new image information; Generate the target grasping control information of the target type of end effector according to the environmental information and the new position information; Control the target type of end effector to grasp the target object using the target grasping control information.

6. The method according to claim 1, wherein The determining the grasping control information of the end effector according to the position information and orientation information of the target object includes: Obtain the current task information of the humanoid robot; Determine the candidate grasping points of the end effector according to the position information and orientation information of the target object; Filter out the target grasping points from the candidate grasping points according to the current task information; the grasping control information includes the target grasping points; Determine the grasping control information of the end effector according to the target grasping point.

7. The method according to claim 6, characterized in that, The screening of the target grasping point from the candidate grasping points according to the current task information includes: Obtain the density information and volume information of the target object; Screen the target grasping point from the candidate grasping points according to the current task information, the density information and the volume information.

8. The method according to claim 7, wherein The screening of the target grasping point from the candidate grasping points according to the current task information, the density information and the volume information includes: Obtain the joint limit information of the humanoid robot; Screen the target grasping point from the candidate grasping points according to the current task information, the density information, the volume information and the joint limit information of the humanoid robot.

9. The method according to any one of claims 1 to 8, characterized in that The determination of the position information and orientation information of the target object according to the first point cloud information corresponding to the target object and the state includes: Determine the position information of the target object according to the first point cloud information corresponding to the target object and the state; Perform principal component analysis on the first point cloud information corresponding to the target object to obtain the first sub-orientation information and the second sub-orientation information of the target object; Calculate the third sub-orientation information according to the first sub-orientation information and the second sub-orientation information; Generate the orientation information according to the first sub-orientation information, the second sub-orientation information and the third sub-orientation information.

10. The method according to any one of claims 1 to 8, characterized in that, The obtaining of the target point cloud information corresponding to the image information, and the clustering process of the target point cloud information based on the target clustering algorithm to obtain the first point cloud information corresponding to the target object includes: Obtain the initial point cloud information in the image information; Identify the bounding box of the target object in the image information; Crop the initial point cloud information with the bounding box to obtain the valid point cloud information; Determine the normal vector of each point in the valid point cloud information; If the angle between the normal vector of the target point in the valid point cloud information and the preset reference normal is less than the preset angle threshold, and the ordinate of the target point is within the preset ordinate range, then remove the target point from the valid point cloud information; Downsample the remaining valid point cloud information to obtain the target point cloud information; Perform clustering on the target point cloud information to obtain the first point cloud information corresponding to the target object.

Citation Information

Patent Citations

  • Dynamic target grabbing posture rapid detection method based on heterogeneous deep network fusion

    CN111598172A

  • Robot vision system

    CN111775154A