Object grabbing control method and system

The acquisition of object images through the camera device and the determination of the grab position using the attitude detection model is solved, and the problem of low efficiency and accuracy of robot grasping operations on objects of different shapes is achieved, achieving more efficient and flexible grasping operations.

CN119987284APending Publication Date: 2025-05-13SHENZHENSHI YUZHAN PRECISION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411985236.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-28
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When existing robot grasping operations face objects of different shapes and sizes, they have low efficiency and accuracy and insufficient flexibility.

Method used

Two-dimensional and three-dimensional images of the object are obtained through the camera device, the image feature points are extracted using the attitude detection model, the point cloud of the object is determined, and the transformation matrix is ​​calculated, multiple grasping positions on the object are predicted, the grasping posture and position of the manipulator are determined, and the grasping angle is optimized to determine the target grasping position.

Benefits of technology

It improves the accuracy and efficiency of the robot's grasping on objects of different shapes, reduces the action of adjusting operating parameters, and enhances the flexibility of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987284A_ABST
    Figure CN119987284A_ABST
Patent Text Reader

Abstract

The invention provides an object grabbing control method and system. The method comprises the steps that a first image and a second image are received; inputting the first image into a preset attitude detection model to obtain a plurality of image feature points; determining a point cloud corresponding to the object based on the plurality of image feature points; calculating a conversion matrix according to the point cloud and the model feature points of the second image; according to the conversion matrix, a plurality of first grabbing positions on the object are predicted through the virtual position marked on the object model by the manipulator model and the position of the camera device relative to the manipulator; a first grabbing posture, corresponding to each first grabbing position, of the manipulator is determined; based on the current posture of the robot, the grabbing angle corresponding to the first grabbing position corresponding to each first grabbing posture is determined; based on the grabbing angle, a target grabbing position is determined from the multiple first grabbing positions; and driving the robot to control the manipulator to grab the object at the target grabbing position. According to the method, the object grabbing efficiency and accuracy of the robot can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of automation control technology, and in particular to an object grasping control method and system. Background Art

[0002] Robotic grasping is a robot system that can simulate human arm movements and perform various tasks. The robot can effectively grasp, move and place objects by acquiring environmental information through cameras or sensors. At present, the robot's grasping operation depends on the pre-set robot first grasping posture or fixed first grasping position. When facing objects of different shapes and sizes, the operating parameters need to be adjusted accordingly, which affects the accuracy and efficiency of grasping and has low flexibility. Summary of the invention

[0003] The embodiments of the present application disclose an object grasping control method and system, which solve the technical problem of low efficiency and accuracy of current robot grasping operations when facing objects of various shapes.

[0004] The present application provides an object grasping control method, which is applied to a system including a robot and a camera device, and the method includes: receiving a first image and a second image sent by the camera device, wherein the first image is a two-dimensional image generated when the camera device photographs the object, and the second image is a three-dimensional image generated when the camera device photographs the object; inputting the first image into a preset posture detection model to obtain a plurality of image feature points; based on the plurality of image feature points, determining a point cloud corresponding to the object; calculating a transformation matrix according to the point cloud and the model feature points of the second image; and calculating a transformation matrix according to the transformation matrix and the robot arm model on the object. The virtual position marked on the model and the position of the camera device relative to the manipulator are used to predict multiple first grasping positions on the object, wherein the manipulator model is a three-dimensional model corresponding to the manipulator of the robot, and the object model is a three-dimensional model corresponding to the object; a first grasping posture of the manipulator corresponding to each first grasping position is determined; based on the current posture of the robot, a grasping angle corresponding to the first grasping position corresponding to each first grasping posture is determined; based on the grasping angle, a target grasping position is determined from the multiple first grasping positions; and the robot is driven to control the manipulator to grasp the object at the target grasping position.

[0005] In some embodiments of the present application, the first image is input into a preset posture detection model to obtain multiple image feature points, including: determining multiple contour points corresponding to the first image through the preset posture detection model; based on the occlusion information of each contour point, screening out multiple target contour points from the multiple contour points, the occlusion information of each target contour point indicating that it is not occluded; based on the multiple faces formed by the multiple contour points, determining the corresponding center point of each face; based on the multiple target contour points and the corresponding center point of each face, obtaining the multiple image feature points.

[0006] In some embodiments of the present application, determining the point cloud corresponding to the object based on the multiple image feature points includes: obtaining the depth value of each image feature point and the camera parameters of the camera device; obtaining a candidate point cloud corresponding to each image feature point based on the depth value, the camera parameters and the image coordinates corresponding to each image feature point; and denoising the candidate point cloud using a preset clustering algorithm to obtain a point cloud corresponding to the object.

[0007] In some embodiments of the present application, the transformation matrix is ​​calculated based on the point cloud and the model feature points of the second image, including: based on the preset numbers of the point cloud, determining the target model feature points with the preset numbers from the model feature points; based on the point cloud and the target model feature points with the same preset numbers, calculating the scaling matrix, the rotation matrix and the translation matrix; and obtaining the transformation matrix based on the scaling matrix, the rotation matrix and the translation matrix.

[0008] In some embodiments of the present application, the method also includes training the posture detection model, and the training method includes: obtaining training samples and corresponding sample labels; using the training samples to train the posture detection model; calculating the loss value of the posture detection model based on the prediction results output by the posture detection model and the sample labels; when the loss value is less than or equal to a preset threshold, obtaining a posture detection model that has completed training.

[0009] In some embodiments of the present application, the obtaining of training samples and corresponding sample labels includes: obtaining a three-dimensional training model for training, and a first skeleton point corresponding to the three-dimensional training model, the first skeleton point being represented by a preset sphere equation; processing the first skeleton point using a principal component analysis algorithm to obtain multiple eigenvalues ​​and an eigenvector corresponding to each eigenvalue; obtaining a second skeleton point based on the eigenvalue and the sphere equation; obtaining a third skeleton point corresponding to the three-dimensional model based on the second skeleton point and the eigenvector; using a two-dimensional image of the three-dimensional training model under multiple perspectives as the training sample; and using a two-dimensional image marked with occlusion information of the third skeleton point as the sample label.

[0010] In some embodiments of the present application, the method further includes: determining a corresponding second grasping posture of the manipulator at each second grasping position based on a plurality of preset second grasping positions; and determining a grasping angle of each second grasping position based on the current posture of the robot and each second grasping posture.

[0011] In some embodiments of the present application, the method further includes: determining the target grasping position from the plurality of second grasping positions based on a grasping angle corresponding to each second grasping position.

[0012] In some embodiments of the present application, determining a target grasping position from the plurality of first grasping positions based on the grasping angle includes: taking a first grasping position corresponding to a minimum grasping angle as the target grasping position.

[0013] The present application also provides an object grasping control system, which includes a robot and a camera device. The robot includes a manipulator, a processor and a memory. The processor is used to execute a computer program stored in the memory to implement the object grasping control method.

[0014] In the object grasping control method provided by the present application, a two-dimensional first image and a three-dimensional second image are obtained by using a camera device. Comprehensive and accurate object information can be obtained through the first image and the second image, providing data support for subsequent analysis of the object. A plurality of image feature points of the first image are extracted by using a posture detection model, so that a point cloud corresponding to the object is determined based on the plurality of image feature points, and the surface features of the object can be determined through the point cloud. Then, a transformation matrix is ​​calculated according to the point cloud and the model feature points of the second image, and the transformation matrix is ​​used to describe the linear transformation relationship between the point cloud and the model feature points. According to the transformation matrix, the virtual position marked on the object model by the manipulator model, and the position of the camera device relative to the manipulator, a plurality of first grasping positions on the object are predicted, so that the first grasping posture corresponding to each first grasping position of the manipulator can be determined, thereby improving the determination efficiency of the first grasping posture. Based on the current posture of the robot, the grasping angle corresponding to the first grasping position corresponding to each first grasping posture is determined; based on the grasping angle, the target grasping position is determined from the plurality of first grasping positions, thereby improving the calculation efficiency of the target grasping position. The robot is driven to control the manipulator to grasp the object at the target grasping position. The above method can reduce the actions of adjusting operating parameters when facing objects of different shapes, and improve the accuracy and efficiency of the robot in grasping objects. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a structural diagram of an object grasping control system provided in an embodiment of the present application.

[0016] Figure 2 It is a flow chart of the object grasping control method provided in an embodiment of the present application.

[0017] Figure 3 It is a skeleton diagram of the posture detection model provided in the embodiment of the present application.

[0018] Figure 4 It is a schematic diagram of skeleton points on an object provided in an embodiment of the present application.

[0019] Figure 5 It is a schematic diagram of the process of determining multiple image feature points provided in an embodiment of the present application.

[0020] Figure 6 It is a schematic diagram of multiple contour points and multiple target contour points provided in an embodiment of the present application.

[0021] Figure 7 This is a schematic diagram for determining the center point provided in an embodiment of the present application.

[0022] Figure 8 It is a schematic diagram of the second image provided in an embodiment of the present application.

[0023] Fig. 9 It is a schematic diagram of the point cloud and model feature points provided in an embodiment of the present application.

[0024] Fig.10 It is a schematic diagram of a manipulator model and an object model provided in an embodiment of the present application.

[0025] Fig.11 It is a schematic diagram of an application scenario of the object grasping control method provided in an embodiment of the present application.

[0026] Fig.12 It is a schematic diagram of grabbing an object at a target grabbing position provided in an embodiment of the present application.

[0027] Fig.13 It is a schematic diagram of a robot arm grasping an object provided in an embodiment of the present application.

[0028] Fig.14 It is a schematic diagram of the human body model provided in the embodiment of the present application.

[0029] Fig.15 It is a training flow chart of the posture detection model provided in the embodiment of the present application. DETAILED DESCRIPTION

[0030] To facilitate understanding, some illustrations of concepts related to the embodiments of the present application are given by way of example for reference.

[0031] It should be noted that in this application, "at least one" means one or more, and "more than one" means two or more than two. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and drawings of this application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0032] The current robot grasping operation relies on a pre-set first grasping posture of the robot, or a fixed first grasping position. When facing objects of different shapes and sizes, the operating parameters need to be adjusted accordingly, which affects the accuracy and efficiency of grasping and has low flexibility. In order to solve the technical problem of low efficiency and accuracy of the current robot grasping operation, the embodiment of the present application provides an object grasping control method and system, which can reduce the action of adjusting the operating parameters when grasping objects of different shapes, and improve the accuracy and efficiency of the robot grasping objects by analyzing the two-dimensional image (first image) and the three-dimensional image (second image) of the object.

[0033] In order to better understand the object grasping control method and system provided in the embodiments of the present application, the object grasping control system of the present application is first described below.

[0034] Figure 1 : is a structural diagram of an object grasping control system provided in an embodiment of the present application. The object grasping control method provided in an embodiment of the present application is applied to an object grasping control system 10, and the object grasping control system 10 includes a robot 110 and a camera device 120. The robot 110 and the camera device 120 are connected in communication. The communication connection mode includes wired network communication and wireless network communication. Among them, the wired network can be any one of a local area network, a metropolitan area network and a wide area network, and the wireless network can be any one of Bluetooth Technology, Wireless LAN (Wireless Fidelity, Wi-Fi), Near Field Communication (Near Field Communication, NFC), ZigBee Wireless Networks (ZigBee) technology, Infrared Data Association (Infrared Data Association, IrDA) technology, Ultra Wideband (UltraWideband, UWB) technology, wireless Universal Serial Bus (Universal Serial Bus, USB) and the like.

[0035] The robot 110 includes but is not limited to a manipulator 1101 , a memory 1102 , and a processor 1103 .

[0036] Among them, the manipulator 1101 can be a gripper, and the robot 110 can flexibly adjust the posture, angle and size of the manipulator 1101 to grasp, carry or manipulate objects.

[0037] The memory 1102 may include one or more random access memories (RAM) and one or more non-volatile memories (NVM). The random access memory can be directly read and written by the processor 1103, and can be used to store executable programs (such as machine instructions) of the operating system or other running programs, and can also be used to store user and application data. The random access memory may include static random-access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), etc.

[0038] The non-volatile memory may also store executable programs and user and application data, etc., and may be loaded into the random access memory in advance for direct reading and writing by the processor 1103. The non-volatile memory may include a disk storage device and a flash memory.

[0039] The memory 1102 is used to store one or more computer programs. The one or more computer programs are configured to be executed by the processor 1103. The one or more computer programs include multiple instructions. When the multiple instructions are executed by the processor 1103, the object grasping control method executed on the robot 110 can be implemented.

[0040] In other embodiments, the robot 110 further includes an external memory interface for connecting to an external memory to expand the storage capacity of the robot 110 .

[0041] The processor 1103 may include one or more processing units, for example, the processor 1103 may include an application processor (AP), a modem processor, a graphics processor (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.

[0042] The processor 1103 provides computing and control capabilities. For example, the processor 1103 is used to execute a computer program stored in the memory 1102 .

[0043] The camera device 120 may be a 3D (Dimension) camera.

[0044] In some embodiments of the present application, a first image and a second image of an object are acquired by a camera device 120, and the first image and the second image are processed by a processor 1103, so that a target grasping position of the manipulator 1101 of the robot 110 for grasping the object can be obtained.

[0045] In other embodiments of the present application, the object grasping control system 10 may also include electronic devices such as computers and servers for receiving images acquired by the camera device 120 and processing the images to control the robot 110 to perform operations of grasping objects.

[0046] The indication Figure 1 This is only an example of the object grasping control system 10 and does not constitute a limitation on the object grasping control system 10. It may include more or fewer components than shown in the figure, or combine certain components, or different components. In other embodiments of the present application, the object grasping control system 10 may also include electronic devices such as computers and servers, which are used to receive images captured by the camera device 120 and process the images to control the robot 110 to perform the operation of grasping the object.

[0047] Figure 2 is a flowchart of an object grasping control method provided in an embodiment of the present application, which is applied to a system including a robot and a camera device (eg Figure 1 According to different requirements, the order of the steps in the flowchart can be changed, and some steps can be omitted.

[0048] Step S201: receiving a first image and a second image sent by a camera device.

[0049] In some embodiments of the present application, the camera device may be a 3D camera, which generally includes: a laser transmitting end, a receiving end, a time to digital converter (TDC), and a control system. In practical applications, the camera device may be fixed in a position to prevent shaking caused by movement from affecting the quality of image acquisition.

[0050] When photographing an object using a camera, the camera may be calibrated to ensure shooting accuracy and stability. A conventional shooting function of the camera is used to capture a two-dimensional image of the object, and the two-dimensional image is preprocessed such as denoising and contrast enhancement to obtain a first image.

[0051] The depth sensing function of the camera device is used to capture the depth data of the object surface. The depth data is represented in the form of a point cloud, and each point cloud contains its position information in space. The depth data is processed, filtered, and registered to generate a three-dimensional model. The image corresponding to the three-dimensional model is used as the second image.

[0052] In addition to using the depth perception function of the camera device to obtain the second image, a three-dimensional second image can also be generated based on multiple two-dimensional images. Specifically, the angle of the object is changed, and the camera device is used to capture two-dimensional images of the object at different angles. Common feature points are extracted from the two-dimensional images at different angles. By comparing the position differences of the feature points in different two-dimensional images, the positions of these feature points in three-dimensional space can be calculated. Based on the calculated positions of the feature points, interpolation and other algorithms are used to fill the three-dimensional surface of the entire object and construct a three-dimensional model, so that the image corresponding to the three-dimensional model is used as the second image.

[0053] The above is only an example, and the present application does not limit the generation process of the second image.

[0054] Step S202: input the first image into a preset posture detection model to obtain a plurality of image feature points.

[0055] In some embodiments of the present application, the posture detection model can be yolov8n-pose, a skeleton recognition AI model. yolov8n-pose designs a universal skeleton, such as Figure 3 As shown, the structure of the skeleton is a spherical structure formed by 17 points, and a preset number is set for each point according to a preset numbering rule, such as Figure 3 The preset numbers shown are 1 to 17. Figure 4As shown in Figure 1, these 17 points will be evenly distributed on the object to achieve accurate posture recognition and positioning of the object. The training process of the posture detection model can refer to Fig.15 The embodiment shown. In the embodiment of the present application, a plurality of image feature points of a first image are determined by using a posture detection model. The process of determining the plurality of image feature points is as follows: Figure 5 As shown, the steps are as follows: Step S2021 to Step S2024:

[0056] Step S2021: Determine a plurality of contour points corresponding to the first image through a preset posture detection model.

[0057] In some embodiments of the present application, the first image is input into a posture detection model to obtain a bounding box and multiple contour points of the object. The number of the multiple contour points can be 17. Each contour point has three numerical values, each of which has a corresponding meaning. The first two numerical values ​​represent coordinates, and the last numerical value represents occlusion information, and the occlusion information includes unobstructed and obstructed. Unobstructed contour points represent points on the object that can be captured when the object is captured by a camera device to generate the first image, and belong to points on the object that are visualized in the first image. Obstructed contour points represent points on the object that are not captured when the object is captured by a camera device to generate the first image, and belong to points on the object that are not visible in the first image.

[0058] The occlusion information may be presented in the form of a percentage of occlusion. In one example, if the occlusion information is greater than or equal to 98%, it means that the probability of the contour point being occluded is 98%, and the contour point is marked as occluded. If the occlusion information is less than 98%, it means that the contour point is not occluded, and the contour point is marked as not occluded.

[0059] In one example, there is a contour point A (A1, A2, A3), then A1 and A2 represent the coordinates of the contour point in the image coordinate system, and A3 represents the occlusion information of each contour point. If the occlusion information is 98%, the contour point A is marked as occluded.

[0060] Step S2022: based on the occlusion information of each contour point, a plurality of target contour points are selected from the plurality of contour points.

[0061] In some embodiments of the present application, based on the occlusion information of each contour point, the contour point that is not occluded by the occlusion information is used as the target contour point, and multiple target contour points that are not occluded can be screened out from multiple contour points. Figure 6 As shown, Figure 6 (1) is a schematic diagram of multiple contour points and bounding boxes, such as Figure 6 (2) is a schematic diagram of multiple target contour points and bounding boxes. Figure 6After removing the occluded contour points based on the occlusion information, the following is obtained: Figure 6 Schematic diagram of (2).

[0062] Step S2023, based on the multiple surfaces formed by the multiple contour points, determine the corresponding center point of each surface.

[0063] In some embodiments of the present application, 17 contour points are obtained through the posture detection model. After removing the occluded contour points based on the occlusion information, the number of remaining unoccluded target contour points is about 8 or less. In order to avoid the influence of a small number of target contour points on the subsequent calculation accuracy, multiple faces are formed based on multiple contour points, and the center point of each face is determined. For example, 20 faces are formed based on 17 contour points, and the center points of the 20 faces are obtained. Figure 7 As shown, a skeleton based on multiple contour points forms multiple faces on the object, and for a triangular face, the center point of the triangle is determined. For a square face, two triangles are cut obliquely from the square face to obtain the center point of each triangle.

[0064] The number of multiple contour points can be increased from 17 to 47 based on the corresponding center point of each face.

[0065] Step S2024, obtaining a plurality of image feature points based on a plurality of target contour points and a center point corresponding to each face.

[0066] In some embodiments of the present application, when multiple target contour points and the corresponding center point of each face body are obtained, multiple target face bodies including multiple target contour points are obtained, and the center points corresponding to the multiple target face bodies are used as target center points. The multiple target contour points and the multiple target center points are used as multiple image feature points, and the number of points increases from 8 to 23.

[0067] In the above embodiment, the unobstructed target contour points on the object can be accurately obtained through the posture detection model, and the target center point is added on the basis of the determined target contour points, which can improve the precision and accuracy of subsequent calculations.

[0068] Step S203: determining a point cloud corresponding to the object based on the multiple image feature points.

[0069] In some embodiments of the present application, in order to better understand the object features such as the shape, size, and surface features of the object, the point cloud corresponding to the object can be determined by multiple image feature points. Specifically: obtain the depth value of each image feature point and the camera parameters of the camera device. The depth value represents the distance between the image feature point and the camera device, and the camera parameters include intrinsic parameters (such as focal length, optical center position) and extrinsic parameters (such as the position of the camera device). Obtain the image coordinates of each image feature point in the image coordinate system. Based on the depth value, camera parameters, and image coordinates, calculate the candidate point cloud corresponding to each image feature point.

[0070] Since some candidate point clouds may not be on the object, a preset clustering algorithm can be used to denoise the candidate point clouds to obtain the corresponding point clouds on the object. The clustering algorithm can be a machine learning algorithm with advantages such as fewer adjustment parameters and fast calculation rate, for example, a density-based clustering algorithm (such as DBSCAN) and a distance-based clustering algorithm (such as K-Means). This application does not limit the type of clustering algorithm.

[0071] Step S204, calculating a transformation matrix according to the point cloud and the model feature points of the second image.

[0072] In some embodiments of the present application, since each contour point carries a preset number set based on a preset numbering rule, the point cloud determined based on multiple contour points will also carry a corresponding preset number. Among them, the target center point included in the multiple image feature points will also carry a corresponding preset number, and the preset number of the target center point can be generated according to the determination order of the target center point, and the present application does not limit this.

[0073] The second image is a three-dimensional image generated when the camera device captures the object. According to the size of the object in the three-dimensional image, the principal component analysis algorithm is used to calculate the sphere surrounding the object and mark 47 spherical feature points on the surface of the sphere, such as Figure 8 As shown in FIG. 1 , a schematic diagram of an object including 47 spherical feature points is shown. Through 3D engine simulation, the spherical feature points are projected from the spherical surface into the center point of the object, and the intersection points of each spherical feature point and the spherical surface are determined, and these intersection points are used as model feature points on the object surface. The 47 model feature points are numbered according to a preset numbering rule, which is the same as the numbering rule corresponding to the preset numbering carried by the point cloud.

[0074] After the model feature points of the second image are determined, target model feature points with preset numbers are determined from the model feature points according to the preset numbers of the point cloud. Fig. 9 As shown, Fig. 9 (1) is a schematic diagram of the point cloud, such as Fig. 9 (2) is a schematic diagram of the model feature points. Based on the preset number of the point cloud, traverse Fig. 9 All model feature points corresponding to (2) are calculated until all model feature points with the same preset number as each point cloud are found. Based on the point cloud and target model feature points with the same number, the scaling matrix, rotation matrix and translation matrix are calculated, which are expressed as follows:

[0075] P Vision =S 3*3 *R 3*3 *P model +T;

[0076] Among them, P Vision Represents the coordinates of the point cloud (x Vision ,y Vision , z Vision ), S 3*3 represents a 3×3 scaling matrix, R 3*3 represents a 3×3 rotation matrix, T represents a translation matrix, P model Represents the coordinates of the model feature points (x model ,y model , z model ).

[0077] Typically, S x , S y and S z are the scaling factors along the x-axis, y-axis, and z-axis respectively, and the determinant of the scaling matrix det(S) is not necessarily equal to 1. The scaling matrix and rotation matrix are denoted as The translation matrix is ​​written as After expanding the formulas for calculating the scaling matrix, rotation matrix, and translation matrix, we get the following formula:

[0078] x Vision =ax model +by model +cz model +j;

[0079] y Vision =dx model +ey model +fz mode +k;

[0080] z Vision =gx model +hy model +iz model +l;

[0081] Use the least square solution to find abcj and use the following formula for calculation:

[0082]

[0083] The same algorithm is used to calculate defk and ghil, so that the scaling matrix, rotation matrix and translation matrix can be obtained. Based on the scaling matrix, rotation matrix and translation matrix, the transformation matrix is ​​calculated. as follows:

[0084]

[0085] The rotation matrix is ​​used to describe the transformation relationship between the point cloud and the model feature points.

[0086] Step S205 , predicting a plurality of first grasping positions on the object according to the transformation matrix, the virtual positions marked on the object model by the manipulator model, and the position of the camera device relative to the manipulator.

[0087] In some embodiments of the present application, a manipulator model of a manipulator may be pre-stored in the object grasping control system, and the manipulator model is a three-dimensional model. After acquiring the second image containing the object model, the manipulator model may be driven to grasp the object model around the object model to generate a virtual position. Based on the virtual position, the position that can be grasped with the object mark is a three-dimensional model. Fig.10 As shown in (1), it is a robot model. Fig.10 As shown in (2), it is the object model.

[0088] In practical applications, the robot and the camera device are fixed at a certain position. Fig.11 As shown, the robot 110 includes a manipulator 1101, and the robot 110 and the camera device 120 are fixed at a position. The virtual position of the manipulator model of the robot 110 marked on the object model is obtained. The virtual position can include the 360° position of the object. Get the position of the camera device relative to the manipulator Predict multiple first grasping positions on the object according to the transformation matrix, the virtual position of the manipulator model marked on the object model, and the position of the camera device relative to the manipulator.

[0089] The grasping position is the position in the robot coordinate system. It can be expressed as follows:

[0090]

[0091] in, represents the virtual position of the robot model marked on the object model, represents the transformation matrix, Indicates the position of the camera device relative to the robot, Indicates the first grasp position.

[0092] Step S206, determining a first grasping posture of the manipulator corresponding to each first grasping position.

[0093] In some embodiments of the present application, the first grasping posture represents the posture of the manipulator when the robot grasps the object through the first grasping position. According to the posture of the manipulator model when grasping the virtual object at the virtual position, the first grasping posture corresponding to the manipulator at each first grasping position is determined.

[0094] Step S207: determining a grasping angle corresponding to a first grasping position corresponding to each first grasping posture based on the current posture of the robot.

[0095] In some embodiments of the present application, the current posture of the robot and the first grasping posture are presented in quaternions, which are an extended complex number system used in mathematics and computer science to represent rotations. The grasping angle is calculated using the following formula:

[0096] θ Different =2*cos(|q robot ·q grip |);

[0097] Among them, θ Different represents the grasping angle, q robot represents the current posture of the robot, q grip Indicates the first grasping posture of the manipulator when grasping an object at the first position.

[0098] Step S208: determining a target grasping position from a plurality of first grasping positions based on the grasping angle.

[0099] In some embodiments of the present application, in order to reduce the movement of the robot arm corresponding to the robot hand, the smallest gripping angle is selected from multiple gripping angles. Different The corresponding first grasping position is taken as the target grasping position. Fig.12 As shown, based on the multiple first grasping positions determined by the manipulator model, a target grasping position is determined from the multiple first grasping positions, and the manipulator 1101 is driven to grasp the object at the target grasping position.

[0100] Step S209: driving the robot to control the manipulator to grasp the object at the target grasping position.

[0101] In some embodiments of the present application, Fig.12 As shown, after the target grasping position is determined, the robot can be driven to control the manipulator 1101 to grasp the object at the target grasping position. Fig.13 As shown, Fig.13 (1) is the initial posture of the robot's manipulator grasping the object, such as Fig.13(2) is the posture of the robot grasping the object based on the target grasping position.

[0102] In addition, in practical applications, the objects can also be human models of various shapes, such as Fig.14 After the target grasping position is calculated through the above embodiment, the human body model can be grasped accurately and effectively through the target grasping position. Fig.14 The human body model shown.

[0103] In other embodiments of the present application, in order to reduce the amount of calculation, it is possible not to mark the virtual position on the object model through the manipulator model, but to pre-set multiple second grasping positions based on empirical values. Such objects for which the second grasping position is set based on empirical values ​​may be objects with relatively simple structures or objects that have been grasped in the past. The present application does not limit the type of objects for which the second grasping position is set based on empirical values.

[0104] The corresponding second grasping posture of the manipulator can be determined according to the multiple second grasping positions, and the grasping angle of each second grasping position can be determined according to the current posture of the robot and each second grasping posture. Thus, the minimum grasping angle can be determined from the grasping angles of each second grasping position, and the second grasping position corresponding to the lowest grasping angle among the grasping angles corresponding to each second grasping position is determined, that is, the target grasping position.

[0105] In the above embodiment, a camera device is used to obtain a two-dimensional first image and a three-dimensional second image. Comprehensive and accurate object information can be obtained through the first image and the second image, providing data support for subsequent analysis of the object. A plurality of image feature points of the first image are extracted using a posture detection model, so that a point cloud corresponding to the object is determined based on the plurality of image feature points, and the surface features of the object can be determined through the point cloud. Then, a transformation matrix is ​​calculated based on the point cloud and the model feature points of the second image, and the transformation matrix is ​​used to describe the linear transformation relationship between the point cloud and the model feature points. According to the transformation matrix, the virtual position marked on the object model by the manipulator model, and the position of the camera device relative to the manipulator, a plurality of first grasping positions on the object are predicted, so that the first grasping posture corresponding to each first grasping position of the manipulator can be determined, thereby improving the efficiency of determining the first grasping posture. Based on the current posture of the robot, the grasping angle corresponding to the first grasping position corresponding to each first grasping posture is determined; based on the grasping angle, the target grasping position is determined from the plurality of first grasping positions, thereby improving the calculation efficiency of the target grasping position. The driving robot controls the manipulator to grasp the object at the target grasping position. The above method can reduce the actions of adjusting operating parameters when facing objects of different shapes, and improve the accuracy and efficiency of the robot grasping objects. In addition, through the above embodiment, different objects in different environments can be accurately grasped.

[0106] Fig.15 is a training flow chart of the posture detection model provided in the embodiment of the present application, such as Fig.13 As shown, the following steps are included:

[0107] Step S1501, obtaining training samples and corresponding sample labels.

[0108] In some embodiments of the present application, a three-dimensional training model for training is obtained. The three-dimensional training model may be a three-dimensional model corresponding to objects of various shapes, for example, a three-dimensional model of a cylinder. A first skeleton point corresponding to the three-dimensional training model is obtained. The structure of the first skeleton point is as follows: Figure 3 The first skeleton point can be represented by a preset sphere equation as follows:

[0109] Let r be the radius of the sphere, i represent the horizontal coordinate of the skeleton point, and the value range of i is pre-set to 1-3 according to experience, and j represent the vertical coordinate of the skeleton point, and the value range of j is pre-set to 0-4 according to experience. Figure 3 The coordinates of the first point shown are x=0, y=0, z=r.

[0110] The coordinates of points 2 to 16 are:

[0111]

[0112] The coordinates of the 17th point are: x=0, y=0, z=-r.

[0113] In some embodiments of the present application, after determining the first skeleton point, the first skeleton point is processed using a principal component analysis algorithm to obtain multiple eigenvalues ​​and an eigenvector corresponding to each eigenvalue. The eigenvalues ​​are sorted in descending order and recorded as λ 1 ,λ 2 ,λ 3 . The eigenvector is recorded as a 3×3 matrix [v 1 , v 2 , v 3 ].

[0114] According to the eigenvalue and the sphere equation, the second skeleton point is obtained as follows:

[0115] Let a = 2*λ 1 , b=2*λ 2 , c=2*λ 3 , then the coordinates of the first point are x=0, y=0, z=a.

[0116] The coordinates of points 2 to 16 are:

[0117]

[0118] The coordinates of the 17th point are:

[0119] x=0, y=0, z=-a.

[0120] In some embodiments of the present application, after the second skeleton point is determined, the second skeleton point is multiplied by the feature vector to obtain a third skeleton point, wherein the first skeleton point is a point on the periphery surrounding the three-dimensional training model, and after processing, the third skeleton point obtained is a point on the three-dimensional training model.

[0121] Calculate the third skeleton point (x new ,y new , z new ) is as follows:

[0122]

[0123] In some embodiments of the present application, after the third skeleton point of the three-dimensional training model is determined, two-dimensional images of the three-dimensional training model under multiple viewing angles are obtained as training samples. Since the two-dimensional image cannot display all the third skeleton points, the third skeleton points on each two-dimensional image carry corresponding occlusion information. The occlusion information is used to determine whether the third skeleton point is visible in the two-dimensional image. The occlusion information includes unoccluded and occluded. The two-dimensional image marked with the occlusion information of the third skeleton point is used as a sample label.

[0124] Step S1502: training the posture detection model using training samples.

[0125] Step S1503, calculating the loss value of the posture detection model according to the prediction result output by the posture detection model and the sample label.

[0126] In some embodiments of the present application, the prediction result is the prediction result output by the forward propagation of each posture detection model during the training process. According to the prediction result output by the posture detection model and the sample label, the loss value of the posture detection model is calculated to determine whether the posture detection model has reached a convergence state during the training process.

[0127] Step S1504, when the loss value is less than or equal to the preset threshold, a posture detection model that has completed training is obtained.

[0128] In some embodiments of the present application, when the loss value is greater than a preset threshold, it indicates that the posture detection model has not reached a convergence state, and training continues along the back propagation process. When the loss value is less than or equal to the preset threshold, it indicates that the posture detection model has completed training, and it is determined that the posture detection model has reached a convergence state, and the training of the posture detection model is completed.

[0129] In the above embodiment, since the posture detection model is composed of multiple skeleton points, and each skeleton point carries a preset number, the posture detection model can be used to determine the preset numbers of multiple contour points to be grasped in the subsequent use, so as to predict the complete posture of the object to be grasped in space through the preset numbers. This improves the efficiency and accuracy of grasping objects to a certain extent.

[0130] The embodiment of the present application also provides a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. The processor of the server reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the server executes the data extraction method described above in the embodiment of the present application.

[0131] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0132] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0133] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. An object grasping control method, applied to a system including a robot and a camera device, characterized in that: The method comprises: receiving a first image and a second image sent by the camera device, wherein the first image is a two-dimensional image generated when the camera device photographs an object, and the second image is a three-dimensional image generated when the camera device photographs the object; Inputting the first image into a preset posture detection model to obtain a plurality of image feature points; Determining a point cloud corresponding to the object based on the multiple image feature points; Calculating a transformation matrix according to the point cloud and the model feature points of the second image; predicting a plurality of first grasping positions on the object according to the transformation matrix, the virtual positions marked on the object model by the manipulator model, and the position of the camera device relative to the manipulator, wherein the manipulator model is a three-dimensional model corresponding to the manipulator of the robot, and the object model is a three-dimensional model corresponding to the object; Determining a first grasping posture of the manipulator corresponding to each first grasping position; Based on the current posture of the robot, determining a grasping angle corresponding to a first grasping position corresponding to each first grasping posture; Based on the grasping angle, determining a target grasping position from the plurality of first grasping positions; The robot is driven to control the manipulator to grasp the object at the target grasping position.

2. The object grasping control method according to claim 1, characterized in that: The step of inputting the first image into a preset posture detection model to obtain a plurality of image feature points includes: Determining a plurality of contour points corresponding to the first image by using a preset posture detection model; Based on the occlusion information of each contour point, a plurality of target contour points are selected from the plurality of contour points, and the occlusion information of each target contour point indicates that the target contour point is not occluded; Based on the multiple faces formed by the multiple contour points, determine the corresponding center point of each face; The plurality of image feature points are obtained based on the plurality of target contour points and the center point corresponding to each of the face bodies.

3. The object grasping control method according to claim 1, characterized in that: The step of determining a point cloud corresponding to the object based on the plurality of image feature points includes: Obtaining the depth value of each image feature point and the camera parameters of the camera device; Obtaining a candidate point cloud corresponding to each image feature point according to the depth value, the camera parameters, and the image coordinates corresponding to each image feature point; The candidate point cloud is denoised using a preset clustering algorithm to obtain a point cloud corresponding to the object.

4. The object grasping control method according to claim 1, characterized in that: The step of calculating a transformation matrix according to the point cloud and the model feature points of the second image includes: Based on the preset number of the point cloud, determining a target model feature point having the preset number from the model feature points; Calculate the scaling matrix, rotation matrix and translation matrix based on the point cloud and target model feature points with the same preset number; The conversion matrix is ​​obtained according to the scaling matrix, the rotation matrix and the translation matrix.

5. The object grasping control method according to claim 1, characterized in that: The method further includes training the posture detection model, and the training method includes: Obtain training samples and corresponding sample labels; Using the training samples to train the posture detection model; Calculating a loss value of the posture detection model according to the prediction result output by the posture detection model and the sample label; When the loss value is less than or equal to a preset threshold, a posture detection model that has completed training is obtained.

6. The object grasping control method according to claim 5, characterized in that: The obtaining of training samples and corresponding sample labels includes: Acquire a three-dimensional training model for training, and a first skeleton point corresponding to the three-dimensional training model, wherein the first skeleton point is represented by a preset sphere equation; Processing the first skeleton point using a principal component analysis algorithm to obtain a plurality of eigenvalues ​​and an eigenvector corresponding to each eigenvalue; Based on the eigenvalue and the sphere equation, obtaining a second skeleton point; Obtaining a third skeleton point corresponding to the three-dimensional model according to the second skeleton point and the feature vector; Using two-dimensional images of the three-dimensional training model under multiple viewing angles as the training samples; The two-dimensional image marked with the occlusion information of the third skeleton point is used as the sample label.

7. The object grasping control method according to claim 1, characterized in that: The method further comprises: Based on a plurality of preset second grasping positions, determining a corresponding second grasping posture of the manipulator at each second grasping position; Based on the current posture of the robot and each second grasping posture, a grasping angle of each second grasping position is determined.

8. The object grasping control method according to claim 7, characterized in that: The method further comprises: The target grasping position is determined from the plurality of second grasping positions based on the grasping angle corresponding to each second grasping position.

9. The object grasping control method according to claim 1, characterized in that: The step of determining a target grasping position from the plurality of first grasping positions based on the grasping angle comprises: The first grasping position corresponding to the minimum grasping angle is used as the target grasping position.

10. An object grasping control system, comprising a robot and a camera device, wherein the robot comprises a manipulator, a processor and a memory, wherein: The processor is used to execute a computer program stored in the memory to implement the object grasping control method according to any one of claims 1 to 9.

Citation Information

Cited By

  • Humanoid robot intelligent grabbing control system and method, electronic equipment and storage medium

    CN121223807A

  • Robot grabbing control method and device

    CN122274992A