Grabbing point prediction model training method and object grabbing point determination method and device
By introducing image feature extraction, instance feature generation, and grasp point generation models into the neural network model, and adjusting the model using a loss function, the problem of low training efficiency in existing technologies is solved, and efficient and accurate grasp point prediction is achieved.
Patent Information
- Application Number
- CN202310015157.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-04
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2043-01-04
AI Technical Summary
Existing neural network training schemes need to be improved in terms of training efficiency, especially in the fields of computer vision technology and intelligent robot grasping. In existing schemes, the calculation of 2D grasping points requires accurate instance segmentation masks, and incorrect segmentation masks will lead to incorrect grasping point predictions.
By acquiring training images and inputting them into the target neural network model, a loss function is generated using an image feature extraction model, an instance feature generation model, and a grasp point generation model to adjust the neural network model, reduce the prediction search range of grasp points, and directly output 2D grasp points from RGB images, thus simplifying the prediction process.
It improves the training efficiency and accuracy of the grasp point prediction model, reduces the impact of erroneous intermediate results on 2D grasp point prediction, and enables accurate acquisition of object grasp points from RGB images.
Smart Images

Figure CN115984668B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer vision, and in particular to a grasping point prediction model training method and an object grasping point determination method and device. BACKGROUND
[0002] This section is intended to provide background information to facilitate an understanding of embodiments of the application as set forth in the claims. The description herein does not constitute an admission that any of the information provided herein is prior art.
[0003] In the field of computer vision technology and intelligent robot grasping, a grasping point can be obtained by training a deep neural network model. At present, the training efficiency of existing neural network training schemes needs to be improved. SUMMARY
[0004] The present application provides a grasping point prediction model training method and device to at least solve the problem of improving model training efficiency.
[0005] According to an aspect of the present application, a grasping point prediction model training method is provided, comprising: obtaining a training picture; the training picture includes one or more objects and two-dimensional grasping point position labeling information of the objects; inputting the training picture into an image feature extraction model of a target neural network model to obtain image feature information; the target neural network model includes the image feature extraction model, an instance feature generation model and a grasping point generation model; inputting the image feature information into the instance feature generation model to obtain instance feature information; generating position reference point information of the instance feature information; inputting the instance feature information into the grasping point generation model to obtain predicted grasping point information, and generating a loss function using the predicted grasping point information, the two-dimensional grasping point position labeling information and the position reference point information; adjusting the target neural network model using the loss function to obtain a grasping point prediction model.
[0006] According to another aspect of the present application, an object grasping point determination method is provided, comprising: obtaining a target picture; inputting the target picture into a grasping point prediction model to obtain a two-dimensional grasping point of an object in the target picture; wherein the grasping point prediction model is obtained by training the above method.
[0007] According to another aspect of the present application, a grasp point prediction model training apparatus is provided, comprising: a data module configured to obtain a training picture; the training picture comprising one or more objects and two-dimensional grasp point position annotation information of the objects; an image module configured to input the training picture into an image feature extraction model of a target neural network model to obtain image feature information; the target neural network model comprising the image feature extraction model, an instance feature generation model and a grasp point generation model; an instance module configured to input the image feature information into the instance feature generation model to obtain instance feature information; a reference point module configured to generate position reference point information of the instance feature information; a grasp point module configured to input the instance feature information into the grasp point generation model to obtain predicted grasp point information, and generate a loss function using the predicted grasp point information, the two-dimensional grasp point position annotation information and the position reference point information; and a training module configured to adjust the target neural network model using the loss function to obtain a grasp point prediction model.
[0008] According to another aspect of the present application, an object grasp point determination apparatus is provided, comprising: an acquisition module configured to obtain a target picture; and a determination module configured to input the target picture into a grasp point prediction model to obtain a two-dimensional grasp point of an object in the target picture; wherein the grasp point prediction model is obtained by training the method described above.
[0009] According to another aspect of the present application, an electronic device is also provided, comprising: a processor; and a memory storing a program, wherein the program comprises instructions that, when executed by the processor, cause the processor to perform the method described above.
[0010] According to another aspect of the present application, a non-transitory computer-readable storage medium storing computer instructions is also provided, wherein the computer instructions are configured to cause the computer to perform the method steps described above.
[0011] In the embodiment of the present application, the image feature extraction model of the target neural network model is input with the training picture including one or more objects and the two-dimensional grasping point position labeling information of the objects, to obtain image feature information, the target neural network model including the image feature extraction model, an instance feature generation model and a grasping point generation model; the image feature information is input into the instance feature generation model to obtain instance feature information; position reference point information of the instance feature information is generated; the instance feature information is input into the grasping point generation model to obtain predicted grasping point information, and a loss function is generated by using the predicted grasping point information, the two-dimensional grasping point position labeling information and the position reference point information; the target neural network model is adjusted by using the loss function to obtain a grasping point prediction model. Through the position reference point information of the instance feature information, the embodiment of the present application reduces the prediction search range of the grasping point, and improves the efficiency of the grasping point prediction model training. BRIEF DESCRIPTION OF DRAWINGS
[0012] In the following description of the exemplary embodiments in conjunction with the accompanying drawings, more details, features and advantages of the present application are disclosed, in which:
[0013] Figure 1 A flowchart of a grasping point prediction model training method according to an exemplary embodiment of the present application is shown;
[0014] Figure 2 A schematic diagram of the overall flow of a grasping point prediction model according to an exemplary embodiment of the present application is shown;
[0015] Figure 3 A schematic diagram of a grasping point generation module according to an exemplary embodiment of the present application is shown;
[0016] Figure 4 A schematic diagram of multi-task parallel training according to an exemplary embodiment of the present application is shown;
[0017] Figure 5 A structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present application is shown;
[0018] Figure 6 A structural block diagram of a grasping point prediction model training device according to an exemplary embodiment of the present application is shown. DETAILED DESCRIPTION
[0019] Embodiments of the present application will be described herein below with reference to the drawings. Although certain embodiments of the present application are shown in the drawings, it is understood that the present application can be implemented in various forms and should not be interpreted as being limited to the embodiments set forth herein, but rather these embodiments are provided so as to more thoroughly and completely understand the present application. It is understood that the drawings and embodiments of the present application are for exemplary purposes only and are not intended to limit the scope of protection of the present application.
[0020] It should be understood that each of the steps recited in the method embodiments of the present application can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit performing the steps shown. The scope of the present application is not limited in this respect.
[0021] The term "comprising" and variations thereof as used herein are open-ended, that is "including but not limited to". The term "based on" is "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related terms are defined as follows. It should be noted that references made in the present application to "first", "second", etc. concepts merely serve to distinguish different apparatuses, modules or units from each other, and are not intended to limit the order or interdependence of the functions performed by these apparatuses, modules or units.
[0022] It should be noted that the modification of "one" or "multiple" mentioned in the present application is illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise explicitly indicated in the context, it should be understood as "one or more".
[0023] The names of the messages or information exchanged between the plurality of apparatuses in the embodiments of the present application are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0024] In the existing grasp point determination scheme, the image to be processed is input into a deep learning network, and a segmentation mask of the instance in the image is output. The obtained mask is subjected to a series of post-processing to improve the accuracy of the mask. The corrected mask is subjected to morphological processing to obtain a 2D grasp point position, such as calculating the center of the corrected mask. In combination with other information (such as depth, point cloud, etc.), a 3D grasp point result is obtained according to the 2D grasp point position.
[0025] In the existing scheme, the calculation of the 2D grasp point requires an accurate instance segmentation mask, so after obtaining the segmentation mask of the object by the deep neural network, a series of correction steps need to be performed on the mask result. The prediction of the 2D grasp point depends on the shape of the mask, and an incorrect segmentation mask will lead to incorrect prediction of the grasp point.
[0026] Based on this, the application provides a grasping point prediction model training method, an object grasping point determination method and device. The grasping point prediction model training method can provide training efficiency and accuracy of the model, and the obtained grasping point prediction model can accurately obtain the grasping point of each object in the image through two-dimensional input (RGB picture). For each object (instance) in the input image, a corresponding 2D grasping point is predicted, and then a 3D grasping point can be obtained through simple post-processing.
[0027] Firstly, the terms involved are explained.
[0028] Grasping point: refers to the best grasping position point or the center point of the best grasping area of an object to be grasped in an intelligent robot grasping system under the system conditions.
[0029] Depth map: a kind of data describing the distance (depth) information of objects in space to the camera.
[0030] Point cloud: a kind of data describing the position (three-dimensional coordinates) of points on the surface of an object in space.
[0031] Deep learning network: a kind of deep neural network model based on convolution and the like.
[0032] NMS (Non-Maximum Suppression, Non-Maximum Suppression): a kind of detection result post-processing method, which removes repeated detected results according to the scores of the detection boxes.
[0033] According to an aspect of an embodiment of the application, a grasping point prediction model training method is provided, Figure 1 The flowchart of the grasping point prediction model training method provided by the embodiment of the application is shown in Figure 1 The method comprises the following steps:
[0034] Step S202, obtaining a training picture; the training picture comprises one or more objects and two-dimensional grasping point position labeling information of the objects;
[0035] In this step, the training picture can be an RGB (a kind of color standard in the industry) image. The training picture comprises one or more objects and two-dimensional grasping point position labeling information of the objects, wherein the two-dimensional grasping point position labeling information is pre-labeled in the training picture, and each object corresponds to one grasping point position labeling information.
[0036] Step S204, inputting the training picture into an image feature extraction model of a target neural network model to obtain image feature information; the target neural network model comprises the image feature extraction model, an instance feature generation model and a grasping point generation model;
[0037] In this step, the training picture is input into a first model of the target neural network model, i.e., an image feature extraction model, to obtain image feature information. The target neural network model includes the image feature extraction model, an instance feature generation model, and a grasp point generation model, the output of the image feature extraction model is input into the instance feature generation model, and the output of the instance feature generation model is input into the grasp point generation model.
[0038] It should be noted that the image feature extraction model can be any deep neural network structure. Generally speaking, in an application scenario with low efficiency requirement, the image feature extraction model uses a network model with large amount of calculation, so as to obtain higher accuracy. In a scenario with relatively high efficiency requirement, the image feature extraction model uses a network model with small amount of calculation, so that the efficiency of the overall scheme is higher, but the prediction accuracy is slightly lost. The specific structure used can be determined according to the actual application requirements of efficiency and accuracy, and the embodiments of the present application do not make specific limitations.
[0039] In step S206, the image feature information is input into the instance feature generation model to obtain instance feature information.
[0040] In this step, the image feature information is input into the instance feature generation model, the instance feature generation model identifies the instances in the image, and outputs the features corresponding to each instance.
[0041] In an optional embodiment, in the instance feature generation model, the image feature information is first input into a deep learning network to obtain the score of each point on the feature map, i.e., the probability that the current point exists an object. According to the score, the K features with the highest scores are selected, the detection frame and score corresponding to these results are predicted, and then the final M results (M<=K) are obtained after NMS post-processing. According to the results, M features are selected from the K features as the output of the instance feature generation module.
[0042] In step S208, position reference point information of the instance feature information is generated.
[0043] In this step, the position reference point refers to a point that can reflect the position characteristics of the instance. This reference point can be a value that can be directly obtained, or a value that is calculated based on the instance feature information.
[0044] Directly predicting the position of a grasp point in the full image range has a large search range of prediction results, and it is difficult to obtain a relatively accurate grasp point position. Considering that the two-dimensional grasp point of the instance must be on the surface of the instance, a reference point related to the instance can be determined to reduce the search space of the grasp point prediction.
[0045] Step S210, input the instance feature information into the grasp point generation model to obtain predicted grasp point information, and generate a loss function by using the predicted grasp point information, the two-dimensional grasp point position label information and the position reference point information.
[0046] In this step, based on the feature of each instance, the grasp point generation model outputs the 2D grasp point coordinates (x, y) of the object, i.e., the position of the 2D grasp point in the RGB image. For each instance feature, the grasp point generation model is used to predict its grasp point. A loss function is generated by using the predicted grasp point information, the two-dimensional grasp point position label information and the position reference point information.
[0047] Step S212, adjust the target neural network model by using the loss function to obtain a grasp point prediction model.
[0048] In this step, the parameters of the target neural network are adjusted by minimizing the loss function to obtain a grasp point prediction model. By using the grasp point prediction model, the 2D grasp point of the instance can be directly output from the RGB image without additional processing steps, and the overall scheme is more efficient and simple.
[0049] In the embodiment of the present application, an image feature extraction model of a target neural network model is input with a training picture including one or more objects and two-dimensional grasp point position label information of the objects to obtain image feature information, the target neural network model includes the image feature extraction model, an instance feature generation model and a grasp point generation model; the image feature information is input into the instance feature generation model to obtain instance feature information; position reference point information of the instance feature information is generated; the instance feature information is input into the grasp point generation model to obtain predicted grasp point information, and a loss function is generated by using the predicted grasp point information, the two-dimensional grasp point position label information and the position reference point information; the target neural network model is adjusted by using the loss function to obtain a grasp point prediction model. The position reference point information of the instance feature information reduces the prediction search range of the grasp point and improves the efficiency of the grasp point prediction model training.
[0050] Considering that the training of a deep learning network needs a certain amount of data to support, a 2D (two-dimensional) grasp point calibration method is proposed in this scheme. In an optional implementation, a training picture can be obtained, and the following steps can be performed:
[0051] Obtaining an RGB picture and depth map data corresponding to the RGB picture, generating first point cloud data of the RGB picture by using the RGB picture and the depth map data; obtaining a calibration instruction, generating a segmentation mask of an object in the RGB picture according to the calibration instruction, determining second point cloud data of the object by using the segmentation mask and the first point cloud data; receiving a grabbing point generation instruction, determining a three-dimensional grabbing point of the object according to the grabbing point generation instruction and the second point cloud data, labeling a two-dimensional grabbing point position of the object in the RGB picture according to the three-dimensional grabbing point, and obtaining a training picture.
[0052] In the optional implementation, the calibration instruction can be manually controlled to be sent, and the position and shape of the segmentation mask and other attributes can be determined by using the calibration instruction. The grabbing point generation instruction can be manually controlled to be sent according to actual requirements, or selected in a pre-set instruction. A batch of RGB pictures are calibrated by using the above steps, so that a training picture with object grabbing point labeling is obtained.
[0053] In specific implementation, for example, a batch of RGBD data, that is, RGB pictures and corresponding depth map data, can be obtained based on an RGBD camera. Corresponding point cloud data can be obtained according to the RGB image and the depth data. Next, the RGB pictures are calibrated by instances, and the segmentation mask corresponding to each instance in the picture can be obtained by manual calibration. In combination with the instance segmentation mask in the RGB image and the point cloud information, the point cloud of the corresponding instance, that is, the information in the 3D space, can be obtained. According to the preference of the current grabbing system, the generation strategy of the grabbing point is specified, and then the 3D grabbing point corresponding to each instance is generated. For example, for some cuboid objects, the center of the upper surface is specified as the grabbing point, and the center (3D position) of the upper surface of the object can be obtained through the point cloud information of the object, and then the 2D grabbing point of the current instance in the RGB image can be obtained by mapping the center to the RGB image.
[0054] Based on the RGBD information, the 3D grabbing point of the object is calculated and projected into the 2D image, and the 2D grabbing point in the RGB image is labeled, so that the data preparation for directly predicting the 2D grabbing point by using the deep learning network is realized. A large amount of labeled data with 2D grabbing point true value is generated, and a large amount of data is the basis for training the deep learning network, so that the prediction of the 2D grabbing point and other functions related to the 2D grabbing point can be effectively trained.
[0055] In an optional implementation, the position reference point information of the instance feature information can be generated according to the following steps: receiving a reference point generation instruction, and determining reference point information by using the reference point generation instruction and the instance feature information; the reference point generation instruction is determined based on the feature position of the instance feature information.
[0056] In the optional embodiment, the position reference point can be a feature position of the instance feature, such as a coordinate position (x, y) corresponding to the instance feature on the feature map, which is a value that can be directly obtained.
[0057] In an optional embodiment, the position reference point information of the instance feature information can be generated according to the following steps: generating reference point information by using a reference point generation model; the reference point generation model is obtained by training a convolutional network by using training data; and the training data includes reference point position information.
[0058] In the optional embodiment, the position reference point can be a center point of a detection box corresponding to the instance, or a centroid of an instance mask, etc., which can be obtained by prediction. In order to predict the position reference point by using the network, an additional reference point generation model needs to be added after the instance feature generation model, as shown in FIG. 6. Figure 3 The reference point generation submodule in the figure is used to perform data calculation related to the reference point generation model. The model can be composed of a simple convolutional network with several layers, and the output of the model needs to be supervised by the reference point position (such as the center of the detection box) of the training data.
[0059] In an optional embodiment, the loss function can be generated by using the predicted grasping point information, the two-dimensional grasping point position label information, and the position reference point information according to the following steps:
[0060] The offset prediction result is generated by using the predicted grasping point information and the reference point information; the offset reference result is generated by using the two-dimensional grasping point position label information and the reference point information; and the loss function is generated according to the offset reference result and the offset prediction result.
[0061] In the optional embodiment, the offset prediction result can be obtained by subtracting the predicted grasping point information from the reference point information, and the offset reference result can be obtained by subtracting the two-dimensional grasping point position label information from the reference point information. The loss function is determined based on the difference between the offset reference result and the offset prediction result.
[0062] In the specific implementation, the offset of the 2D grasping point relative to the reference point can be obtained according to the feature output of each instance. The data with the instance corresponding 2D grasping point label prepared in advance is used as a supervision signal to supervise the output of the grasping point generation module. When the network is trained, the true value of the offset relative to the reference point can be obtained by subtracting the calibrated 2D grasping point position from the reference point position, and the parameters of the entire grasping point prediction model are trained by using the true value. Assuming that the offset prediction output by the network is d, and the offset calculated according to the calibration data is D, the loss function of the grasping point is the L1 distance between the prediction value and the true value, that is, the loss function can be L p= |D - d|1, where L p represents a loss function, is an L1 distance of D and d, D represents an offset reference result, and d represents an offset prediction result.
[0063] In an optional embodiment, the target neural network model further comprises a mask generation model and a bounding box generation model; after obtaining the instance feature information, the following steps can be further performed:
[0064] The instance feature information is input into the mask generation model to obtain predicted mask information; the instance feature information is input into the bounding box generation model to obtain predicted bounding box information; a second loss function is generated by using the predicted mask information, the predicted bounding box information, and the loss function; and the target neural network model is adjusted by using the second loss function to obtain a second grasp point prediction model.
[0065] In this optional embodiment, the target neural network model further comprises a mask generation model and a bounding box generation model; after obtaining the instance feature information, the grasp point generation model, the mask generation model, and the bounding box generation model can be calculated in parallel, as shown in a multi-task parallel training schematic diagram shown in Figure 4 The grasp point generation module can be used to perform data calculation of the grasp point generation model, the mask generation module can be used to perform data calculation of the mask generation model, and the bounding box generation module can be used to perform data calculation of the bounding box generation model. Then, a second loss function is generated by using the obtained predicted mask information, predicted bounding box information, predicted grasp point information, two-dimensional grasp point position label information, and position reference point information. The predicted grasp point information, two-dimensional grasp point position label information, and position reference point information can be used to generate a loss function according to the above steps, the predicted mask information and the predicted bounding box information are used to generate a loss function of the mask generation model and the bounding box generation model respectively, and the three obtained loss functions are combined to obtain the second loss function.
[0066] It should be noted that, in addition to the grasp point generation model, mask generation model, and bounding box generation model, the target neural network model can also include many other parallel models. Parallel prediction not only ensures the independence between individual predictions but also allows the use of multi-task supervision signals. For example, for the same instance feature, it can be supervised simultaneously by grasp points, segmentation masks, and bounding boxes. Therefore, this instance feature receives guidance from multiple sources, each with its own emphasis (e.g., bounding boxes emphasize the instance's position information, segmentation masks emphasize the instance's shape information, and grasp points need to consider both). This results in a better instance feature representation after training, effectively improving the overall model's prediction accuracy. During multi-task parallel training, the overall network's loss function is a weighted sum of the loss functions of all tasks. The specific weights depend on the specific form of the loss function used by the other tasks, i.e. Where L represents the second loss function, α i L represents the coefficient. i This represents the loss function of the model participating in parallel computation, and N represents the number of models participating in parallel computation.
[0067] In practical implementation, taking multi-task parallel training of three tasks—predicting grasp points, segmenting masks, and detecting bounding boxes—as an example, if the grasp point loss function uses the aforementioned L1 loss function, the segmentation mask loss function uses the commonly used Dice Loss, and the bounding box loss function uses the commonly used IoU Loss, then the weights of the three loss functions are 0.5, 3, and 1, respectively. That is, L = 0.5 * L p +3*L m +1*L b , where L p L m L b These represent the loss functions for the capture point, segmentation mask, and detection box, respectively.
[0068] In robotic grasping systems, due to practical constraints, the theoretically optimal grasping point may not be the actual optimal grasping point in certain special cases. Therefore, this grasping point prediction model can be extended to predict multiple grasping points, meaning that for each instance, more than one grasping point is predicted. In one optional implementation, the instance feature information is input into the grasping point generation model to obtain predicted grasping point information. A loss function is then generated using the predicted grasping point information, the two-dimensional grasping point location annotation information, and the location reference point information. This can be performed according to the following steps:
[0069] The instance feature information is input into the grabbing point generation model multiple times to obtain multiple sets of predicted grabbing point information; the reference point information and the multiple sets of predicted grabbing point information are used to generate multiple sets of offset prediction results; the two-dimensional grabbing point position labeling information and the reference point information are used to generate an offset reference result; and the offset reference result and the multiple sets of offset prediction results are used to generate a loss function.
[0070] In the optional embodiment, by increasing the prediction of multiple sets of offsets, the function of predicting multiple grabbing points for the same instance can be realized. After obtaining the multiple sets of predicted grabbing point information, the multiple sets of predicted grabbing point information are used to obtain multiple sets of offset prediction results, and a loss function is generated for each offset prediction result and offset reference result, so that the parameters of the target neural network model are adjusted based on multiple loss functions, so that multiple alternative grabbing points can be output according to image information.
[0071] Based on the grabbing point prediction model, the prediction of the two-dimensional grabbing point does not need the result of the task intermediate variable, the prediction of the grabbing point is more independent, and is more easily extended to the prediction of multiple grabbing points.
[0072] The embodiment of the application further provides a grabbing point prediction model training method, which can provide training efficiency and accuracy of the grabbing point prediction model.
[0073] According to another aspect of the embodiment of the application, a method for determining an object grabbing point is further provided, which comprises: acquiring a target picture; inputting the target picture into a grabbing point prediction model to obtain a two-dimensional grabbing point of an object in the target picture; wherein the grabbing point prediction model is obtained by training the grabbing point prediction model training method.
[0074] In the embodiment of the application, referring to Figure 2 The image feature extraction module can be used to perform data calculation of the image feature extraction model, the instance feature generation module can be used to perform data calculation of the instance feature generation model, and the grabbing point generation module can be used to perform data calculation of the grabbing point generation model. By using the grabbing point prediction model for prediction, a two-dimensional grabbing point of all objects in an RGB image can be directly output by inputting the RGB image. There can be multiple objects in a picture, and the network will also output multiple grabbing points, each object corresponding to its own grabbing point position.
[0075] The flow of 2D grasping point prediction can be implemented by a mature and optimized deep learning network structure, which improves the speed of 2D grasping point prediction and thus improves the efficiency of the grasping system. Meanwhile, in the scheme, the grasping point does not need to be generated first. The segmentation mask is calculated according to the segmentation mask. The prediction of the grasping point is only related to the input image, and the steps are more direct and simple, and do not depend on the results of intermediate quantities such as instance segmentation mask, which reduces the influence of incorrect intermediate quantity results on the accuracy of 2D grasping point prediction, and improves the correctness of 2D grasping point prediction.
[0076] In an optional embodiment, the method can further perform the following steps:
[0077] Obtaining the spatial position information of the object in the target picture; using the spatial position information to determine a target two-dimensional grasping point among the plurality of two-dimensional grasping points of the same object in the target picture.
[0078] In this optional embodiment, the spatial position information can be the height of the instance, the distance from the grasping placement destination, and other dimensional information obtained by combining other sensors such as binocular infrared cameras, etc. By combining these additional information and the positions of the plurality of grasping points, the best grasping point under this grasping system and scene, i.e. the target two-dimensional grasping point, is determined.
[0079] In an optional embodiment, the method can further perform the following steps:
[0080] Generating a three-dimensional grasping point of the object in the target picture according to the two-dimensional grasping point and the point cloud information of the target picture.
[0081] In this optional embodiment, after obtaining the 2D grasping point of the object, the 3D grasping point of the object can be directly obtained by combining the depth (or point cloud) information of the RGB image. The mechanical arm or grasping system can actually grasp the object according to the 3D grasping point.
[0082] According to another aspect of the embodiment of the present application, a grasping point prediction model training device is also provided, Figure 6 The schematic diagram of the grasping point prediction model training device provided by the embodiment of the present application is shown in Figure 6 The grasping point prediction model training device includes: a data module 61, an image module 62, an instance module 63, a reference point module 64, a grasping point module 65, and a training module 66. The data analysis device is described in detail below.
[0083] Data module 61 is used to acquire training images; the training images include one or more objects and two-dimensional grasping point location annotation information of the objects; image module 62 is used to input the training images into the image feature extraction model of the target neural network model to obtain image feature information; the target neural network model includes the image feature extraction model, the instance feature generation model, and the grasping point generation model; instance module 63 is used to input the image feature information into the instance feature generation model to obtain instance feature information; reference point module 64 is used to generate location reference point information of the instance feature information; grasping point module 65 is used to input the instance feature information into the grasping point generation model to obtain predicted grasping point information, and generate a loss function using the predicted grasping point information, the two-dimensional grasping point location annotation information, and the location reference point information; training module 66 is used to adjust the target neural network model using the loss function to obtain a grasping point prediction model.
[0084] It should be noted that the data module 61, image module 62, instance module 63, reference point module 64, capture point module 65, and training module 66 mentioned above correspond to steps S102 to S112 in the method embodiment. The examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in the above method embodiment.
[0085] In one optional implementation, acquiring training images includes: acquiring an RGB image and corresponding depth map data; generating first point cloud data of the RGB image using the RGB image and the depth map data; acquiring a calibration instruction; generating a segmentation mask of an object in the RGB image according to the calibration instruction; determining second point cloud data of the object using the segmentation mask and the first point cloud data; receiving a grasping point generation instruction; determining three-dimensional grasping points of the object according to the grasping point generation instruction and the second point cloud data; and marking the two-dimensional grasping point positions of the object in the RGB image based on the three-dimensional grasping points to obtain a training image.
[0086] In one optional implementation, generating the location reference point information of the instance feature information includes: receiving a reference point generation instruction, and determining the reference point information using the reference point generation instruction and the instance feature information; the reference point generation instruction is determined based on the feature position of the instance feature information.
[0087] In one optional implementation, generating the location reference point information of the instance feature information includes: generating reference point information using a reference point generation model; the reference point generation model is obtained by training a convolutional network with training data; the training data includes reference point location information.
[0088] In an optional implementation, the loss function is generated by using the predicted grasp point information, the two-dimensional grasp point position label information, and the position reference point information, including: generating an offset prediction result by using the predicted grasp point information and the reference point information; generating an offset reference result by using the two-dimensional grasp point position label information and the reference point information; and generating a loss function according to the offset reference result and the offset prediction result.
[0089] In an optional implementation, the target neural network model further includes a mask generation model and a bounding box generation model; after obtaining the instance feature information, the method further includes: inputting the instance feature information into the mask generation model to obtain predicted mask information; inputting the instance feature information into the bounding box generation model to obtain predicted bounding box information; generating a second loss function by using the predicted mask information, the predicted bounding box information, and the loss function; and adjusting the target neural network model by using the second loss function to obtain a second grasp point prediction model.
[0090] In an optional implementation, the instance feature information is input into the grasp point generation model to obtain predicted grasp point information, and the loss function is generated by using the predicted grasp point information, the two-dimensional grasp point position label information, and the position reference point information, including: inputting the instance feature information into the grasp point generation model multiple times to obtain multiple groups of predicted grasp point information; generating multiple groups of offset prediction results by using the reference point information and the multiple groups of predicted grasp point information; generating an offset reference result by using the two-dimensional grasp point position label information and the reference point information; and generating a loss function according to the offset reference result and the multiple groups of offset prediction results.
[0091] According to another aspect of the embodiment of the present application, an object grasp point determination device is also provided, including: an acquisition module configured to acquire a target picture; and a determination module configured to input the target picture into a grasp point prediction model to obtain a two-dimensional grasp point of an object in the target picture, wherein the grasp point prediction model is obtained by training the grasp point prediction model training method.
[0092] In an optional implementation, the device further includes: a calculation module configured to acquire spatial position information of the object in the target picture; and determine a target two-dimensional grasp point from multiple two-dimensional grasp points of the same object in the target picture by using the spatial position information.
[0093] In an optional implementation, the device further includes: a three-dimensional module configured to generate a three-dimensional grasp point of the object in the target picture according to the two-dimensional grasp point and point cloud information of the target picture.
[0094] The exemplary embodiments of the present application further provide an electronic device, comprising: at least one processor; and a memory connected with the at least one processor in communication. The memory stores a computer program capable of being executed by the at least one processor, and the computer program, when executed by the at least one processor, is configured to cause the electronic device to perform the method according to the embodiments of the present application.
[0095] The exemplary embodiments of the present application further provide a non-transitory computer readable storage medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is configured to cause the computer to perform the method according to the embodiments of the present application.
[0096] Reference Figure 5 The structure block diagram of the electronic device 500 which can be a server or a client of the present application will now be described, which is an example of a hardware device that can be applied to various aspects of the present application. The electronic device is intended to represent a wide variety of digital electronic computer devices such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computer devices. The electronic device can also represent a wide variety of mobile devices such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present application described and / or claimed in this document.
[0097] As Figure 5 shown, the electronic device 500 includes a computing unit 501 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded into a random access memory (RAM) 503 from a storage unit 508. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0098] A plurality of components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, an output unit 507, a storage unit 508, and a communication unit 509. The input unit 506 can be any type of device that can input information to the electronic device 500, and can receive inputted digital or character information, and generate key signal inputs related to user settings and / or function controls of the electronic device. The output unit 507 can be any type of device that can present information, and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 508 can include, but is not limited to, a magnetic disk, an optical disk. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, e.g., a Bluetooth device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0099] The computing unit 501 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 performs various methods and processes described above. For example, in some embodiments, the above-described grasp point prediction model training method or object grasp point determination method can be implemented as a computer software program, which is tangibly embodied in a machine-readable medium, e.g., the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 500 via the ROM 502 and / or the communication unit 509. In some embodiments, the computing unit 501 can be configured to perform the above-described grasp point prediction model training method or object grasp point determination method by any other appropriate means, e.g., by means of firmware.
[0100] Program code for carrying out the methods of the present application can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be embodied in whole or in part within a machine, executed partially on the machine, partially on the machine and partially on a remote machine or entirely on a remote machine or server.
[0101] In the context of this application, a machine-readable medium can be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine- readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium will include one or more lines of electrical conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0102] As used in this application, the terms "machine-readable medium" and "computer- readable medium" refer to any computer program product, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal that can be used to provide machine instructions and / or data to a programmable processor.
[0103] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0104] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0105] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
Claims
1. A method for training a grasping point prediction model, comprising: Acquire training images; the training images include one or more objects and the two-dimensional grasping point position annotation information of the objects; The training images are input into the image feature extraction model of the target neural network model to obtain image feature information; the target neural network model includes the image feature extraction model, the instance feature generation model, and the grasp point generation model. The image feature information is input into the instance feature generation model to obtain instance feature information; The location reference point information for generating the instance feature information is determined based on the feature location or instance feature information. The instance feature information is input into the grasping point generation model to obtain predicted grasping point information. A loss function is generated using the predicted grasping point information, the two-dimensional grasping point location annotation information, and the location reference point information. Generating the loss function using the predicted grasping point information, the two-dimensional grasping point location annotation information, and the location reference point information includes: generating an offset prediction result using the predicted grasping point information and the reference point information; generating an offset reference result using the two-dimensional grasping point location annotation information and the reference point information; and generating a loss function based on the offset reference result and the offset prediction result. The target neural network model is adjusted using the loss function to obtain the grab point prediction model.
2. The method as described in claim 1, wherein, Obtain training images, including: Obtain an RGB image and the corresponding depth map data of the RGB image, and use the RGB image and the depth map data to generate the first point cloud data of the RGB image; Obtain calibration instructions, generate a segmentation mask of the object in the RGB image according to the calibration instructions, and determine the second point cloud data of the object using the segmentation mask and the first point cloud data; The system receives a gripping point generation instruction, determines the three-dimensional gripping points of the object based on the gripping point generation instruction and the second point cloud data, and marks the two-dimensional gripping point positions of the object in the RGB image based on the three-dimensional gripping points to obtain a training image.
3. The method as described in claim 1, wherein, The location reference point information for generating the instance feature information includes: Receive a reference point generation instruction, and use the reference point generation instruction and the instance feature information to determine reference point information; the reference point generation instruction is determined based on the feature position of the instance feature information.
4. The method of claim 1, wherein, The location reference point information for generating the instance feature information includes: Reference point information is generated using a reference point generation model; the reference point generation model is obtained by training a convolutional network with training data; the training data includes reference point location information.
5. The method as described in claim 1, wherein the target neural network model further comprises a mask generation model and a detection box generation model; After obtaining the instance feature information, it also includes: The instance feature information is input into the mask generation model to obtain the predicted mask information; The instance feature information is input into the detection box generation model to obtain the predicted detection box information; A second loss function is generated using the predicted mask information, the predicted detection box information, and the loss function. The target neural network model is adjusted using the second loss function to obtain the second grasping point prediction model.
6. The method as described in claim 1, wherein the instance feature information is input into the grasping point generation model to obtain predicted grasping point information, and a loss function is generated using the predicted grasping point information, the two-dimensional grasping point location annotation information, and the location reference point information, comprising: The instance feature information is input into the grab point generation model multiple times to obtain multiple sets of predicted grab point information; Multiple sets of offset prediction results are generated using the reference point information and the multiple sets of predicted capture point information; An offset reference result is generated using the two-dimensional grasping point location annotation information and the reference point information; A loss function is generated based on the offset reference result and the multiple sets of offset prediction results.
7. A method for determining the object grasping point, comprising: Obtain the target image; The target image is input into the grasping point prediction model to obtain the two-dimensional grasping points of the objects in the target image; The grab point prediction model is obtained by training the method described in any one of claims 1-6.
8. The method of claim 7, further comprising: Obtain the spatial location information of objects in the target image; Using the spatial location information, the target two-dimensional grasping point is determined from multiple two-dimensional grasping points of the same object in the target image.
9. The method of claim 7, further comprising: Based on the two-dimensional grasping points and the point cloud information of the target image, generate three-dimensional grasping points for the objects in the target image.
10. A training device for a grasping point prediction model, comprising: The data module is used to acquire training images; The training image includes one or more objects and the two-dimensional gripping point location annotation information of the objects; The image module is used to input the training images into the image feature extraction model of the target neural network model to obtain image feature information; the target neural network model includes the image feature extraction model, the instance feature generation model, and the grasp point generation model; The instance module is used to input the image feature information into the instance feature generation model to obtain instance feature information; The reference point module is used to generate location reference point information for the instance feature information, wherein the location reference point information is determined based on the feature position or instance feature information. The grasping point module is used to input the instance feature information into the grasping point generation model to obtain predicted grasping point information, and to generate a loss function using the predicted grasping point information, the two-dimensional grasping point position annotation information, and the position reference point information; wherein, generating the loss function using the predicted grasping point information, the two-dimensional grasping point position annotation information, and the position reference point information includes: generating an offset prediction result using the predicted grasping point information and the reference point information; generating an offset reference result using the two-dimensional grasping point position annotation information and the reference point information; and generating a loss function based on the offset reference result and the offset prediction result; The training module is used to adjust the target neural network model using the loss function to obtain the grab point prediction model.
11. A device for determining an object grasping point, comprising: The acquisition module is used to acquire the target image; The determination module is used to input the target image into the grasping point prediction model to obtain the two-dimensional grasping points of the objects in the target image; The grab point prediction model is obtained by training the method described in any one of claims 1-6.
12. An electronic device, comprising: processor; as well as Stored program memory, The program includes instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1-9.
13. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-9.
Citation Information
Patent Citations
Method for detecting the grabbing position of a robot target object
CN109658413A
Object grabbing point visual positioning method and device, storage medium and electronic equipment
CN112258567A
Positioning method and device, computer equipment and computer readable storage medium
CN113591841A