Image Data Processing Method, Apparatus, Electronic Device, and Storage Medium

Through color pictures and deep learning network identification masks and morphological processing, combined with template matching algorithm correction masks, the problem of insufficient point cloud data in industrial scenarios such as black transparent glass bottles is solved, accurate acquisition of point information and automatic conversion of grabbing point information is achieved, and the reliability of robot grabbing is improved.

CN114092428BActive Publication Date: 2025-07-04MECH MIND ROBOTICS TECH LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111338191.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-10
Publication Date
2025-07-04
Estimated Expiration
2041-11-10

AI Technical Summary

Technical Problem

In industrial scenarios, especially for items such as black transparent glass bottles, it is difficult to collect clear point cloud data, which makes the robot unable to accurately identify the crawling point, resulting in the problem of crawling failure or drop.

Method used

By using other image data such as color pictures, combining deep learning networks to identify the mask and perform morphological processing, the mask is corrected using the template matching algorithm to obtain accurate grab point information, and automatically convert the two-dimensional grab point into three-dimensional grab point.

Benefits of technology

In the absence of point cloud data, the crawling point information can be accurately obtained, avoiding the problem of inaccurate crawling or dropping, reducing manual intervention, and improving the reliability of robot crawling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114092428B_ABST
    Figure CN114092428B_ABST
Patent Text Reader

Abstract

The present application discloses an image data processing method, apparatus, electronic device, and storage medium. The image data processing method includes: receiving image data containing an item to be processed; identifying the item to be processed from the image data and generating a mask of the item to be processed; performing morphological processing on the generated mask of the item to be processed; based on a pre-stored template of the item to be processed, using a template matching algorithm to process the mask after morphological processing; and obtaining a corrected mask of the item to be processed and / or grasping point association information based on the processing result of the template matching algorithm. The present invention enables, in the case where an accurate item mask cannot be obtained, correcting an inaccurate mask to obtain as accurate a mask as possible and further obtaining accurate grasping point information, effectively avoiding the problem of inaccurate grasping points caused by an inaccurate mask, which in turn leads to inaccurate grasping or dropping during grasping.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of automatic control and program control B25J of robotic arms or jigs. More specifically, it particularly relates to an image data processing method, apparatus, electronic device, and storage medium. Background Art

[0002] In recent years, industries such as e-commerce and express delivery have rapidly emerged, creating good development opportunities for the logistics industry. With the maturity of technologies such as artificial intelligence and machine vision, in the most basic parts of logistics such as warehousing, handling, and sorting, more and more automated devices are being used, and it is an inevitable trend for industrial robots to replace humans. As an important part of industrial robots, jigs are increasingly being used in the logistics industry, sorting, the 3C industry, and food feeding.

[0003] Currently, the automatic control of robotic arms or jigs requires point cloud data of the item to be grasped. The usual approach is to first use a 3D camera to collect point cloud data, then determine trajectory points of the robotic arm's movement, or information such as grasping points based on the point cloud data, and then control the robotic arm to perform a grasping operation at an appropriate position along an appropriate trajectory based on the trajectory points or grasping points, etc. The robotic arm operates in three-dimensional space, and its movement trajectory is a three-dimensional trajectory in the three directions of length, width, and height. Therefore, the trajectory points or grasping points should also have information in three dimensions of the X-axis, Y-axis, and Z-axis. The point cloud data collected by the 3D camera includes information in these three dimensions. However, in some industrial scenarios, it is difficult to obtain clear point cloud information of the grasping object. For example, in the cosmetics industry, glass, especially black glass bottles, are used to hold liquids. Such glass bottles have weak self-light signals, and the reflective material is easily interfered by multiple reflections from surrounding objects. Moreover, since the glass is transparent, there will be a lot of diffuse reflection and multiple reflection data, making it difficult to collect appropriate point cloud data. Specifically, in such industrial scenarios, it may be impossible to obtain the point cloud of the item to be grasped, or the point cloud of the item to be grasped collected may be missing, or there may be other poor point cloud conditions. This results in the robot either being unable to recognize the item to be grasped, or calculating incorrect grasping points based on incorrect point cloud data and performing grasping using the incorrect grasping points, leading to failure to grasp, or even the bottle dropping. Currently, there is no solution in the prior art for automatically controlling a robot to perform grasping based on robot vision in such industrial scenarios with missing point clouds. Summary of the Invention

[0004] In view of the above problems, the present invention is proposed to overcome or at least partially solve the above problems. Specifically, first, the present invention can obtain the grasping point information of the item to be grasped by means of other image data such as color pictures in the case where the point cloud of the item to be grasped cannot be obtained, enabling the robot or fixture to directly rely on the grasping point information to achieve the grasping of the item to be grasped without relying on the point cloud of the item, effectively solving the problem of item grasping in the environment where the point cloud is missing; second, the present invention proposes a method for correcting the mask and obtaining two-dimensional grasping point information, so that in the case where an accurate item mask cannot be obtained, the inaccurate mask can be corrected to obtain an accurate mask as much as possible and further obtain accurate grasping point information, effectively avoiding the problem that inaccurate grasping points caused by inaccurate masks lead to inaccurate grasping or dropping during grasping; third, the present invention proposes a method for automatically converting the input two-dimensional grasping point information into three-dimensional grasping point information by the robot. This method can automatically obtain the reference information that can convert the two-dimensional grasping point into a three-dimensional grasping point according to the environmental characteristics of the item to be grasped, and obtain the grasping point information based on the reference information for the robot to grasp. This solution enables the complete grasping point information to be supplemented based on the existing information in the case where the grasping point information is missing, and reduces manual intervention.

[0005] All the solutions disclosed in the claims and the specification of this application have one or more of the above innovations, and correspondingly, can solve one or more of the above technical problems. Specifically, this application provides an image data processing method, device, electronic device, and storage medium.

[0006] The image data processing method of the embodiment of this application includes:

[0007] Receiving image data including the item to be processed;

[0008] Identifying the item to be processed from the image data and generating a mask of the item to be processed;

[0009] Performing morphological processing on the generated mask of the item to be processed;

[0010] Based on the template of the item to be processed saved in advance, using the template matching algorithm to process the mask after morphological processing;

[0011] Obtaining the corrected mask and / or grasping point association information of the item to be processed based on the processing result of the template matching algorithm.

[0012] In some embodiments, the morphological processing includes morphological dilation processing.

[0013] In some embodiments, the matching algorithm includes a shape-based matching algorithm.

[0014] In some embodiments, the grasping point includes the center point of the graspable area of the article.

[0015] In some embodiments, identifying the article to be processed from the image data and generating a mask of the article to be processed includes: processing the image data based on deep learning to identify the article to be processed and generating a mask of the article to be processed.

[0016] In some embodiments, the grasping point association information includes two-dimensional information of the grasping point.

[0017] The image data processing device according to an embodiment of the present application includes:

[0018] An image data receiving module, configured to receive image data including an article to be processed;

[0019] A mask generation module, configured to identify the article to be processed from the image data and generate a mask of the article to be processed;

[0020] A mask processing module, configured to perform morphological processing on the generated mask of the article to be processed;

[0021] A processing module, configured to process the morphologically processed mask based on a template of the article to be processed pre-stored, using a template matching algorithm, and obtain a corrected mask and / or grasping point association information of the article to be processed based on the processing result of the template matching algorithm.

[0022] In some embodiments, the morphological processing includes morphological dilation processing.

[0023] In some embodiments, the matching algorithm includes a shape-based matching algorithm.

[0024] In some embodiments, the grasping point includes the center point of the graspable area of the article.

[0025] In some embodiments, identifying the article to be processed from the image data and generating a mask of the article to be processed includes: processing the image data based on deep learning to identify the article to be processed and generating a mask of the article to be processed.

[0026] In some embodiments, the grasping point association information includes two-dimensional information of the grasping point.

[0027] The electronic device according to an embodiment of the present application includes a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the image data processing method according to any one of the above embodiments is implemented.

[0028] A computer-readable storage medium according to an embodiment of the present application stores a computer program, and when the computer program is executed by a processor, it implements the image data processing method according to any one of the above embodiments.

[0029] Additional aspects and advantages of the present application will be given in part in the following description, will become apparent in part from the following description, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0031] Figure 1 is a flowchart of a method for obtaining grasping point information in a poor point cloud scenario according to some embodiments of the present application;

[0032] Figure 2 is a flowchart of a method for obtaining a correction mask and grasping point association information according to some embodiments of the present application;

[0033] Figure 3 is a flowchart of a method for converting two-dimensional grasping point information into three-dimensional grasping point information according to some embodiments of the present application;

[0034] Figure 4 is a schematic diagram of an item to be grasped and an item mask obtained using a non-specialized deep learning network according to some embodiments of the present application;

[0035] Figure 5 is a schematic diagram of the structure of a device for obtaining grasping point information in a poor point cloud scenario according to some embodiments of the present application;

[0036] Figure 6 is a schematic diagram of the structure of a device for obtaining a correction mask and grasping point association information according to some embodiments of the present application;

[0037] Figure 7 is a schematic diagram of the structure of a device for converting two-dimensional grasping point information into three-dimensional grasping point information according to some embodiments of the present application;

[0038] Figure 8 is a schematic diagram of the structure of an electronic device according to some embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.

[0040] Figure 1 The flowchart of the method for obtaining the object grasping point information according to an embodiment of the present invention is shown. As Figure 1 shown, the method includes:

[0041] Step S100, obtaining a non-point cloud image including the object to be grasped;

[0042] Step S110, processing the non-point cloud image to obtain a mask of the object to be grasped;

[0043] Step S120, processing the mask of the object to be grasped to obtain grasping point association information, where the grasping point association information includes information for calculating the grasping point information;

[0044] Step S130, based on the grasping point association information, obtaining the grasping point information for controlling the robot to grasp the object to be grasped.

[0045] For step S100, in an industrial scenario, it is usually necessary to use a robot to grasp a large number of objects. Therefore, there may be multiple objects to be grasped, rather than just one. As a preferred embodiment, the object to be grasped can be a group of black cosmetic bottles placed in a bin. The robot needs to grasp the cosmetic bottles from the bin and transport them to other positions. Since the black glass itself has a weak optical signal and is easily interfered by multiple reflections of surrounding objects, there is a lot of diffuse reflection and multiple reflection data. When using an industrial camera to take pictures, it is very likely that no point cloud can be obtained, and it is difficult to collect the 3D point cloud data usually required for robot control. Therefore, the method of the present invention is particularly suitable for use in such a scenario to identify the object to be grasped by using image data other than point cloud data, calculate the grasping point information, and control the robot to perform grasping. The non-point cloud image data in this embodiment can preferably be a 2D color picture of the object. Different from the point cloud, the 2D color picture can more clearly identify the object contained therein. Even if the object is an object such as a black transparent glass bottle that cannot generate a point cloud, it can be clearly photographed and identified. It can be photographed by an industrial camera, and the group of objects to be grasped is placed below the visual sensor for photographing to obtain the image data of the object to be grasped.

[0046] For step S110, any existing method can be used to calculate the mask of the object. Preferably, the mask of the object can be obtained based on a deep learning network. Image recognition is a conventional application of deep learning networks, and there are already general deep learning networks in the prior art that can perform image recognition. As a specific implementation, an existing deep learning network for identifying objects or extracting object masks can be used to identify the object and extract the mask. To improve the recognition accuracy, the network can be pre-trained. The obtained 2D color image is input into the deep learning network, and the color image is processed by the deep learning network to identify the region of interest of the image for which the mask is to be generated. An image mask is generated in this region, and the original image region is masked using the image mask. In this way, the mask of the object to be grasped can be obtained based on the color image.

[0047] For step S120, after obtaining the mask of the object, the grasping point association information of the object can be calculated in the mask region. The so-called grasping point association information lacks some necessary information required for grasping point information and is not complete grasping point information, so it cannot be directly used by the robot to perform grasping. However, the grasping point association information can be used to calculate the grasping point information. As a specific implementation, the grasping point association information can be the two-dimensional information of the grasping point, such as the coordinate information of the X-axis and Y-axis. Assuming that the current control system needs to determine the information of the three dimensions of the X-axis, Y-axis, and Z-axis of the grasping point before it can control the grasping position of the robot (depending on the different fixtures used, more information may be required in the actual industrial scenario, such as the rotation angle of the fixture, etc.), then after obtaining the two-dimensional grasping point information, the grasping cannot be performed based on this information, and the two-dimensional grasping point information needs to be converted into three-dimensional grasping point information in the subsequent steps. Among them, the specific position of the grasping point is related to the fixture used and the object to be grasped. For example, when the object to be grasped is a black cosmetic bottle, the center point of the bottle mouth can be selected as the grasping point.

[0048] A non-customized deep learning network that can be used and identify objects in various situations, when processing the photographed photo, identifying the object in the photo and extracting the object mask, the generated mask usually does not fit exactly with the photographed object. Sometimes it may slightly exceed the graspable area of the object, sometimes it may be slightly smaller than the graspable area of the object, and the shape of the mask is usually inconsistent with the graspable area of the object. Figure 4 Shows the object to be grasped and a typical mask of the object to be grasped generated by the deep learning network. Specifically, Figure 4Shown is a scenario where the item to be grasped is the above-mentioned cosmetic bottle. The circular part therein is the bottle mouth area of the glass bottle to be grasped, that is, the grasping area of the robotic arm. The shaded part is the mask extracted by a general deep learning network. It can be seen that in this case, the mask of the item is inaccurate, and the grasping points calculated therefrom are naturally inaccurate. Inaccurate grasping points may lead to problems such as the inability to grasp the bottle or the bottle falling during grasping. One solution is to design a deep learning network dedicated to this scenario and repeatedly train it to improve the accuracy of the processing result after inputting the image into the network. In this way, accurate masks and grasping points can be obtained after inputting the image into this deep learning network. However, this solution has a high cost and is dedicated to a specific scenario, with low flexibility. Another solution is to process the inaccurate mask and obtain the grasping point information from the processed mask.

[0049] The present invention preferably uses the latter solution. Specifically, the applicant has developed a method for mask correction and grasping point acquisition with relatively low cost and high generality, which is also one of the focuses of the present invention.

[0050] Figure 2 The flowchart of an image data processing method for mask correction and grasping point association information acquisition according to an embodiment of the present invention is shown. As Figure 2 shown, the method includes:

[0051] Step S200, receiving image data containing the item to be processed;

[0052] Step S210, identifying the item to be processed from the image data and generating a mask of the item to be processed;

[0053] Step S220, performing morphological processing on the generated mask of the item to be processed;

[0054] Step S230, further processing the mask after morphological processing to obtain a corrected mask of the item to be processed and / or grasping point association information.

[0055] For step S200, the type of the image data is not limited in this embodiment. Any type of image data, as long as it can identify the item contained therein and generate a mask through existing deep learning algorithms, can be applicable to this embodiment. Preferably, this embodiment can obtain the image data by a method similar to step S100, and the image data can be a 2D color image.

[0056] For step S210, the mask of the item to be processed can be identified and generated in a manner similar to step S120.

[0057] For step S220, as Figure 4As shown, for the object mask obtained by a general deep learning network, the mask area usually cannot perfectly fit the object contour ( Figure 4 the grasping area in

[0058] is the area where the bottle mouth is located. It is not difficult to see that there is a large gap between the mask and the actual bottle mouth), showing a skewed state, and there may also be many holes inside. For general object recognition applications, this has little impact. However, in the industrial scenario of this invention for grasping objects, the accuracy requirement is relatively high, and such errors are intolerable. Therefore, it is necessary to process the obtained mask. The first thing to do is morphological processing to change the graphic form of the mask area. In one implementation, the morphological processing can be dilation processing. After obtaining the two-dimensional image information, perform dilation processing on the image to fill in the missing, irregular and other defects of the image. For example, for each pixel point on the mask, a certain number of points around this point, such as 8 - 25 points, can be set to have the same color as this point. This step is equivalent to filling the surroundings of each pixel point. Therefore, if there are missing parts in the object mask, this operation will fill all the missing parts. After such processing, the object mask will become complete without missing parts, and at the same time, the overall mask will also become slightly "fatter" due to dilation. Appropriate dilation helps subsequent further image processing operations.

[0059] For step S230, as an example, the present invention discloses three processing methods to obtain the corrected mask of the object and obtain the grasping point association information.

[0060] Step S240, obtain the circumscribed rectangle of the mask after dilation processing;

[0061] Step S241, generate an inscribed circle of the circumscribed rectangle based on the circumscribed rectangle of the mask;

[0062] Step S242, obtain the corrected mask of the object to be processed and / or the grasping point association information based on the inscribed circle.

[0063] For step S240, any circumscribed rectangle algorithm can be used to obtain the circumscribed rectangle of the mask. As a specific implementation, the X coordinate value and Y coordinate value of each pixel point in the mask can be calculated, and the minimum X value, the minimum Y value, the maximum X value and the maximum Y value are respectively selected; then, the 4 values are combined into the coordinates of a point, that is, the minimum X value and the minimum Y value form the coordinates (X min , Y min ), the maximum X value and Y value form the coordinates (X max , Y max ), the minimum X value and the maximum Y value form the coordinates (X min , Y max), and the maximum X value and the minimum Y value to form the coordinates (X max , Y min ). Taking the points (X min , Y min ), (X max , Y max ), (X min , Y max ), (X max , Y min ) as the four corner points of the circumscribed rectangle and connecting them, the circumscribed rectangle is obtained.

[0064] For step S241, the key of the present invention is to use the inscribed circle algorithm as a part of calculating the correction mask and the grasping point information, without making any improvement to the inscribed circle algorithm. Therefore, the specific inscribed circle algorithm is not limited, and any inscribed circle algorithm can be used in the present invention.

[0065] For step S242, after obtaining the inscribed circle, the mask part enclosed by the inscribed circle can be calculated, and this part of the mask is used as the correction mask of the item to be processed. The contour of the obtained correction mask is the same as the shape and size of the inscribed circle. Calculate the position of the center of the circle on the inscribed circle correction mask and obtain the information of the center of the circle, and this information of the center of the circle is used as the grasping point association information. Specifically, the two-dimensional position information of the center of the circle, for example, the X-axis and Y-axis information of the center of the circle, can be used as the X-axis position information and Y-axis position information of the grasping point.

[0066] The correction mask obtained in this way has the same shape as the bottle mouth shape, but the covered area may still be different from the actual area of the bottle mouth.

[0067] The second processing method includes:

[0068] Step S250, using a circle detection algorithm to process the mask after dilation processing;

[0069] Step S251, obtaining the correction mask of the item to be processed and / or the grasping point association information based on the processing result of the circle detection algorithm.

[0070] For step S250, the circle detection algorithm is also called the circle finding algorithm, which can be used to detect circular features in an irregular graph and find the circles contained in the graph. Commonly used algorithms include the circular Hough transform algorithm, the random Hough transform algorithm, the random circle detection algorithm, etc. The focus of this embodiment is to use the circle detection algorithm to find circles from the mask after morphological processing, without limiting the specific circle detection algorithm used. Since the bottle mouth itself is circular and the collected mask contains some features of the bottle mouth shape, the circle found in the mask area is roughly the position of the bottle mouth.

[0071] For step S251, after finding the circle in the mask through the circle detection algorithm, the part of the mask enclosed by the circle can be used as the corrected mask of the item to be processed. Then, calculate its center, and use the information of the center as the grasping point association information. Similar to Method 1, the grasping point association information can be the two-dimensional position information of the center, and this information is used as the X-axis position information and Y-axis position information of the grasping point.

[0072] The second method does not require calculating the circumscribed rectangle and inscribed circle that do not originally exist, and only needs to find the circular part from the existing mask area. In this way, the calculation accuracy is higher than that of the first method.

[0073] The third processing method includes:

[0074] Step S260, based on the pre-saved template of the item to be processed, use the template matching algorithm to process the dilated mask;

[0075] Step S261, obtain the corrected mask and / or grasping point association information of the item to be processed based on the processing result of the template matching algorithm.

[0076] For step S260, the template of the item to be processed can be the template of the whole item to be processed or the template of the grasping area of the item to be processed. For example, in the scenario where the item to be grasped is a black glass cosmetic bottle and the fixture needs to grasp the bottle mouth of the cosmetic bottle, the template of the item to be processed can be a three-dimensional template established for the whole black glass cosmetic bottle, or only a template established for the graspable area of the bottle mouth. After obtaining the uncorrected mask area, based on the pre-stored template, use the matching algorithm to match the template within the mask area. Simply put, the template is equivalent to a known small image, and the template matching algorithm is equivalent to searching for a target in a large image including the small image. It is known that there is a target to be found in the picture, and the target has the same size, direction, and image elements as the template. Through this template matching algorithm, the target, that is, the small image, can be found in the picture and its pose can be determined. In this embodiment, no specific limitation is imposed on the matching algorithm. Since the mask itself will lose color information, the key is the shape matching rather than the color matching. Therefore, the present invention preferably uses a shape-based matching algorithm for matching. In addition, considering the matching efficiency and accuracy comprehensively, when performing template matching, when the shape similarity reaches 70-95%, it can be considered that the matching is successful. Which specific value to select can be selected and adjusted according to the needs of the actual application scenario.

[0077] For step S261, after finding the shape in the mask that matches the pre-stored template, the mask surrounded by this shape can be used as the correction mask, and the grasping point correlation information can be further calculated. Since the template is adopted, no matter what shape the item to be grasped presents within the area to be grasped, the item can be matched and grasped, not limited to the scenario where the grasping area of the item to be grasped is circular. Correspondingly, when grasping different items, the positions of the grasping points are also different. In one embodiment, the grasping point can be the center point of the correction mask of the item to be grasped, and the grasping point correlation information can be the two-dimensional position information of this center point, that is, the X-axis position information and Y-axis position information of the grasping point.

[0078] In engineering practice, the inventor found that the third method can meet relatively high standards both in terms of accuracy and operation speed. It is the optimal among the three implementation methods of the present invention. Moreover, the first two methods are actually only applicable to industrial scenarios where the area to be grasped is circular, while the third method can be used for grasping any item. Therefore, the third implementation method is also one of the focuses of the present invention.

[0079] In step S130, as described above, the grasping point correlation information lacks some necessary information required by the grasping point information and is not complete grasping point information. Therefore, it cannot be directly used by the robot to perform grasping, but the grasping point correlation information can be used to calculate the grasping point information. In one implementation, the grasping point correlation information can be two-dimensional information of the grasping point, while the grasping point information required by the robot is three-dimensional information. In order to be used by the robot, the two-dimensional grasping point correlation information should be changed into three-dimensional information, so as to grasp the item to be grasped based on the three-dimensional information. In the case where the three-dimensional data of the grasping point cannot be directly obtained, in order to convert the two-dimensional data into three-dimensional data, one method is to manually input the third-dimensional data of the item to be grasped. For example, when the item to be grasped is a black glass cosmetic bottle for which three-dimensional grasping point information cannot be obtained, the specific height information can be pre-input by means of manual entry. In this way, after processing the two-dimensional image through the foregoing solution and obtaining the two-dimensional information of the grasping point, the two-dimensional grasping point information can be further converted into three-dimensional grasping point information for the robot to use based on the height information. This method requires manual input of the height information before grasping. However, in many robot-based application scenarios, the goal is to reduce manual participation. And with the method of manually inputting the height information, the material box for loading the bottles must always be in the same position. If the material box is placed in different positions, the height information will change and the height information needs to be re-entered, otherwise correct grasping cannot be achieved. Therefore, a more automated method is desired.

[0080] To achieve this automated processing method, the applicant proposes a method for converting two-dimensional grasping point information into three-dimensional grasping point information without manual intervention, which is also one of the key points of the present invention.

[0081] Figure 3 FIG. shows a schematic flowchart of a method for converting two-dimensional grasping point information into three-dimensional grasping point information according to an embodiment of the present invention. As Figure 3 shown, the method includes:

[0082] Step S300, obtaining reference object information of the item to be grasped;

[0083] Step S310, obtaining two-dimensional grasping point information of the item to be grasped;

[0084] Step S320, processing the reference object information of the item to be grasped to obtain reference information, where the reference information includes information that the two-dimensional grasping point information does not have and can convert two-dimensional information into three-dimensional information;

[0085] Step S330, generating three-dimensional grasping point information of the item to be grasped based on the reference information and the two-dimensional grasping point information of the item to be grasped.

[0086] For step S300, the present invention is applied to a scenario where the point cloud of the item to be grasped is in poor condition. Therefore, the reference object should be an item with a qualified point cloud. In the present invention, a qualified point cloud means that information about the dimension missing from the two-dimensional grasping point information of the item can be obtained through the point cloud. For example, in a case where the X-axis information and Y-axis information of the grasping point can be obtained, an item with recognizable Z-axis information can be used as a reference object. The reference object can be an item relatively close to the item to be grasped, or other similar items to be grasped placed together with the item to be grasped. Specifically, for an industrial scenario where a large number of items to be grasped, such as cosmetic bottles, are placed in a material box, since the point cloud of the box is usually complete and as long as the whole box does not have strong deformation, its height at each position is the same, that is, at each position, the height of the material box is the same as the height of the item to be grasped or has a fixed height difference from the item to be grasped. Therefore, it is suitable as a reference object. Therefore, when the point cloud of the item to be grasped cannot be obtained in this scenario, the 2D color image data of the item and the point cloud data of the recognized box can be collected for subsequent steps. In other embodiments, when using a camera for shooting, since multiple items to be grasped are in different positions, in the overall point cloud data obtained by shooting at a certain position, the point cloud quality of each item to be grasped is different. Specifically, it may not be possible to obtain suitable point cloud data for some items to be grasped, but suitable point cloud data for other items to be grasped can be obtained. In this case, the item to be grasped with better point cloud among them can be selected as the reference object for other items to be grasped.

[0087] The point cloud information can be obtained through a 3D industrial camera. The 3D industrial camera is generally equipped with two lenses, which capture the group of items to be grasped from different angles respectively. After processing, the three-dimensional image of the object can be displayed. Place the group of items to be grasped under the visual sensor, and the two lenses take pictures simultaneously. According to the relative pose parameters of the two obtained images, the X, Y, and Z coordinate values and the coordinate orientations of each point of the object to be filled are calculated using a general binocular stereo vision algorithm, and then converted into the point cloud data of the group of items to be grasped. In specific implementation, components such as a laser detector, a visible light detector such as an LED, an infrared detector, and a radar detector can also be used to generate the point cloud, and the present invention does not limit the specific implementation manner.

[0088] As an example, a two-dimensional color map corresponding to the three-dimensional item area and a depth map corresponding to the two-dimensional color map can also be obtained along the depth direction perpendicular to the item. Among them, the two-dimensional color map corresponds to the image of the plane area perpendicular to the preset depth direction; each pixel point in the depth map corresponding to the two-dimensional color map corresponds one by one to each pixel point in the two-dimensional color map, and the value of each pixel point is the depth value of the pixel point. In one implementation manner, the obtained reference object information can be the point cloud of the reference object or the depth map of the reference object.

[0089] For step S310, in this embodiment, two-dimensional grasping point information needs to be obtained, but the key does not lie in the obtaining method, so the specific method for obtaining this information is not limited either. Preferably, the method for obtaining the grasping point association information in any of the foregoing embodiments can be used to obtain the two-dimensional grasping point information.

[0090] For step S320, taking an industrial scenario where multiple black glass cosmetic bottles to be grasped are arranged and placed in a material box as an example, after obtaining the point cloud of the whole group of items, the point cloud of a relatively clear reference object can be further identified from the obtained overall point cloud. For example, the point cloud of the material box or the point cloud of the bottle mouth with relatively clear points can be identified from the overall point cloud. Then, the identified point cloud is processed to extract the height information therein as the reference information. Although in this embodiment, the missing information of the two-dimensional grasping point information is taken as the height information as an example, those skilled in the art can understand that when the missing information is not the height information, the reference information may not be the height information either.

[0091] For step S330, if the point cloud of the bottle mouth is used, since the types of bottles in the material box are the same, the height of this point cloud is the same as the heights of the bottle mouths of all bottles. After obtaining the height information of the bottle mouth, combine it with the two-dimensional grasping point information to obtain three-dimensional grasping point information. Then the fixture can perform grasping based on this three-dimensional grasping point information. If the point cloud of the material box is used, the height obtained from the point cloud may be the same as that of the bottle or different from that of the bottle. If they are different, an adjustment value can be preset according to the height difference between the two. After obtaining the height information of the material box and the two-dimensional grasping point information, the three-dimensional grasping point information can be determined by combining this adjustment value. For example, if the height of the box is 10 cm and the adjustment value is -2 cm, the height of the bottle mouth can be calculated as 10 - 2 = 8 cm, and then combined with the X-axis and Y-axis information of the grasping point to obtain the three-dimensional grasping point information.

[0092] The robot or fixture mentioned in the above embodiments may include various general fixtures. General fixtures refer to fixtures with standardized structures and a wide range of applications. For example, three-jaw chucks and four-jaw chucks used in lathes, vise and dividing heads used in milling machines, etc. Also, according to the clamping power source used by the fixture, the fixture can be divided into manually clamped fixtures, pneumatically clamped fixtures, hydraulically clamped fixtures, pneumatic-hydraulic combined clamped fixtures, electromagnetic fixtures, vacuum fixtures, etc., or other bionic instruments capable of picking up items. The present invention does not limit the specific type of the fixture, as long as it can achieve the item grasping operation.

[0093] In addition, it should be noted that although each embodiment of the present invention has a specific combination of features, further combinations and cross-combinations of these features between embodiments are also feasible.

[0094] According to the above embodiments, first, the present invention can obtain the grasping point information of the item to be grasped by means of other image data such as color pictures in the case where the point cloud of the item to be grasped cannot be obtained, so that the robot or fixture can directly rely on the grasping point information to grasp the item to be grasped without relying on the point cloud of the item, effectively solving the problem of item grasping in the environment where the point cloud is missing; second, the present invention proposes three methods for correcting the mask and obtaining the two-dimensional grasping point information, so that in the case where an accurate item mask cannot be obtained, the inaccurate mask can be corrected to obtain an accurate mask as much as possible and obtain the grasping point information, effectively avoiding the problem that the inaccurate grasping point caused by the inaccurate mask leads to inaccurate grasping or dropping during grasping; third, the present invention proposes a method for automatically converting the input two-dimensional grasping point information into three-dimensional grasping point information by the robot. This method can automatically obtain the reference information that can convert the two-dimensional grasping point into a three-dimensional grasping point according to the environmental characteristics of the item to be grasped, and obtain the grasping point information based on the reference information for the robot to grasp. This solution enables the complete grasping point information to be supplemented based on the existing information in the case where the grasping point information is missing, and reduces manual intervention.

[0095] Figure 5 Fig. shows a grasping point information acquisition device according to another embodiment of the present invention, the device comprising:

[0096] An image acquisition module 400, configured to acquire a non-point cloud image including the item to be grasped, that is, to implement step S100;

[0097] A mask generation module 410, configured to process the non-point cloud image to obtain a mask of the item to be grasped, that is, to implement step S110;

[0098] A mask processing module 420, configured to process the mask of the item to be grasped to obtain grasping point association information, that is, to implement step S120;

[0099] A grasping point information generation module 430, configured to obtain grasping point information for controlling the robot to grasp the item to be grasped based on the grasping point association information, that is, to implement step S130.

[0100] Figure 6 Fig. shows an image data processing device according to another embodiment of the present invention, the device comprising:

[0101] An image data receiving module 500, configured to receive image data including the item to be processed, that is, to implement step S200;

[0102] A mask generation module 510, configured to identify the item to be processed from the image data and generate a mask of the item to be processed, that is, to implement step S210;

[0103] A mask processing module 520, which is used to perform morphological processing on the mask of the item to be processed generated, that is, to implement step S220;

[0104] A processing module 530, which is used to further process the mask after morphological processing to obtain the corrected mask of the item to be processed and / or the grasping point association information, that is, to implement step S230.

[0105] Figure 7 The grasping point information acquisition device according to another embodiment of the present invention is shown. The device includes:

[0106] A reference object information acquisition module 600, which is used to acquire the reference object information of the item to be grasped, that is, to implement step S300;

[0107] A two-dimensional information acquisition module 610, which is used to acquire the two-dimensional grasping point information of the item to be grasped, that is, to implement step S310;

[0108] A reference information acquisition module 620, which is used to process the reference object information of the item to be grasped, acquire reference information, the reference information includes information that the two-dimensional grasping point information does not have, and can convert two-dimensional information into three-dimensional information, that is, to implement step S320;

[0109] A grasping point information generation module 630, which is used to generate three-dimensional grasping point information of the item to be grasped based on the reference information and the two-dimensional grasping point information of the item to be grasped, that is, to implement step S330.

[0110] The above Figures 5 - 7 In the device embodiments shown above, only the main functions of the modules are described. All functions of each module correspond to the corresponding steps in the method embodiments. The working principles of each module can also be referred to the descriptions of the corresponding steps in the method embodiments, which will not be elaborated here. In addition, although the corresponding relationship between the functions of the functional modules and the method is defined in the above embodiments, those skilled in the art can understand that the functions of the functional modules are not limited to the above corresponding relationship, that is, a specific functional module can also implement other method steps or a part of the method steps. For example, the above embodiments describe that the grasping point information generation module 630 is used to implement the method of step S330. However, according to the actual situation, the grasping point information generation module 630 can also be used to implement the method of step S300, S310 or S320 or a part of the method.

[0111] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method of any of the above embodiments. It should be noted that the computer program stored in the computer-readable storage medium of the embodiments of the present application can be executed by the processor of an electronic device. In addition, the computer-readable storage medium can be a storage medium built into the electronic device or a storage medium that can be plugged into the electronic device. Therefore, the computer-readable storage medium of the embodiments of the present application has high flexibility and reliability.

[0112] Figure 8 FIG. shows a schematic structural diagram of an electronic device according to an embodiment of the present invention. The electronic device can be a control system / electronic system configured in an automobile, a mobile terminal (such as a smart mobile phone, etc.), a personal computer (PC, such as a desktop computer or a notebook computer, etc.), a tablet computer, a server, etc. The specific embodiments of the present invention do not limit the specific implementation of the electronic device.

[0113] As Figure 8 shown, the electronic device may include: a processor 1202, a communication interface 1204, a memory 1206, and a communication bus 1208.

[0114] Wherein:

[0115] The processor 1202, the communication interface 1204, and the memory 1206 communicate with each other through the communication bus 1208.

[0116] The communication interface 1204 is used to communicate with network elements of other devices such as clients or other servers.

[0117] The processor 1202 is used to execute the program 1210, and specifically can execute the relevant steps in the above method embodiments.

[0118] Specifically, the program 1210 may include program code, and the program code includes computer operation instructions.

[0119] The processor 1202 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention. One or more processors included in the electronic device can be of the same type of processor, such as one or more CPUs; or different types of processors, such as one or more CPUs and one or more ASICs.

[0120] A memory 1206 for storing a program 1210. The memory 1206 may include a high-speed RAM memory and may also include a non-volatile memory, such as at least one disk memory.

[0121] The program 1210 can be downloaded and installed from a network through a communication interface 1204, and / or installed from a removable medium. When the program is executed by a processor 1202, the processor 1202 can be caused to perform various operations in the above method embodiments. Generally speaking, the disclosure of the present invention includes:

[0122] A method for obtaining grasping point information, comprising:

[0123] Obtaining a non-point cloud image including an item to be grasped;

[0124] Processing the non-point cloud image to obtain a mask of the item to be grasped;

[0125] Processing the mask of the item to be grasped to obtain grasping point association information;

[0126] Based on the grasping point association information, obtaining grasping point information for controlling a robot to grasp the item to be grasped.

[0127] Optionally, the non-point cloud image includes a color image.

[0128] Optionally, the processing of the non-point cloud image includes processing the non-point cloud image based on deep learning.

[0129] Optionally, the processing of the non-point cloud image based on deep learning includes pre-constructing a deep learning network and inputting the non-point cloud image into the deep learning network for processing.

[0130] Optionally, the processing of the mask of the item to be grasped includes performing morphological processing on the mask of the item to be grasped.

[0131] Optionally, the grasping point association information includes two-dimensional information of the grasping point.

[0132] Optionally, the grasping point includes the center point of the graspable area of the item.

[0133] Optionally, the grasping point information includes three-dimensional information of the grasping point.

[0134] Optionally, obtaining the grasping point information for controlling the robot to grasp the item to be grasped based on the grasping point association information includes presetting the dimension information missing from the grasping point association information, and obtaining the grasping point information for controlling the robot to grasp the item to be grasped based on the grasping point association information and the preset dimension information missing from the grasping point association information.

[0135] A grasping point information acquisition device includes:

[0136] An image acquisition module, configured to acquire a non-point cloud image including the item to be grasped;

[0137] A mask generation module, configured to process the non-point cloud image to obtain a mask of the item to be grasped;

[0138] A mask processing module, configured to process the mask of the item to be grasped to obtain grasping point association information;

[0139] A grasping point information generation module, configured to obtain the grasping point information for controlling the robot to grasp the item to be grasped based on the grasping point association information.

[0140] Optionally, the non-point cloud image includes a color image.

[0141] Optionally, the mask generation module processes the non-point cloud image based on deep learning.

[0142] Optionally, a deep learning network is pre-constructed, and the mask generation module inputs the non-point cloud image into the deep learning network for processing.

[0143] Optionally, the mask processing module performs morphological processing on the mask of the item to be grasped.

[0144] Optionally, the grasping point association information includes two-dimensional information of the grasping point.

[0145] Optionally, the grasping point includes the center point of the graspable area of the item.

[0146] Optionally, the grasping point information includes three-dimensional information of the grasping point.

[0147] Optionally, the dimension information missing from the grasping point association information is preset, and the grasping point information generation module obtains the grasping point information for controlling the robot to grasp the item to be grasped based on the grasping point association information and the preset dimension information missing from the grasping point association information.

[0148] A grasping point information acquisition method includes:

[0149] Obtaining the reference object information of the item to be grasped;

[0150] Obtain the two-dimensional grasping point information of the item to be grasped;

[0151] Process the reference object information of the item to be grasped to obtain reference information, where the reference information includes information not possessed by the two-dimensional grasping point information and can convert two-dimensional information into three-dimensional information;

[0152] Generate the three-dimensional grasping point information of the item to be grasped based on the reference information and the two-dimensional grasping point information of the item to be grasped.

[0153] Optionally, the reference object information includes the point cloud and / or depth map of the reference object.

[0154] Optionally, the reference object has a qualified point cloud.

[0155] Optionally, the reference object includes other items to be grasped and / or bins.

[0156] Optionally, the two-dimensional grasping point information includes the X-axis information and Y-axis information of the grasping point.

[0157] Optionally, the reference information includes Z-axis information.

[0158] Optionally, the generating the three-dimensional grasping point information of the item to be grasped based on the reference information and the two-dimensional grasping point information of the item to be grasped includes presetting a reference information adjustment value, adjusting the reference information using the reference information adjustment value, and then generating the three-dimensional grasping point information of the item to be grasped based on the adjusted reference information and the two-dimensional grasping point information of the item to be grasped.

[0159] Optionally, the grasping point includes the center point of the graspable area of the item.

[0160] A device for obtaining grasping point information, comprising:

[0161] A reference object information acquisition module for obtaining the reference object information of the item to be grasped;

[0162] A two-dimensional information acquisition module for obtaining the two-dimensional grasping point information of the item to be grasped;

[0163] A reference information acquisition module for processing the reference object information of the item to be grasped to obtain reference information, where the reference information includes information not possessed by the two-dimensional grasping point information and can convert two-dimensional information into three-dimensional information;

[0164] A grasping point information generation module for generating the three-dimensional grasping point information of the item to be grasped based on the reference information and the two-dimensional grasping point information of the item to be grasped.

[0165] Optionally, the reference object information includes the point cloud and / or depth map of the reference object.

[0166] Optionally, the reference object has qualified point cloud.

[0167] Optionally, the reference object includes other items to be grasped and / or a material box.

[0168] Optionally, the two-dimensional grasping point information includes the X-axis information and Y-axis information of the grasping point.

[0169] Optionally, the reference information includes Z-axis information.

[0170] Optionally, a preset reference information adjustment value is set. After the grasping point information generation module adjusts the reference information using the reference information adjustment value, three-dimensional grasping point information of the item to be grasped is generated based on the adjusted reference information and the two-dimensional grasping point information of the item to be grasped.

[0171] Optionally, the grasping point includes the center point of the graspable area of the item.

[0172] An image data processing method includes:

[0173] Receiving image data including an item to be processed;

[0174] Identifying the item to be processed from the image data and generating a mask of the item to be processed;

[0175] Performing morphological processing on the generated mask of the item to be processed;

[0176] Performing further processing on the mask after morphological processing to obtain a corrected mask of the item to be processed and / or grasping point association information.

[0177] Optionally, the morphological processing includes morphological dilation processing.

[0178] Optionally, the performing further processing on the mask after morphological processing to obtain a corrected mask of the item to be processed and / or grasping point association information includes:

[0179] Obtaining the circumscribed rectangle of the mask after dilation processing;

[0180] Generating an inscribed circle of the circumscribed rectangle based on the circumscribed rectangle of the mask;

[0181] Obtaining the corrected mask of the item to be processed and / or grasping point association information based on the inscribed circle.

[0182] Optionally, the obtaining the circumscribed rectangle of the mask after dilation processing includes: generating 4 corner points of the circumscribed rectangle based on the mask after dilation processing, and then generating the circumscribed rectangle based on the corner points.

[0183] Optionally, further processing the morphologically processed mask to obtain a corrected mask and / or grasping point association information of the item to be processed, including:

[0184] Processing the dilated mask using a circle detection algorithm;

[0185] Obtaining a corrected mask and / or grasping point association information of the item to be processed based on the processing result of the circle detection algorithm.

[0186] Optionally, the circle detection algorithm includes: circle Hough transform algorithm, random Hough transform algorithm, and / or random circle detection algorithm.

[0187] Optionally, further processing the morphologically processed mask to obtain a corrected mask and / or grasping point association information of the item to be processed, including:

[0188] Based on a pre-stored template of the item to be processed, processing the dilated mask using a template matching algorithm;

[0189] Obtaining a corrected mask and / or grasping point association information of the item to be processed based on the processing result of the template matching algorithm.

[0190] Optionally, the matching algorithm includes a shape-based matching algorithm.

[0191] Optionally, identifying the item to be processed from the image data and generating a mask of the item to be processed includes: processing the image data based on deep learning to identify the item to be processed and generating a mask of the item to be processed.

[0192] An image data processing device, including:

[0193] An image data receiving module, configured to receive image data including an item to be processed;

[0194] A mask generation module, configured to identify the item to be processed from the image data and generate a mask of the item to be processed;

[0195] A mask processing module, configured to perform morphological processing on the generated mask of the item to be processed;

[0196] A processing module, configured to further process the morphologically processed mask to obtain a corrected mask and / or grasping point association information of the item to be processed.

[0197] Optionally, the morphological processing includes morphological dilation processing.

[0198] Optionally, the processing module is specifically configured to:

[0199] Obtain the circumscribed rectangle of the mask after dilation processing;

[0200] Generate an inscribed circle of the circumscribed rectangle based on the circumscribed rectangle of the mask;

[0201] Obtain the correction mask and / or the grasping point association information of the item to be processed based on the inscribed circle.

[0202] Optionally, the obtaining of the circumscribed rectangle of the mask after dilation processing includes: generating 4 corner points of the circumscribed rectangle based on the mask after dilation processing, and then generating the circumscribed rectangle based on the corner points.

[0203] Optionally, the processing module is specifically configured to:

[0204] Process the mask after dilation processing using a circle detection algorithm;

[0205] Obtain the correction mask and / or the grasping point association information of the item to be processed based on the processing result of the circle detection algorithm.

[0206] Optionally, the circle detection algorithm includes: the circular Hough transform algorithm, the probabilistic Hough transform algorithm, and / or the random circle detection algorithm.

[0207] Optionally, the processing module is specifically configured to:

[0208] Process the mask after dilation processing using a template matching algorithm based on a template of the item to be processed saved in advance;

[0209] Obtain the correction mask and / or the grasping point association information of the item to be processed based on the processing result of the template matching algorithm.

[0210] Optionally, the matching algorithm includes a shape-based matching algorithm.

[0211] Optionally, the identifying the item to be processed from the image data and generating the mask of the item to be processed includes: processing the image data based on deep learning to identify the item to be processed and generating the mask of the item to be processed.

[0212] An image data processing method includes:

[0213] Receive image data including the item to be processed;

[0214] Identify the item to be processed from the image data and generate the mask of the item to be processed;

[0215] Perform morphological processing on the generated mask of the item to be processed;

[0216] Process the mask after morphological processing using a template matching algorithm based on a template of the item to be processed saved in advance;

[0217] Obtain the correction mask and / or the grasping point association information of the item to be processed based on the processing result of the template matching algorithm.

[0218] Optionally, the morphological processing includes morphological dilation processing.

[0219] Optionally, the matching algorithm includes a shape-based matching algorithm.

[0220] Optionally, the grasping point includes the center point of the graspable area of the item.

[0221] Optionally, identifying the item to be processed from the image data and generating the mask of the item to be processed includes: processing the image data based on deep learning to identify the item to be processed and generating the mask of the item to be processed.

[0222] Optionally, the grasping point association information includes two-dimensional information of the grasping point.

[0223] An image data processing device, comprising:

[0224] An image data receiving module, configured to receive image data including an item to be processed;

[0225] A mask generation module, configured to identify the item to be processed from the image data and generate a mask of the item to be processed;

[0226] A mask processing module, configured to perform morphological processing on the generated mask of the item to be processed;

[0227] A processing module, configured to process the mask after morphological processing using a template matching algorithm based on a template of the item to be processed pre-stored, and obtain the correction mask and / or the grasping point association information of the item to be processed based on the processing result of the template matching algorithm.

[0228] Optionally, the morphological processing includes morphological dilation processing.

[0229] Optionally, the matching algorithm includes a shape-based matching algorithm.

[0230] Optionally, the grasping point includes the center point of the graspable area of the item.

[0231] Optionally, identifying the item to be processed from the image data and generating the mask of the item to be processed includes: processing the image data based on deep learning to identify the item to be processed and generating the mask of the item to be processed.

[0232] Optionally, the grasping point association information includes two-dimensional information of the grasping point.

[0233] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the said embodiments or examples are included in at least one embodiment or example of this application. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiments or examples. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0234] Any process or method description shown in a flowchart or described in other ways herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a way that may not be in the order shown or discussed, including in a substantially simultaneous manner according to the functions involved or in a reverse order, which should be understood by those skilled in the technical field to which the embodiments of this application belong.

[0235] The logic and / or steps represented in a flowchart or described in other ways herein, for example, can be considered as an ordered list of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processing module, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with such instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.

[0236] The processor can be a Central Processing Unit (CPU), or can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0237] It should be understood that each part of the embodiments of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with suitable combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0238] Those of ordinary skill in the art of the present technology can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program. The said program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0239] In addition, in each of the embodiments of the present application, the functional units can be integrated in a processing module, or each unit can exist physically alone, or two or more units can be integrated in a module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0240] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc.

[0241] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. An image data processing method, characterized in that, including: Receiving image data including an item to be processed; Identifying the item to be processed from the image data and generating a mask of the item to be processed; Performing morphological processing on the generated mask of the item to be processed; Based on a pre-stored template of the item to be processed, using a template matching algorithm to process the morphologically processed mask; The template of the item to be processed is a template of the grasping area of the item to be processed; Obtaining a corrected mask of the item to be processed based on the processing result of the template matching algorithm, and calculating grasping point association information based on the corrected mask; The processing based on a pre-stored template of the item to be processed and using a template matching algorithm to process the morphologically processed mask includes: based on the pre-stored template, using a matching algorithm to match the template within the mask area.

2. The image data processing method according to claim 1, wherein The morphological processing includes morphological dilation processing.

3. The image data processing method according to claim 1, characterized in that: The matching algorithm includes a shape-based matching algorithm.

4. The image data processing method according to any one of claims 1-3, characterized in that: The grasping point includes the center point of the graspable area of the item.

5. The image data processing method according to any one of claims 1-3, characterized in that: The identifying the item to be processed from the image data and generating a mask of the item to be processed includes: processing the image data based on deep learning to identify the item to be processed and generating a mask of the item to be processed.

6. The image data processing method according to any one of claims 1 to 3, characterized in that: The grasping point association information includes two-dimensional information of the grasping point.

7. An image data processing device, characterized in that, including: An image data receiving module for receiving image data including an item to be processed; A mask generating module for identifying the item to be processed from the image data and generating a mask of the item to be processed; A mask processing module for performing morphological processing on the generated mask of the item to be processed; A processing module for, based on a pre-stored template of the item to be processed, using a template matching algorithm to process the morphologically processed mask, obtaining a corrected mask of the item to be processed based on the processing result of the template matching algorithm, and calculating grasping point association information based on the corrected mask; the template of the item to be processed is a template of the grasping area of the item to be processed; The processing based on a pre-stored template of the item to be processed and using a template matching algorithm to process the morphologically processed mask includes: based on the pre-stored template, using a matching algorithm to match the template within the mask area.

8. The image data processing device according to claim 7, wherein The morphological processing includes morphological dilation processing.

9. The image data processing device according to claim 7, wherein: The matching algorithm includes a shape-based matching algorithm.

10. The image data processing apparatus according to any one of claims 7-9, characterized in that: The grasping point includes the center point of the graspable area of the item.

11. The image data processing apparatus according to any one of claims 7-9, characterized in that: The identifying the item to be processed from the image data and generating a mask of the item to be processed includes: processing the image data based on deep learning to identify the item to be processed and generating a mask of the item to be processed.

12. The image data processing device according to any one of claims 7-9, characterized in that: The grasping point association information includes two-dimensional information of the grasping point.

13. An electronic device, characterized in that, including: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the image data processing method according to any one of claims 1 to 6 is implemented.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the image data processing method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Method, device, equipment for segmentation of lesion in biological image and storage medium

    CN108682015A

  • Method and device for disorderly sorting of strip-shaped agricultural products

    CN112883881A

  • Image mask generation method and device, electronic equipment and storage medium

    CN113378948A