A method and device for vision-based object recognition and localization

Through visual recognition and positioning methods, combined with deep learning and camera hole imaging model, multi-category and multi-pose recognition and three-dimensional positioning of tableware are achieved, solving the problems of low efficiency and hygiene quality of tableware, and reducing labor costs.

CN114842468BActive Publication Date: 2025-08-01SHENZHEN YUTONG INNOVATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210554413.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-19
Publication Date
2025-08-01
Estimated Expiration
2042-05-19

AI Technical Summary

Technical Problem

The prior art is difficult to achieve accurate identification and positioning of various categories and different placement postures of tableware, especially the three-dimensional positioning of the center point of tableware, and the recognition effect of transparent tableware is poor, resulting in low manual sorting efficiency, high cost and difficult to ensure hygiene quality.

Method used

The vision-based target recognition and positioning method is adopted, and two-dimensional images are processed through deep learning networks, combined with the camera small hole imaging model to correct and compensate the initial center point, obtain the three-dimensional center point of the tableware, and grab it using a pneumatic suction cup.

Benefits of technology

It realizes accurate identification and three-dimensional positioning of tableware of various categories and different placement postures, reduces the intensity of manual labor, and improves sorting efficiency and sanitary quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114842468B_ABST
    Figure CN114842468B_ABST
Patent Text Reader

Abstract

The present application relates to a method and device for visual-based target recognition and positioning, including the following steps: S1. Process the two-dimensional image of the target to be sorted, and identify the area of the target and the information of the target in the two-dimensional image; S2. Extract the initial center point of the target from the area of the target according to the information of the target; S3. Use the height or diameter of the target as prior information, and correct and compensate the initial center point of the target by using the camera pinhole imaging model to obtain the planar center point of the target; then use the thickness, height or diameter of the target as the value of the planar center point of the target in the third dimension Z-axis direction to obtain the three-dimensional center point of the target. Through the present invention, the recognition and three-dimensional positioning of tableware of different categories and placement postures can be realized, and it has wide applicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of automation technology, and particularly relates to a method and device for target recognition and positioning based on vision. Background Art

[0002] In the industry of centralized cleaning and disinfection of tableware, after the tableware is cleaned, disinfected and dried, it is also necessary to sort different types of tableware in an orderly manner. Due to the large quantity and variety of tableware, the labor intensity of the whole working process is relatively high. At present, the sorting operation of tableware is mainly completed manually, which has problems such as low efficiency and high labor cost. Moreover, manual sorting is likely to cause secondary pollution to the disinfected tableware, thus affecting the hygienic quality of the tableware. Realizing the automatic sorting of tableware is an effective way to solve the above problems. To achieve the automatic sorting of tableware, it is first necessary to achieve the accurate recognition and positioning of tableware.

[0003] Patent application CN109684942A discloses a fully automatic tableware sorting method based on visual recognition. Based on a two-dimensional visual image, it uses the YOLO V3 deep learning network to identify and position tableware, and then uses a robotic arm for sorting. However, this application only realizes the recognition of tableware categories and the positioning of the center point of the tableware, and cannot provide the pose information of the tableware (upright, upside down, horizontal and the direction when horizontal, etc.). At present, most of the existing technologies can only realize the positioning of the center point of the tableware or the positioning and recognition of a single type of tableware. Some positioning requires the use of a depth camera, which is costly and has a poor recognition effect on transparent tableware such as glass cups. Summary of the Invention

[0004] The purpose of the present invention is to solve the deficiencies existing in the prior art, and provide a method and device for target recognition and positioning based on vision. Through the present invention, it is possible to realize the recognition of various types of tableware with different placement postures and the three-dimensional positioning of the center point of the tableware (including the x, y, z coordinates and direction of the grasping point), which has wide applicability.

[0005] Specifically, on the one hand, an embodiment of the present invention provides a method for target recognition and positioning based on vision, including the following steps:

[0006] S1. Process the two-dimensional image of the target to be sorted, and identify the area of the target and the information of the target in the two-dimensional image;

[0007] S2. According to the information of the target, extract the initial center point of the target (i.e., a two-dimensional plane point, which is also the intersection point of the x-axis and the y-axis) from the area of the target;

[0008] S3. Using the height or diameter of the target as prior information, correct and compensate the initial center point of the target by using the camera pinhole imaging model to obtain the planar center point of the target (i.e., a two-dimensional planar point); then use the thickness, height or diameter of the target as the value of the planar center point of the target in the third dimension Z-axis direction to obtain the three-dimensional center point of the target (i.e., the center position of the target); the three-dimensional center point of the target is the target grasping point when using a pneumatic suction cup to grasp the target to be sorted.

[0009] In step S1,

[0010] As a preferred embodiment, the processing is performed through a deep learning network. In this application, the deep learning network is a pre-trained deep learning network offline.

[0011] As a preferred embodiment, the deep learning network is a network such as YOLO, R-CNN or SSD. As a preferred embodiment, the region of the target is defined as the circumscribed rectangle of the outer contour of the target.

[0012] As a preferred embodiment, the information of the target includes the target category, the target placement posture and the local area of the target.

[0013] Further, the target types include plates, bowls, tea cups or glass cups, etc.

[0014] Further, the target placement postures include upright, upside down or lying horizontally, etc.

[0015] Further, the local area of the target includes the bottom, the body part and the top, etc.

[0016] In the embodiment of this application, the information of the target is one of the following options: upright plate, upright bowl, upright tea cup, upright glass cup, upside down plate, upside down bowl, upside down tea cup, upside down glass cup, horizontally lying tea cup body, horizontally lying glass cup body, horizontally lying tea cup bottom or horizontally lying glass cup bottom, etc. Some tableware such as bowls can be further subdivided into large bowls, small bowls, etc.

[0017] In step S2,

[0018] As a preferred embodiment, if the information of the target in step S1 is an upright tableware, extract the opening circle of the upright tableware from the area of the tableware, and use the center of the circle as the initial center point of the upright tableware (i.e., a two-dimensional planar point);

[0019] Or, directly use the deep learning network to identify the area of the opening circle of the tableware (i.e., the circumscribed rectangle of the opening circle) from the area of the tableware, and use the center of the area of the opening circle as the initial center point of the upright tableware (i.e., a two-dimensional planar point).

[0020] Further, the upright tableware includes an upright plate, an upright bowl, an upright teacup or an upright glass.

[0021] As a preferred embodiment, if the information of the target in step S1 is the upside-down tableware, a circle at the bottom of the upside-down tableware is fitted and extracted from the area of the tableware, and the center of the circle is used as the initial center point (i.e., a two-dimensional plane point) of the upside-down tableware;

[0022] Alternatively, the area of the bottom circle of the tableware (i.e., the circumscribed rectangle of the bottom circle) is directly identified from the area of the tableware by using a deep learning network, and the center of the area of the bottom circle is used as the initial center point (i.e., a two-dimensional plane point) of the upside-down tableware.

[0023] Further, the upside-down tableware includes an upside-down plate, an upside-down bowl, an upside-down teacup or an upside-down glass.

[0024] As a preferred embodiment, if the information of the target in step S1 is the lying tableware, the center point of the cup body area of the lying tableware is used as the initial center point (i.e., a two-dimensional plane point) of the lying tableware.

[0025] In step S3,

[0026] As a preferred embodiment, for the upright tableware, taking the height of the upright tableware as prior information, a camera pinhole imaging model is used to correct and compensate the initial center point of the upright tableware to obtain the planar center point (i.e., a two-dimensional plane point) of the upright tableware; then, taking the thickness of the upright tableware as the value of the planar center point of the upright tableware in the third dimension Z-axis direction, the three-dimensional center point of the upright tableware is obtained.

[0027] In the present application, since the tableware has a certain height, there will be a certain deviation between the obtained initial center point of the upright tableware and the actual center point of the tableware. By using a camera pinhole imaging model to correct and compensate the initial center point, a planar center point (two-dimensional plane point) closer to the actual center point of the tableware is obtained, which is more conducive to the positioning of the tableware.

[0028] As a preferred embodiment, for the upside-down tableware, taking the height of the upside-down tableware as prior information, a camera pinhole imaging model is used to correct and compensate the initial center point of the upside-down tableware to obtain the planar center point (i.e., a two-dimensional plane point) of the upside-down tableware; then, taking the height of the upside-down tableware as the value of the planar center point of the upside-down tableware in the third dimension Z-axis direction, the three-dimensional center point of the upside-down tableware is obtained.

[0029] In the present application, due to the fact that the tableware has a certain height, there will be a certain deviation between the initial center point of the inverted tableware obtained and the actual center point of the tableware. By using the camera pinhole imaging model to correct and compensate the initial center point, a planar center point (two-dimensional planar point) closer to the actual center point of the tableware is obtained, which is more conducive to the positioning of the tableware.

[0030] As a preferred embodiment, for the horizontally placed tableware, taking the diameter of the horizontally placed tableware as prior information, the camera pinhole imaging model is used to correct and compensate the initial center point of the horizontally placed tableware to obtain the planar center point (i.e., two-dimensional planar point) of the horizontally placed tableware; then, taking the diameter of the horizontally placed tableware as the value of the planar center point of the horizontally placed tableware in the third-dimensional Z-axis direction, the three-dimensional center point of the horizontally placed tableware is obtained.

[0031] As a preferred embodiment, for the horizontally placed tableware, the cup body of the horizontally placed tableware and the cup bottom of the horizontally placed tableware are paired according to rules such as the distance between the center points, the sizes of the two regions, and the positional relationship of the corner points of the two regions, and the orientation of the opening of the horizontally placed tableware is determined by the connection direction of the center points of the paired cup body region and the cup bottom region.

[0032] As a preferred embodiment, the three-dimensional center point of the tableware is the target grasping point when using a pneumatic suction cup to grasp the tableware to be sorted. Also, according to actual needs, other feature points or regions of the tableware can be obtained based on the geometric shape and other characteristics of the tableware using the above center points.

[0033] On the other hand, an embodiment of the present invention further provides a device for target recognition and positioning by using the above method for target recognition and positioning based on vision.

[0034] Compared with the prior art, the present invention has the following beneficial effects: Through the present invention, the recognition of tableware of various categories and different placement postures and the three-dimensional positioning of the center point of the tableware (including the x, y, z coordinates and direction of the grasping point) can be realized, and it has wide applicability.

[0035] The realization of the object of the present invention, functional features and advantages will be further described in conjunction with embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is a schematic flow chart of the method for target recognition and positioning based on vision according to an embodiment of the present application;

[0037] Figure 2 is Figure 1 a detailed schematic flow chart of the method for target recognition and positioning based on vision. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] Next, in combination with the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0039] It should be noted that if there are directional indications (such as up, down, left, right, front, back, top, bottom...) involved in the embodiments of the present invention, then such directional indications are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the attached drawings). If this specific posture changes, then the directional indications will also change accordingly.

[0040] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, then such descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0041] Aiming at the defects in the prior art that mostly only the positioning of the center point of tableware or the positioning and recognition of single-category tableware can be achieved; some positioning requires the use of a depth camera, with high costs and poor recognition effects on transparent tableware such as glass cups, etc., the present invention provides a method for visual-based target recognition and positioning. Through the present invention, the recognition of various categories of tableware, tableware in different placement postures, and the three-dimensional positioning of the center point of the tableware (including the x, y, z coordinates and direction of the grasping point) can be realized, with wide applicability.

[0042] Specifically, as Figures 1 to 2 shown, the embodiments of the present invention provide a method for visual-based target recognition and positioning, including the following steps:

[0043] S1. Process the two-dimensional image of the target to be sorted, and identify the area of the target and the information of the target in the two-dimensional image;

[0044] S2. According to the information of the target, extract the initial center point of the target (i.e., a two-dimensional plane point, that is, the intersection point of the x coordinate axis and the y coordinate axis) from the area of the target;

[0045] S3. Using the height or diameter of the target as prior information, correct and compensate the initial center point of the target by using the camera pinhole imaging model to obtain the planar center point of the target (i.e., a two-dimensional planar point); then use the thickness, height or diameter of the target as the value of the planar center point of the target in the third-dimensional Z-axis direction to obtain the three-dimensional center point of the target (i.e., the center position of the target).

[0046] In step S1,

[0047] As a preferred embodiment, the processing is performed by a deep learning network. In this application, the deep learning network is a pre-trained offline deep learning network.

[0048] As a preferred embodiment, the deep learning network is a network such as YOLO, R-CNN or SSD, etc.

[0049] As a preferred embodiment, the area of the target is defined as the circumscribed rectangle of the outer contour of the target.

[0050] As a preferred embodiment, the information of the target includes the target type, the target placement posture and the local area of the target.

[0051] Further, the target types include plates, bowls, tea cups or glass cups, etc.

[0052] Further, the target placement postures include upright, upside down or lying horizontally, etc.

[0053] Further, the local area of the target includes the bottom, the body part and the top, etc.

[0054] In the embodiments of this application, the information of the target includes: upright plate, upright bowl, upright tea cup, upright glass cup, upside-down plate, upside-down bowl, upside-down tea cup, upside-down glass cup, lying horizontally tea cup body, lying horizontally glass cup body, lying horizontally tea cup bottom or lying horizontally glass cup bottom, etc. Some tableware such as bowls can be further subdivided into large bowls, small bowls, etc.

[0055] In step S2,

[0056] As a preferred embodiment, if the information of the target in step S1 is an upright tableware, then fit and extract the opening circle of the upright tableware from the area of the tableware, and use the center of the circle as the initial center point of the upright tableware (i.e., a two-dimensional planar point);

[0057] Or, directly use the deep learning network to identify the area of the opening circle of the tableware (i.e., the circumscribed rectangle of the opening circle) from the area of the tableware, and use the center of the area of the opening circle as the initial center point of the upright tableware (i.e., a two-dimensional planar point).

[0058] Further, the upright tableware includes an upright plate, an upright bowl, an upright teacup or an upright glass.

[0059] As a preferred embodiment, if the information of the target in step S1 is the upside-down tableware, the circle at the bottom of the upside-down tableware is fitted and extracted from the area of the tableware, and the center of the circle is used as the initial center point (i.e., a two-dimensional plane point) of the upside-down tableware.

[0060] Alternatively, the area of the circle at the bottom of the tableware (i.e., the circumscribed rectangle of the bottom circle) is directly recognized from the area of the tableware by using a deep learning network, and the center of the area of the bottom circle is used as the initial center point (i.e., a two-dimensional plane point) of the upside-down tableware.

[0061] Further, the upside-down tableware includes an upside-down plate, an upside-down bowl, an upside-down teacup or an upside-down glass.

[0062] As a preferred embodiment, if the information of the target in step S1 is the horizontally placed tableware, the center point of the cup body area of the horizontally placed tableware is used as the initial center point (i.e., a two-dimensional plane point) of the horizontally placed tableware.

[0063] In step S3,

[0064] As a preferred embodiment, for the upright tableware, taking the height of the upright tableware as a priori information, a camera pinhole imaging model is used to correct and compensate the initial center point of the upright tableware to obtain the planar center point (i.e., a two-dimensional plane point) of the upright tableware; then, taking the thickness of the upright tableware as the value of the planar center point of the upright tableware in the third dimension Z-axis direction, the three-dimensional center point of the upright tableware is obtained.

[0065] In this application, since the tableware has a certain height, there will be a certain deviation between the obtained initial center point of the upright tableware and the actual center point of the tableware. By using a camera pinhole imaging model to correct and compensate the initial center point, a planar center point (two-dimensional plane point) closer to the actual center point of the tableware is obtained, which is more conducive to the positioning of the tableware.

[0066] As a preferred embodiment, for the upside-down tableware, taking the height of the upside-down tableware as a priori information, a camera pinhole imaging model is used to correct and compensate the initial center point of the upside-down tableware to obtain the planar center point (i.e., a two-dimensional plane point) of the upside-down tableware; then, taking the height of the upside-down tableware as the value of the planar center point of the upside-down tableware in the third dimension Z-axis direction, the three-dimensional center point of the upside-down tableware is obtained.

[0067] In the present application, since the tableware has a certain height, there will be a certain deviation between the initial center point of the inverted tableware obtained and the actual center point of the tableware. By using the camera pinhole imaging model to correct and compensate the initial center point, a planar center point (two-dimensional planar point) closer to the actual center point of the tableware is obtained, which is more conducive to the positioning of the tableware.

[0068] As a preferred embodiment, for the horizontally placed tableware, using the diameter of the horizontally placed tableware as prior information, the camera pinhole imaging model is used to correct and compensate the initial center point of the horizontally placed tableware to obtain the planar center point (i.e., two-dimensional planar point) of the horizontally placed tableware; then, using the diameter of the horizontally placed tableware as the value of the planar center point of the horizontally placed tableware in the third-dimensional Z-axis direction, the three-dimensional center point of the horizontally placed tableware is obtained.

[0069] As a preferred embodiment, for the horizontally placed tableware, the cup body and the cup bottom of the horizontally placed tableware are paired according to rules such as the distance between the center points, the sizes of the two regions, and the positional relationship of the corner points of the two regions, and the orientation of the opening of the horizontally placed tableware is determined by the connecting line direction of the center points of the paired cup body region and the cup bottom region.

[0070] For horizontally placed tableware (tea cups and glass cups), the opening direction of the horizontally placed tableware can also be determined by the center point of the cup mouth and the center point of the cup bottom of the horizontally placed tableware; or the opening direction of the horizontally placed tableware can be determined by the center point of the cup mouth and the center point of the cup body of the horizontally placed tableware.

[0071] As a preferred embodiment, the three-dimensional center point of the tableware is the target grasping point when using a pneumatic suction cup to grasp the tableware to be sorted. According to actual needs, other feature points or regions of the tableware can also be obtained based on the above center point according to the geometric shape and other characteristics of the tableware.

[0072] On the other hand, the embodiment of the present invention also provides a device for target recognition and positioning by using the above method for target recognition and positioning based on vision.

[0073] Through the present invention, the recognition of various types of tableware and tableware in different placement postures and the three-dimensional positioning of the center point of the tableware (including the x, y, z coordinates and directions of the grasping point) can be realized, and it has wide applicability.

[0074] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structural transformation made under the inventive concept of the present invention by using the content of the specification and drawings of the present invention, or directly / indirectly applied in other related technical fields, is included in the patent protection scope of the present invention.

Claims

1. A method for vision-based object recognition and localization, characterized in that, It includes the following steps: S1. Process the two-dimensional image of the object to be sorted, and identify the area of the object and the information of the object in the two-dimensional image; S2. Extract the initial center point of the object from the area of the object according to the information of the object; S3. Using the height or diameter of the object as prior information, correct and compensate the initial center point of the object by using the camera pinhole imaging model to obtain the planar center point of the object; then use the thickness, height or diameter of the object as the value of the planar center point of the object in the third dimension Z-axis direction to obtain the three-dimensional center point of the object; In step S1, the processing is performed by a deep learning network; the area of the object is defined as the circumscribed rectangle of the outer contour of the object; the information of the object includes the object type, the placement posture of the object and the local area of the object; The object type is tableware, and the tableware includes plates, bowls, tea cups or glass cups; The placement postures of the object include upright, upside down or lying horizontally; The local area of the object includes the bottom, the body part and the top; If the information of the object in step S1 is upright tableware, fit and extract the circle of the opening of the upright tableware from the area of the tableware, and use the center of the circle as the initial center point of the upright tableware; Or, if the information of the object in step S1 is upright tableware, directly use the deep learning network to identify the area of the circle at the bottom of the tableware from the area of the tableware, and use the center of the area of the bottom circle as the initial center point of the upside-down tableware; If the information of the object in step S1 is upside-down tableware, fit and extract the circle at the bottom of the upside-down tableware from the area of the tableware, and use the center of the circle as the initial center point of the upside-down tableware; Or, if the information of the object in step S1 is upside-down tableware, directly use the deep learning network to identify the area of the circle at the bottom of the tableware from the area of the tableware, and use the center of the area of the bottom circle as the initial center point of the upside-down tableware; If the information of the object in step S1 is horizontally lying tableware, use the center point of the cup body area of the horizontally lying tableware as the initial center point of the horizontally lying tableware; Match the cup body and the cup bottom of the horizontally lying tableware, and determine the orientation of the opening of the horizontally lying tableware according to the connection direction of the center points of the paired cup body area and the cup bottom area; The three-dimensional center point of the tableware is the target grasping point when using a pneumatic suction cup to grasp the tableware to be sorted.

2. The method for vision-based target recognition and positioning according to claim 1, characterized in that The deep learning network is a YOLO, R-CNN or SSD network.

3. The method for vision-based object recognition and localization according to claim 1, characterized in that, In step S3, using the height of the upright tableware as prior information, correct and compensate the initial center point of the upright tableware by using the camera pinhole imaging model to obtain the planar center point of the upright tableware; then use the thickness of the upright tableware as the value of the planar center point of the upright tableware in the third dimension Z-axis direction to obtain the three-dimensional center point of the upright tableware.

4. The method for vision-based object recognition and localization according to claim 1, wherein, In step S3, taking the height of the upside-down tableware as prior information, the camera pinhole imaging model is used to correct and compensate the initial center point of the upside-down tableware to obtain the planar center point of the upside-down tableware; then, taking the height of the upside-down tableware as the value of the planar center point of the upside-down tableware in the third-dimensional Z-axis direction, the three-dimensional center point of the upside-down tableware is obtained.

5. The method for vision-based object recognition and positioning according to claim 1, wherein In step S3, taking the diameter of the lying tableware as prior information, the camera pinhole imaging model is used to correct and compensate the initial center point of the lying tableware to obtain the planar center point of the lying tableware; then, taking the diameter of the lying tableware as the value of the planar center point of the lying tableware in the third-dimensional Z-axis direction, the three-dimensional center point of the lying tableware is obtained.

6. An apparatus for vision-based object recognition and localization, characterized in that, Use the method for visual-based target recognition and positioning according to any one of claims 1-5 for target recognition and positioning.

Citation Information

Patent Citations

  • A full-automatic tableware sorting method based on visual identification

    CN109684942A

  • Image processing method and device, image processing system and storage medium

    CN113661513A