Robot vision positioning method and device based on target detection

By using a single-stage image model based on object detection to detect the support feet of the pallet and combining them with depth images, the problems of high computational cost and environmental impact in pallet detection in existing technologies are solved, and rapid, accurate positioning and stable detection of the pallet position are achieved.

CN115375759BActive Publication Date: 2026-04-17BEIJING GRAY TIANZE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING GRAY TIANZE TECHNOLOGY CO LTD
Filing Date
2021-05-19
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In existing technologies, pallet detection methods based on point cloud processing have high computational costs, expensive data annotation, and are easily affected by the environment, resulting in unstable detection and making it difficult to achieve rapid and accurate detection of industrial pallets.

Method used

A single-stage image model based on object detection is adopted. The pallet's support legs are detected by IR image, and the distance between the pallet and the forklift robot is evaluated by combining depth image. The route is then planned to pick up the pallet and the goods.

Benefits of technology

It enables rapid and accurate pallet positioning, reduces computational costs and data labeling difficulties, improves detection stability and speed, and adapts to warehouse environments with unstable lighting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115375759B_ABST
    Figure CN115375759B_ABST
Patent Text Reader

Abstract

This invention provides a robot visual localization method based on target detection, comprising: an image acquisition module that records in real time IR and depth images of the forklift robot's picking direction; an edge computing module that detects the pallet's border information based on the IR image; a false detection filtering module that filters the pallet's border information based on prior knowledge and establishes internal relationships within the borders; and a distance determination module that performs visual localization of the pallet based on the internal relationships within the borders and the depth image. This invention also provides an apparatus for this method, using a more easily deployable single-stage image target detection model to detect the pallet's support legs, enabling the robot to locate the pallet. Then, by combining the depth image with the detection results, the distance between the pallet and the forklift robot is evaluated, and a route is planned to pick up the pallet along with the goods on it.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial transportation, and in particular to a robot vision localization method and apparatus based on target detection. Background Technology

[0002] For automated forklift products, achieving rapid and accurate detection of industrial pallets is a problem that must be addressed, and this problem is prevalent in the field of smart warehousing.

[0003] The vast majority of existing algorithms in this field are based on point cloud processing, such as the early PointNet and VoxelINet. The main drawbacks of current point cloud-based processing methods are: (1) High computational cost. Due to its density, point cloud processing requires a large amount of computational support. In the tray detection stage, higher computing equipment costs must be paid to achieve real-time detection. (2) High data and price. The cost of labeling 3D point clouds is more than 10 times that of labeling ordinary images, and the labeling difficulty is high, affecting the product iteration time. (3) Easily affected by the environment. Depth cameras are easily affected by external light, reflection and other factors, resulting in unstable depth information and adversely affecting the detection effect.

[0004] Unlike existing technologies, this application provides a robot visual positioning method and device based on target detection. It uses a single-stage image target detection model that is easier to deploy to detect the support legs of the pallet so that the robot can find the position of the pallet. Then, by combining the depth image with the detection results, the distance between the pallet and the forklift robot is evaluated, and a route is planned to realize the forklifting of the pallet and the goods on the pallet. Summary of the Invention

[0005] (I) Purpose of the Invention

[0006] To overcome at least one of the defects in the existing technology, the present invention provides a robot visual localization method and device based on target detection. It uses a single-stage image target detection algorithm that is easier to deploy to detect the support legs of the pallet so that the robot can find the position of the pallet. Then, by combining the depth image with the detection results, the distance between the pallet and the forklift robot is evaluated, and a route is planned to realize the forklifting of the pallet and the goods on the pallet.

[0007] (II) Technical Solution

[0008] As a first aspect of the present invention, the present invention discloses a robot visual localization method based on target detection, comprising:

[0009] The image acquisition step involves real-time recording of IR and depth images of the forklift robot in the forking direction.

[0010] The edge computing step detects the tray's border information based on the IR image;

[0011] The false detection filtering step involves filtering the tray's border information based on prior knowledge and establishing internal connections within the border.

[0012] The distance determination step involves visually positioning the tray based on the internal connections of the border and the depth image.

[0013] In one possible implementation, the tray's border information includes: the overall outer border of the tray and the support leg borders.

[0014] In one possible implementation, the edge computing step includes a single-stage target detection model; the edge computing step inputs the IR image into the single-stage target detection model and outputs the positioning information of the tray, thereby obtaining the overall outline and support leg outline of the tray based on the positioning information of the tray.

[0015] In one possible implementation, the prior knowledge includes the relationship between the support feet and the relationship between the overall outline and the support feet; the internal connection of the outline includes the connection between the overall outline and the support feet.

[0016] In one possible implementation, visual positioning of the pallet includes determining the distance and angle of the pallet relative to the forklift robot.

[0017] As a second aspect of the present invention, the present invention discloses a robot visual localization device based on target detection, comprising:

[0018] The image acquisition module is used to record IR images and depth images of the forklift robot in the picking direction in real time;

[0019] An edge computing module is used to detect the tray's border information based on the IR image;

[0020] The false detection filtering module is used to filter the edge information of the tray based on prior knowledge and establish internal connections within the edge.

[0021] The distance determination module is used to visually locate the tray based on the internal connections of the border and the depth image.

[0022] In one possible implementation, the tray's border information includes: the overall outer border of the tray and the support leg borders.

[0023] In one possible implementation, the edge computing module includes a single-stage target detection model; the edge computing module inputs the IR image into the single-stage target detection model and outputs the positioning information of the tray, thereby obtaining the overall outline and support leg outline of the tray based on the positioning information of the tray.

[0024] In one possible implementation, the prior knowledge includes the relationship between the support feet and the relationship between the overall outline and the support feet; the internal connection of the outline includes the connection between the overall outline and the support feet.

[0025] In one possible implementation, visual positioning of the pallet includes determining the distance and angle of the pallet relative to the forklift robot.

[0026] (III) Beneficial Effects

[0027] This invention provides a robot visual localization method and apparatus based on target detection. An image acquisition module records real-time IR and depth images of the forklift robot's picking direction. An edge computing module detects the pallet's border information based on the IR image. A false detection filtering module filters the pallet's border information based on prior knowledge and establishes internal relationships within the borders. A distance determination module performs visual localization of the pallet based on the internal relationships within the borders and the depth image. Using a more easily deployable single-stage image target detection model, the pallet's support legs are detected, enabling the robot to locate the pallet. Then, by combining the depth image with the detection results, the distance between the pallet and the forklift robot is evaluated, and a route is planned to pick up the pallet along with the goods on it. Attached Figure Description

[0028] The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain and illustrate the present invention, and should not be construed as limiting the scope of protection of the present invention.

[0029] Figure 1 This is a flowchart of a robot visual localization method based on target detection provided by the present invention.

[0030] Figure 2 This is a schematic diagram of the structure of a robot vision positioning device based on target detection provided by the present invention. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions in the embodiments of this invention will be described in more detail below with reference to the accompanying drawings.

[0032] It should be noted that, throughout the accompanying drawings, the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions. The described embodiments are some, but not all, embodiments of the present invention. Unless otherwise specified, the embodiments and features described in this application can be combined with each other. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0033] In the description of this invention, it should be understood that the terms "center", "longitudinal", "lateral", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting the scope of protection of this invention.

[0034] The following is for reference. Figure 1 The first embodiment of the robot visual localization method based on target detection provided by the present invention is described in detail below. For example... Figure 1 As shown, the robot visual localization method based on target detection provided in this embodiment mainly includes: image acquisition step, edge calculation step, false detection filtering step, and distance determination step.

[0035] This application uses a one-stage image target detection method to detect the support legs of an industrial pallet, enabling a forklift robot to locate the pallet. Then, by combining the depth image with the pallet position, the distance between the pallet and the forklift robot is evaluated, facilitating subsequent route planning for picking up the goods.

[0036] The image acquisition step records the IR image and depth image of the forklift robot in the forking direction in real time. The image acquisition step includes a depth camera, which can be installed at the upper middle position of the forklift robot's forks and slightly angled towards the ground. At this time, the depth camera can record the IR image and depth image of the forking direction in real time.

[0037] If the IR image and depth image acquired by the depth camera are frontal images of the pallet in the fork-handling direction, then multiple support legs of the pallet and complete rectangular sockets can be observed.

[0038] In the image acquisition step, the IR image can be sent to the edge computing step via the robot operating system ROS.

[0039] The edge computing step involves detecting the tray's border information based on the IR image. The edge computing step includes an edge computing device, which can detect the tray's border information based on the IR image.

[0040] The pallet's border information includes: the overall outer border of the pallet and the support leg border.

[0041] Edge computing devices are used to detect the overall outline of the tray and the multiple support leg outlines of the tray within the overall outline in the IR image, so as to determine the position of the tray from the IR image.

[0042] The edge computing step includes a single-stage target detection model; the edge computing step inputs the IR image into the single-stage target detection model and outputs the positioning information of the tray, thereby obtaining the overall outer frame and support leg frame of the tray based on the positioning information of the tray.

[0043] The single-stage target detection model includes a BP neural network model; the pre-trained BP neural network model can extract the positioning information of the pallet. First, the single-stage target detection model divides the IR image into multiple grids, and then inputs the IR image divided into multiple grids into the BP neural network model. Then, the BP neural network model outputs the positioning information of the pallet, and the overall outline and support leg outline of the pallet in the IR image can be determined based on the positioning information of the pallet.

[0044] The false positive filtering step filters the tray's edge information based on prior knowledge and establishes internal relationships within the edge. This prior knowledge includes the relationship between the support legs and the relationship between the overall edge shape and the support legs. The internal relationships within the edge include the connection between the overall edge shape and the support legs. The relationship between the support legs can be that the three support legs should be in a straight line; the relationship between the overall edge shape and the support legs can be that the overall edge shape should include the support legs.

[0045] In the false detection filtering step, prior knowledge can be used to filter false detection objects. Based on the principle that the three support legs should be in a straight line and that the overall outline should include the support legs, unreasonable detection objects are filtered out. Furthermore, a connection is established between the support legs and the overall outline that conforms to prior knowledge to determine the tray's pose and shape characteristics. The unreasonable detection objects are those whose overall outline and support leg edges do not conform to the principles that the three support legs should be in a straight line and that the overall outline should include the support legs.

[0046] The distance determination step involves visually locating the pallet based on the internal connections of the border and the depth image. This visual positioning of the pallet includes determining the distance and angle of the pallet relative to the forklift robot.

[0047] The size of each support leg and its position in the image can be determined by the connection between the overall outline and the support legs; the position of the tray in the image can be obtained by the position of each support leg in the image.

[0048] Based on the position of each support leg in the image and combined with depth map information, the distance between the entire pallet and the forklift robot is determined.

[0049] Based on the position of each support leg in the image, combined with depth map information, and according to the relationship between the support legs, the angle of the pallet relative to the forklift robot is obtained.

[0050] The forklift robot uses the distance between itself and the pallet and the angle of the pallet relative to itself to achieve visual positioning of the pallet, and moves to the appropriate position to perform the pallet picking operation, transporting the pallet along with the goods inside to the destination.

[0051] This application can detect a series of pallets with different shape features placed in different spatial positions in real time, and realize the recognition of the position, pose and shape features of the pallets, thereby providing accurate visual positioning for forklift robots.

[0052] This application uses a single-stage target detection model based on IR images: it can be quickly trained and deployed for various pallets; it makes IR images more robust in uncertain warehouse environments such as unstable lighting; and using two-dimensional images instead of three-dimensional point cloud information also makes the detection speed faster and the detection results more stable.

[0053] This application records IR and depth images of the forklift robot's picking direction in real time through an image acquisition step; an edge computing step detects the pallet's border information based on the IR image; a false detection filtering step filters the pallet's border information based on prior knowledge and establishes internal relationships within the borders; and a distance determination step performs visual positioning of the pallet based on the internal relationships within the borders and the depth image. This application uses a more easily deployable single-stage image target detection model to detect the pallet's support legs, enabling the robot to locate the pallet. Then, by combining the depth image with the detection results, the distance between the pallet and the forklift robot is evaluated, and a route is planned to pick up the pallet along with the goods on it.

[0054] The following is for reference. Figure 2 The first embodiment of the robot visual localization device based on target detection provided by the present invention is described in detail below. For example... Figure 2As shown, the robot vision positioning device based on target detection provided in this embodiment mainly includes: an image acquisition module, an edge computing module, a false detection filtering module, and a distance determination module.

[0055] This application uses a one-stage image target detection method to detect the support legs of an industrial pallet, enabling a forklift robot to locate the pallet. Then, by combining the depth image with the pallet position, the distance between the pallet and the forklift robot is evaluated, facilitating subsequent route planning for picking up the goods.

[0056] The image acquisition module is used to record the IR image and depth image of the forklift robot in the forking direction in real time. The image acquisition module includes a depth camera, which can be installed at the upper middle position of the forklift robot's forks and slightly angled towards the ground. At this time, the depth camera can record the IR image and depth image of the forking direction in real time.

[0057] If the IR image and depth image acquired by the depth camera are frontal images of the pallet in the fork-handling direction, then multiple support legs of the pallet and complete rectangular sockets can be observed.

[0058] The image acquisition module can send the IR image to the edge computing module via the robot operating system ROS.

[0059] An edge computing module is used to detect the border information of the tray based on the IR image; the edge computing module includes an edge computing device, which can detect the border information of the tray based on the IR image.

[0060] The tray's border information includes the tray's overall outline border and the support leg borders. An edge computing device is used to detect the tray's overall outline border and multiple support leg borders within the overall outline border in the IR image, in order to determine the tray's position from the IR image.

[0061] The edge computing module includes a single-stage target detection model; the edge computing module inputs the IR image into the single-stage target detection model and outputs the positioning information of the tray, thereby obtaining the overall outline and support leg outline of the tray based on the positioning information of the tray.

[0062] The single-stage target detection model includes a BP neural network model; the pre-trained BP neural network model can extract the positioning information of the pallet. First, the single-stage target detection model divides the IR image into multiple grids, and then inputs the IR image divided into multiple grids into the BP neural network model. Then, the BP neural network model outputs the positioning information of the pallet, and the overall outline and support leg outline of the pallet in the IR image can be determined based on the positioning information of the pallet.

[0063] The false positive filtering module is used to filter the tray's edge information based on prior knowledge and establish internal relationships within the edge. The prior knowledge includes the relationship between the support legs and the relationship between the overall edge shape and the support legs. The internal relationships within the edge include the connection between the overall edge shape and the support legs. The relationship between the support legs can be that the three support legs should be in a straight line; the relationship between the overall edge shape and the support legs can be that the overall edge shape should include the support legs.

[0064] In the false detection filtering module, prior human knowledge can be used to filter false detection objects. Based on the principle that the three support legs should be in a straight line and that the overall outline should include the support legs, unreasonable detection objects are filtered out. Furthermore, a connection is established between the support legs and the overall outline that conforms to prior knowledge to determine the tray's pose and shape characteristics. The unreasonable detection objects are those whose overall outline and support leg edges do not conform to the principles that the three support legs should be in a straight line and that the overall outline should include the support legs.

[0065] A distance determination module is used to visually locate the pallet based on the internal connections of the border and the depth image. The visual positioning of the pallet includes determining the distance and angle of the pallet relative to the forklift robot.

[0066] The size of each support leg and its position in the image can be determined by the connection between the overall outline and the support legs; the position of the tray in the image can then be obtained based on the positions of the support legs.

[0067] Based on the position of each support leg in the image and combined with depth map information, the distance between the entire pallet and the forklift robot is determined.

[0068] Based on the position of each support leg in the image, combined with depth map information, and according to the relationship between the support legs, the angle of the pallet relative to the forklift robot is obtained.

[0069] The forklift robot uses the distance between itself and the pallet and the angle of the pallet relative to itself to achieve visual positioning of the pallet, and moves to the appropriate position to perform the pallet picking operation, transporting the pallet along with the goods inside to the destination.

[0070] This application can detect a series of pallets with different shape features placed in different spatial positions in real time, and realize the recognition of the position, pose and shape features of the pallets, thereby providing accurate visual positioning for forklift robots.

[0071] This application uses a single-stage target detection model based on IR images: it can be quickly trained and deployed for various pallets; it makes IR images more robust in uncertain warehouse environments such as unstable lighting; and using two-dimensional images instead of three-dimensional point cloud information also makes the detection speed faster and the detection results more stable.

[0072] This application uses an image acquisition module to record IR and depth images of the forklift robot's picking direction in real time; an edge computing module detects the pallet's border information based on the IR image; a false detection filtering module filters the pallet's border information based on prior knowledge and establishes internal relationships within the border; and a distance determination module performs visual positioning of the pallet based on the internal relationships within the border and the depth image. This application uses a more easily deployable single-stage image target detection model to detect the pallet's support legs, enabling the robot to locate the pallet. Then, by combining the depth image with the detection results, the distance between the pallet and the forklift robot is evaluated, and a route is planned to pick up the pallet along with the goods on it.

[0073] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A robot vision positioning method based on target detection, characterized in that, Includes the following steps: The image acquisition step involves real-time recording of IR and depth images of the forklift robot in the forking direction. The edge computing step detects the tray's border information based on the IR image; the tray's border information includes: the overall outer border of the tray and the support leg borders; The false positive filtering step filters the tray's border information based on prior knowledge and establishes internal relationships within the border. The prior knowledge includes the relationship between the support legs and the relationship between the tray's overall border and the support legs. The internal relationships within the border include the connection between the tray's overall border and the support legs. Specifically, the relationship between the support legs is that the three support legs should form a straight line; the relationship between the tray's overall border and the support legs is that the tray's overall border should include the support legs. The distance determination step, based on the internal relationships of the border and the depth image, performs visual positioning of the tray, including: The size of each support leg and its position in the image are determined based on the connection between the overall outer frame of the tray and the support legs; the position of the tray in the image is obtained based on the position of each support leg in the image. Based on the position of each support leg in the image and combined with depth map information, the distance between the pallet and the forklift robot is determined. Based on the position of each support leg in the image, combined with depth map information, and according to the relationship between the support legs, the angle of the pallet relative to the forklift robot is obtained.

2. The method of robot vision positioning according to claim 1, wherein, The edge computing step includes a single-stage target detection model; the edge computing step inputs the IR image into the single-stage target detection model and outputs the positioning information of the tray, thereby obtaining the overall outline and support leg outline of the tray based on the positioning information of the tray.

3. A robot vision positioning apparatus based on target detection, characterized by, include: The image acquisition module is used to record IR images and depth images of the forklift robot in the picking direction in real time; An edge computing module is used to detect the border information of the tray based on the IR image; the border information of the tray includes: the overall outer border of the tray and the border of the supporting feet; The false positive filtering module is used to filter the edge information of the tray based on prior knowledge and establish internal relationships within the edge. The prior knowledge includes the relationship between the support legs and the relationship between the overall outer edge of the tray and the support legs. The internal relationships within the edge include the connection between the overall outer edge of the tray and the support legs. Specifically, the relationship between the support legs is that the three support legs should be in a straight line; the relationship between the overall outer edge of the tray and the support legs is that the overall outer edge of the tray should include the support legs. A distance determination module is used to visually locate the tray based on the internal relationships of the border and the depth image; including: The size of each support leg and its position in the image are determined based on the connection between the overall outer frame of the tray and the support legs; the position of the tray in the image is obtained based on the position of each support leg in the image. Based on the position of each support leg in the image and combined with depth map information, the distance between the pallet and the forklift robot is determined. Based on the position of each support leg in the image, combined with depth map information, and according to the relationship between the support legs, the angle of the pallet relative to the forklift robot is obtained.

4. The robotic vision positioning apparatus of claim 3, wherein, The edge computing module includes a single-stage target detection model; the edge computing module inputs the IR image into the single-stage target detection model and outputs the positioning information of the tray, thereby obtaining the overall outer frame and support leg frame of the tray based on the positioning information of the tray.

Citation Information

Patent Citations

  • Warehouse box body identification and positioning method based on contour features

    CN111507390A

  • Tray detecting and positioning method based on depth camera

    CN111986185A