Key point detection-based driving target positioning method and unstacking method

By combining key point detection networks and lidar, high-precision target positioning and stacking/unstacking of bridge cranes in complex environments were achieved, solving the problem of poor adaptability of traditional 2D image processing methods and improving the accuracy and real-time performance of stacking/unstacking.

CN114972488BActive Publication Date: 2026-01-13MATRIXTIME ROBOTICS (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210592472.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-27
Publication Date
2026-01-13
Estimated Expiration
2042-05-27

AI Technical Summary

Technical Problem

Existing methods for destacking and stacking bridge cranes are subject to significant changes in lighting conditions in factory or open-air environments. This results in poor adaptability and low positioning accuracy of traditional 2D image processing methods, making it difficult to achieve automated destacking and stacking.

Method used

This method combines a key point detection network and LiDAR. By acquiring images through a camera, the key point pixel coordinates of the target object are detected. The grasping center position of the target object is calculated by combining geometric transformation and calibration methods. The point cloud data acquired by LiDAR is used for 3D positioning to achieve precise positioning and grasping of the target object.

Benefits of technology

It improves the environmental adaptability and accuracy of target positioning, reduces the consumption of computing resources, and enables fast reasoning and high-precision depalletizing and palletizing operations, and is suitable for special lighting and occlusion conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114972488B_ABST
    Figure CN114972488B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of travelling crane automation, and particularly relates to a travelling crane target object positioning method based on key point detection and a disassembling and stacking method. The travelling crane target object positioning method based on key point detection obtains images through a camera installed on the travelling crane, detects key point pixel coordinates of a target object in the images by using a key point detection network, converts the key point pixel coordinates to image center pixel coordinates, calculates a pixel coordinate position of a target object grabbing center through geometric transformation, converts the pixel coordinate position of the grabbing center to a travelling crane spreader physical coordinate position, and completes target object positioning. The application detects key points of a travelling crane target object by using a key point detection network, can accurately position the target object when calculating the grabbing center, realizes automatic travelling crane disassembling and stacking operation of the target object, and has good environmental adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of vehicle automation technology, specifically relating to a method for locating vehicle target objects and a method for unpacking and stacking based on key point detection. Background Technology

[0002] Overhead cranes (or gantry cranes) are an important type of general-purpose machinery. As heavy-duty logistics equipment, they are used in most factory workshops, such as those in mining, steel, non-ferrous metals, and machinery manufacturing industries. Overhead crane operators need to possess certain technical skills and work long hours in the cramped cab. In production enterprises such as manganese mines, they also work in potentially toxic and hazardous environments. In recent years, recruiting overhead crane operators has become increasingly difficult, and labor costs have been rising year by year.

[0003] Therefore, automation and unmanned operation technologies for bridge cranes have always been a key focus of industry development. Foreign companies such as Siemens, ABB, Demag, and Konecranes, as well as domestic companies such as Taiyuan Heavy Industry, Henan Mining Heavy Industry, Weihua Heavy Industry, and Baosteel Group, are all conducting research on related technologies and products. De-palletizing is a very common operation scenario in overhead crane operations. Automated overhead cranes generally use vision-based measurement for palletizing, using vision or LiDAR to locate the objects to be de-palletized. The problem with existing methods is that traditional 2D image processing methods are highly dependent on the environment and lighting, while the work site is generally a factory or open-air environment, resulting in poor algorithm adaptability and ultimately low positioning accuracy, making palletizing difficult. Summary of the Invention

[0004] The purpose of this invention is to provide a method for locating target objects and a method for depalletizing and palletizing objects by a crane based on key point detection. This method can accurately locate target objects and realize automated crane depalletizing and palletizing operations on target objects, while also having good environmental adaptability.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] A method for locating a target object on a traveling vehicle based on key point detection is characterized by: acquiring images using a camera installed on the traveling vehicle; using a key point detection network to detect the key point pixel coordinates of the target object in the image; transforming the key point pixel coordinates to pixel coordinates in a coordinate system with the image center pixel coordinates as the origin; calculating the pixel coordinate position of the target object's grasping center through geometric transformation; and transforming the pixel coordinate position of the grasping center to the physical coordinate position of the traveling vehicle's crane to complete the target object location.

[0007] Furthermore, the transformation of keypoint pixel coordinates to pixel coordinates in a coordinate system with the image center pixel coordinates as the origin includes:

[0008] When the keypoint pixel coordinates output by the keypoint detection network are ( ), where i is the index of the key point in the output;

[0009] Convert to image center pixel coordinates ( The pixel coordinates in a coordinate system with the origin as the origin are:

[0010]

[0011]

[0012] Furthermore, the pixel coordinate position of the target object grasping center ( )include:

[0013] The output of the keypoint detection network also includes the keypoint pixel coordinates. Corresponding category information , They represent the first The key points are located at the top left, bottom left, top right, and bottom right key points of the rectangular target object. The length and width of the rectangular target object are: ;

[0014] When the keypoint detection network detects only one keypoint The center of the capture is calculated as follows:

[0015]

[0016]

[0017] When two key points are detected , The center of the capture is calculated as follows:

[0018]

[0019]

[0020] When three key points are detected The calculation of the capture center and the detection of the two key points mentioned above Same time;

[0021] When all four key points are detected, The center of the capture is calculated as follows:

[0022]

[0023]

[0024] Furthermore, the pixel coordinates of the center of the capture will be... (Through coefficients) Convert to the physical coordinate position (x, y) of the crane crane;

[0025] Conversion factor The following information is obtained beforehand through calibration: 2D markers are placed on the ground, and the horizontal distance between the center of the marker and the center of the camera is [missing information]. Obtain the pixel coordinates of the marker center in the image. Calculate pixel coordinates With the coordinates of the image center ( The pixel coordinate difference is () ),get:

[0026]

[0027] The horizontal displacement between the grab center and the camera center can be obtained. ):

[0028]

[0029]

[0030] Complete the positioning of the target object's grasping center.

[0031] Furthermore, the method also includes positioning the target object in terms of height, specifically including:

[0032] Point cloud data is acquired by a lidar installed on the vehicle, with the lidar and camera arranged side by side at the same height.

[0033] Point cloud data of the target object's upper surface is extracted using a rectangular region composed of key points. The height of the centroid of the upper surface is then calculated through point cloud segmentation. ;

[0034] The three-dimensional grasping center point of the target object is ( ).

[0035] Furthermore, regarding the coefficients The pixel position is corrected as it changes with height at different heights to obtain new conversion coefficients. for:

[0036]

[0037] The three-dimensional grasping center point of the target object is ( ),in

[0038]

[0039]

[0040] A method for depalletizing and palletizing cranes based on key point detection, characterized in that: based on the above-mentioned three-dimensional grasping center point ( ), control the vehicle's movement to the corresponding horizontal position ( The lifting device descends to height z to grab and place the target object, thus achieving stacking operations.

[0041] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0042] (1) Compared with traditional 2D visual localization methods, the method of key point detection network used in this invention has better environmental adaptability, while taking into account fast inference speed, improving the real-time performance of system detection, and effectively reducing GPU computing resource consumption;

[0043] (2) The present invention provides accurate positioning of the target object, which greatly improves the accuracy of depalletizing and palletizing, and meets the requirements of high accuracy in depalletizing and palletizing.

[0044] (3) The key point detection post-processing algorithm in this invention has good adaptability and can be applied to situations where key points cannot be fully detected due to special lighting, occlusion, etc. It still has high target positioning accuracy and can be used with lifting equipment to achieve target object grabbing and stacking operations. Attached Figure Description

[0045] Figure 1 This is a schematic diagram of the hardware installation in the embodiment.

[0046] Figure 2 This is a schematic diagram showing the key point numbers and dimensions of the steel plate in the embodiment.

[0047] Figure 3 This is a schematic diagram illustrating the relationship between camera pixels and distance in the embodiment.

[0048] Figure 4 This is a flowchart of the overhead crane depalletizing and palletizing method based on key point detection in the embodiment. Detailed Implementation

[0049] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to specific examples. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0050] This embodiment takes the destacking and stacking of rectangular steel plates as an example. It uses key points to detect the features of the target object (the steel plate), designs a specialized post-processing algorithm, and ultimately obtains the precise gripping position, enabling the crane to accurately grip and destacking the target object. Figure 4 As shown.

[0051] I. Hardware Introduction

[0052] like Figure 1 As shown, this solution uses a camera mounted on a trolley to acquire real-time video stream data of the target object. The camera uses a common 4MP fixed-focus IPC network camera video stream as the 2D image input. The camera and LiDAR (at the same height) are mounted on the trolley's crossbeam, vertically downwards at a height H above the ground. Both the camera and LiDAR are calibrated with the lifting device (after calibration, the camera's center corresponds to the lifting device's center, which can then be used as the lifting device's center). For trolleys with a working height less than 10m, a fixed-focus camera is used; for working heights greater than 10m, a zoom camera is used. The camera transmits video streams via RTSP, and the target object image is obtained using GPU decoding. The radar data is acquired directly via standard TCP communication.

[0053] II. Detection of Key Points of Target Object

[0054] This approach uses the EvoPose2D network (keypoint detection network) for target object keypoint detection. By collecting real and simulated images of the target object, a dataset of image samples is generated using data augmentation methods. Keypoints at the four corners of the target object are then labeled within these image samples. Finally, the EvoPose2D network is trained.

[0055] The EvoPose2D network is set up with a runtime environment and model deployment. Images captured by the camera are input into the EvoPose2D network, which can then output classification information and pixel coordinate information of key points of the target object.

[0056] like Figure 2 As shown, the steel plate in this embodiment (length and width are...) There are at most 4 key points (meters), and the pixel coordinates of the detected key points are (meters). ), corresponding classification information , For the key point index of the output, , They represent the first The key points are located at the upper left key point, lower left key point, upper right key point, and lower right key point on the steel plate.

[0057] III. Post-processing Algorithm

[0058] 1. Set the key pixel coordinates ( Convert to image center pixel coordinates ( In a coordinate system with 0 as the origin (the initial origin of the pixel coordinates is at the top left corner), the pixel coordinates are:

[0059]

[0060]

[0061] This involves transforming the key pixel coordinates to pixel coordinates in a coordinate system with the image center pixel coordinates as the origin.

[0062] 2. Due to the influence of ambient light and the possibility of partial occlusion of the target object, in order to improve the robustness of localization, the pixel coordinates of the gripping center of the steel plate are obtained through geometric calculations for the key points detected by the EvoPose2D network. ).

[0063] (1) When the keypoint detection network detects only one keypoint The gripping center of the steel plate is calculated as follows:

[0064]

[0065]

[0066] (2) When two key points are detected , (Indicates any two indexed) (Sum of results) The gripping center of the steel plate is calculated as follows:

[0067]

[0068]

[0069] (3) When three key points are detected At this point, a key point is missing. The gripping center of the steel plate can be calculated using the diagonally opposite key point, which is consistent with the value in step (2) above. The time is the same.

[0070] (4) When all 4 key points are detected The gripping center of the steel plate is calculated as follows:

[0071]

[0072]

[0073] 3. Using the rectangular area formed by key points, extract the lidar point cloud data of the upper surface of the steel plate, and calculate the height of the centroid of the upper surface by point cloud segmentation. , That is, the vertical distance from the lidar (camera) to the upper surface of the steel plate, such as Figure 3 As shown.

[0074] 4. Grasp the pixel coordinates of the center of the steel plate ( (Through coefficients) Convert to the physical coordinate position (x, y) of the overhead crane.

[0075] Conversion factor The following information is obtained beforehand through calibration: 2D markers are placed on the ground, and the horizontal distance between the center of the marker and the center of the camera is [missing information]. Obtain the pixel coordinates of the marker center in the image. Calculate pixel coordinates With the coordinates of the image center ( The pixel coordinate difference is () ),get:

[0076]

[0077] According to the principles of camera imaging, as the height of the target object changes, the displacement represented by each pixel in the image in the x and y directions also varies, such as... Figure 3 As shown. According to the coefficients Calculate the new transformation coefficients for pixel position as a function of height at different heights. for:

[0078]

[0079] The horizontal displacement between the target object's grasping center and the camera center can be obtained. ):

[0080]

[0081]

[0082] Complete the positioning of the steel plate gripping center.

[0083] Based on step 4 above, the final gripping center point of the lifting device can be obtained as ( ), control the vehicle to move to the corresponding position ( Once the lifting device descends to height z, it can grab and place the target object, thereby realizing the stacking operation.

[0084] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for locating a vehicle target based on key point detection, characterized in that: Images are acquired by a camera installed on the crane, and key point detection network is used to detect the key point pixel coordinates of the target object in the image. The key point pixel coordinates are transformed into pixel coordinates in a coordinate system with the image center pixel coordinates as the origin. The pixel coordinate position of the target object's grasping center is calculated through geometric transformation. The pixel coordinate position of the grasping center is transformed into the physical coordinate position of the crane's crane, thus completing the target object positioning. The process of transforming the keypoint pixel coordinates to pixel coordinates in a coordinate system with the image center pixel coordinates as the origin includes: When the key point pixel coordinates output by the key point detection network are , i is the index of the output key point; Transform to image center pixel coordinates Pixel coordinates in the coordinate system with the origin at the center of the image are: calculating a pixel coordinate position of a target object grasp center comprising: The output of the key point detection network further includes key point pixel coordinates Corresponding classification information respectively represent that the i-th key point is located at the upper left key point, the lower left key point, the upper right key point, and the lower right key point of the rectangular target object, The length and width of the rectangular target object are ​​ When the keypoint detection network only detects one keypoint, , the grasp center is computed as follows: ; when two key points are detected, , the center of the grasp is calculated as follows: ; when three key points are detected, , the calculation of the grasp center is the same as described above for the detection of two key points ; When all four key points are detected, The center of the capture is calculated as follows: ; capture the pixel coordinates of the center Through coefficients Convert to the physical coordinate position (x, y) of the crane crane; Conversion factor The following information is obtained beforehand through calibration: 2D markers are placed on the ground, and the horizontal distance between the center of the marker and the center of the camera is [missing information]. Obtain the pixel coordinates of the marker center in the image. Calculate pixel coordinates Image center coordinates The pixel coordinate difference is ,get: The horizontal displacement between the grab center and the camera center can be obtained. : This completes the positioning of the target object's grasping center.

2. The method for locating driving targets based on key point detection according to claim 1, characterized in that: The method also includes positioning the target object in terms of height, specifically including: Point cloud data is acquired by a lidar installed on the vehicle, with the lidar and camera arranged side by side at the same height. Using a rectangular area composed of key points, point cloud data of the upper surface of the target object is extracted, and the height z of the centroid of the upper surface is calculated through point cloud segmentation. The three-dimensional grasping center point of the target object is .

3. The method for locating driving targets based on key point detection according to claim 2, characterized in that: For coefficients The pixel position is corrected as it changes with height at different heights to obtain new conversion coefficients. for: The three-dimensional grasping center point of the target object is then... ,in 。 4. A crane-based method for depalletizing and palletizing based on key point detection, characterized in that: The three-dimensional grasping center point according to claim 2 or 3 Control the vehicle's movement to the corresponding horizontal position. The lifting device descends to height z to grab and place the target object, thus achieving stacking operations.

Citation Information

Patent Citations

  • A line-of-sight estimation method based on key point matching

    CN109344714A

  • Disordered aliasing workpiece grabbing method and system based on key point prediction network

    CN113580149A