Object positioning method based on image recognition and two-dimensional laser radar
By fusing image recognition and two-dimensional lidar data, the problem of insufficient positioning accuracy and real-time performance of objects in the prior art is solved, and the precise recognition and positioning of specific objects in complex environments is realized, which reduces the computational complexity and improves the system processing efficiency.
Patent Information
- Application Number
- CN202510253082.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-11
AI Technical Summary
Existing object positioning methods are difficult to achieve high-precision positioning in complex environments, especially in multi-object scenes, which are prone to misidentification or missed detection. Traditional single sensor systems cannot obtain comprehensive obstacle information, and the calculation complexity is high and the real-time performance is insufficient.
Fusion image recognition and two-dimensional lidar data, initially identify targets and calculate angles through the camera, and use radar to obtain range data for clustering, reducing the calculation amount, and improving positioning accuracy and efficiency.
It realizes accurate identification and positioning of specific objects in complex environments, reduces the system's calculation amount, improves processing efficiency, has the ability to identify objects with lighter colors, and enhances the real-time and positioning accuracy of the system.
Smart Images

Figure CN120298643A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multi-sensor fusion obstacle avoidance, and particularly relates to an object positioning method based on image recognition and two-dimensional lidar. Background Art
[0002] With the rapid development of intelligent mobile platforms such as driverless vehicles and drones, environmental perception and obstacle avoidance capabilities have become core functional requirements. Efficient and accurate environmental perception and flexible and reliable obstacle avoidance capabilities are the prerequisites for intelligent mobile platforms to make autonomous decisions and ensure safe navigation during operation.
[0003] Environmental perception and obstacle avoidance capabilities generally rely on various sensor detection technologies to extract information on the width, height, and depth of obstacles. Traditional single-sensor systems, such as those relying solely on radar or cameras, often have certain limitations. For example, a single two-dimensional lidar can provide accurate depth and width information but cannot obtain the height information of an object. If a three-dimensional lidar is used, the cost and data processing capacity requirements are increased, and accurate detailed information of the object cannot be obtained either. In the case of a camera, when the lighting conditions change greatly or in a complex background, it may be impossible to recognize an object, and an ordinary monocular camera cannot obtain distance information. Therefore, multi-sensor fusion obstacle avoidance technology has gradually attracted wide attention from researchers. The multi-sensor fusion obstacle avoidance technology aims to combine the advantages of different sensors, complement the defects, so as to obtain more comprehensive obstacle information and achieve more accurate environmental perception.
[0004] When there are multiple objects in the camera's field of view, existing object positioning methods are difficult to achieve high-precision positioning of each object, which easily leads to misidentification or missed detection. Traditional edge detection and binarization algorithms have poor robustness to background noise in complex scenes, especially insufficient detection capabilities for objects with lighter colors or low contrast, and are prone to misidentification. Although the clustering-based object positioning method can perform multi-object recognition, due to its high computational complexity and the clustering effect being easily affected by environmental clutter, the positioning accuracy and real-time performance cannot meet actual requirements. Summary of the Invention
[0005] To solve the above technical problems, by fusing image recognition and two-dimensional lidar data, the present application can achieve precise recognition and positioning of specific objects in a relatively complex environment and reduce the computational amount of the system. A method for object positioning based on image recognition and two-dimensional lidar is proposed, and the specific technical solution is as follows:
[0006] A method for object positioning based on image recognition and two-dimensional lidar, comprising the following steps:
[0007] Obtain the original image through a camera;
[0008] Preliminarily identify the target on the original image and generate a recognition box;
[0009] Based on the horizontal position of the recognition box in the original image, calculate the included angle θ of the target relative to the camera view;
[0010] Obtain the range radar data in the direction of the included angle θ through the radar;
[0011] Cluster the range radar data in the direction of the included angle θ to obtain the accurate position data of the target.
[0012] More specifically, use yolov5 to preliminarily identify the target on the original image and generate a recognition box.
[0013] In other embodiments, NanoDet, YOLO-Fastest, MobileNet-SSD, DETR, MMDetection, etc. can be used to replace yolov5 to preliminarily identify the target on the original image and generate a recognition box.
[0014] Among them, the method for calculating the included angle θ of the target relative to the camera view based on the horizontal position of the recognition box in the original image is as follows:
[0015]
[0016] Among them, θ is the included angle between the robot and the center of the detected object;
[0017] C is a constant;
[0018] Δ pixel is the absolute value of the horizontal pixel difference between the center of the recognition box and the center of the camera field of view;
[0019] f 相机 is the focal length of the camera lens.
[0020] The beneficial effects of the present invention are as follows:
[0021] Utilize the YOLOv5 algorithm to perform real-time recognition on the image. When there are many messy items in the image obtained by the camera, it can accurately detect the position of only the specified object, and at the same time has the ability to recognize objects with lighter colors; after obtaining the general orientation of the target object through image recognition, only a small amount of radar data needs to be clustered, without complex clustering and filtering operations, thereby effectively reducing the computational complexity and improving the processing efficiency of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is the original image obtained by the monocular camera and the marked image of the watermelon recognized from the image;
[0023] Figure 2Schematic diagram of the original image obtained by the monocular camera and the recognition frame of the detected object;
[0024] Figure 3 Relative position state of the detected object and the monocular camera during actual shooting;
[0025] Figure 4 State diagram of the radar obtaining data of the detected object and its vicinity at the θ angle. Specific implementation manners
[0026] In the following description, certain specific details are set forth in order to provide a thorough understanding of various embodiments. However, those skilled in the art should understand that the present invention may be practiced without these details. In other instances, well-known structures have not been shown or described in detail to avoid unnecessarily obscuring the description of the embodiments. Unless the context otherwise requires, throughout the specification and the appended claims, the word "comprising" shall be interpreted in an open, inclusive sense, i.e., as "including, but not limited to".
[0027] The object positioning method based on image recognition and two-dimensional lidar provided by the present application is specifically as follows:
[0028] Step 1, Obtain the original image by the monocular camera, and identify the position of the specified item within the camera's field of view through yolov5, as Figure 1 shown is the original image obtained by the monocular camera and the marked image of the watermelon recognized from the image;
[0029] The original image obtained by the monocular camera has a fixed pixel size, such as a pixel image of 320*640. After identifying the target object through yolov5, the generated recognition frame will also have a measurable pixel size.
[0030] Step 2, Calculate the horizontal pixel difference between the center of the recognition frame and the center of the camera's field of view, and further calculate the angle θ between the orientation of the monocular camera and the detected object in combination with the focal length of the camera lens;
[0031] As Figure 2 shown, it is a schematic diagram of the original image obtained by the monocular camera and the recognition frame of the detected object. The image obtained by the monocular camera is 320*640 pixels. The recognition frame is a rectangle, and the pixel positions on both sides of the rectangle are obtained. The horizontal pixel positions on both sides of the rectangle are 100 and 200, and the pixel point position at the center of the rectangle frame is 150, which is 320 - 150 = 70 pixels different from the center of the camera's field of view; Figure 3 Shows the relative position state of the detected object and the monocular camera during actual shooting.
[0032] The calculation of the angle θ of the detected object is based on the following formula:
[0033]
[0034] Wherein, θ is the included angle between the robot and the center of the detected object;
[0035] C is a constant; due to the differences in the optical characteristics of different cameras, the values of C for different cameras are generally different, but it always holds that tanθ is proportional to ; for a fixed camera, the value of C can be obtained by taking multiple values of tanθ and and then performing linear fitting.
[0036] Δ pixel is the absolute value of the horizontal pixel difference between the center of the recognition frame and the center of the camera's field of view.
[0037] Step 3: According to the calculated included angle θ of the detected object, use the radar to obtain the data of the detected object and its vicinity at the θ angle, obtain the distance information from the points on the surface of the detected object to the 2D lidar, and further calculate the coordinate values of the surface of the object to be measured relative to the center of the 2D lidar.
[0038] Figure 4 The figure shows the state diagram of the radar obtaining the data of the detected object and its vicinity at the θ angle. The coordinate values of the surface of the object to be measured relative to the center of the 2D lidar are calculated according to the following formula:
[0039]
[0040] Wherein, d i is the distance from the surface of the object to be measured to the center of the 2D lidar;
[0041] θ i is the actual included angle of the surface of the object to be measured relative to the center of the 2D lidar.
[0042] Step 4: Perform DBSCAN clustering on the obtained series of coordinate values. After removing the coordinate point data that is not on the surface of the detected object, the remaining coordinate points are considered to be on the surface of the object to be measured, and finally the position of the specified object to be measured relative to the 2D lidar is obtained.
[0043] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it.
Claims
1. An object positioning method based on image recognition and 2D lidar, characterized in that, It includes the following steps: Obtain the original image through a camera; Preliminarily identify the target on the original image and generate a recognition box; Based on the horizontal position of the recognition box in the original image, calculate the included angle θ of the target relative to the camera view; Obtain the range radar data in the direction of the included angle θ through a radar; Cluster the range radar data in the direction of the included angle θ to obtain the accurate position data of the target.
2. The object positioning method based on image recognition and two-dimensional lidar according to claim 1, wherein Preliminarily identify the target on the original image and generate a recognition box by one of yolov5, NanoDet, YOLO-Fastest, MobileNet-SSD, DETR, MMDetection; 3. The object positioning method based on image recognition and two-dimensional lidar according to claim 1, characterized in that, The method for calculating the included angle θ of the target relative to the camera view based on the horizontal position of the recognition box in the original image is: where θ is the included angle between the robot and the center of the detected object; C is a constant; Δ pixel is the absolute value of the horizontal pixel difference between the center of the recognition frame and the center of the camera's field of view; f 相机 is the focal length of the camera lens.