Target ranging method, device and equipment integrating monocular camera and high-precision map

By integrating a monocular camera with a high-precision map, the problems of ambient lighting, weather changes, and calibration parameter drift in monocular ranging methods were solved, achieving high-precision 3D ranging under complex road conditions and reducing hardware costs.

CN121702338APending Publication Date: 2026-03-20HIGER
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-27
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing monocular camera target ranging methods are easily affected by changes in ambient lighting, weather, and target appearance, resulting in insufficient stability. They rely on camera calibration parameters and are prone to drift. They also lack effective utilization of road topology, making it difficult to guarantee accuracy under complex road conditions.

Method used

By integrating a monocular camera with a high-precision map, a two-dimensional detection box is obtained through 2D target detection. Global path planning and discretization are then performed using the high-precision map data to generate a set of road segments. The depth value is then calculated using the set of road segments to achieve three-dimensional ranging.

Benefits of technology

It improves the stability and accuracy of monocular ranging, reduces hardware costs, and is suitable for high-precision ranging in complex environments, adapting to complex road conditions such as curves and slopes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121702338A_ABST
    Figure CN121702338A_ABST
Patent Text Reader

Abstract

The invention discloses a target distance measuring method, device and equipment fusing a monocular camera and a high-precision map, and relates to the field of automatic driving. 2D target detection processing is carried out on a current frame image collected by a vehicle-mounted monocular camera, at least one detection target is identified, and a target distance measuring result is obtained; and determining a two-dimensional detection frame corresponding to each detection target under the image coordinate system. And performing global path planning and discretization processing according to the high-precision map data, generating a road line segment set in combination with road boundary information, and converting the road line segment set to an image coordinate system, thereby calculating a depth value of each detection target in a camera coordinate system, and further determining three-dimensional distance measurement information. Because the depth perception of the monocular camera is easily influenced by illumination and texture, the defect that the depth estimation precision of the monocular camera is insufficient is made up through the three-dimensional environment information of the high-precision map, and the spatial constraint of the map provides accurate reference for coordinates, so that the accuracy of three-dimensional positioning and distance measurement is greatly improved. And the reliability of the system in a complex environment is obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of automatic driving, in particular to a target ranging method, device and equipment fusing monocular camera and high-precision map. BACKGROUND

[0002] With the rapid development of automatic driving and advanced auxiliary driving systems, higher requirements are put forward for the accuracy and reliability of environmental perception technology. Among them, target ranging based on vision is one of the key technologies to realize vehicle positioning, path planning and decision control.

[0003] In related technologies, the target ranging method based on monocular camera mainly realizes through target detection in the image combined with camera calibration parameters. This kind of method usually relies on the pixel position of the target in the image, the camera intrinsic parameter and the preset assumption (such as flat ground) to estimate the target distance. Another common solution is to directly estimate the depth information of the target in the image through a deep learning model, but this method has higher requirements for the quality of training data and scene coverage. Therefore, the applicant realizes that the existing monocular ranging method has obvious limitations: on the one hand, simply relying on image features is easily affected by environmental light, weather conditions and target appearance changes, and has insufficient stability; on the other hand, the traditional method is extremely sensitive to the accuracy of the camera calibration parameter, and the drift of the calibration parameter caused by factors such as vehicle vibration and temperature change in actual application will introduce significant ranging error. In addition, the existing method generally lacks effective use of road topological structure, and it is difficult to maintain accuracy in complex road conditions such as curves and slopes. SUMMARY

[0004] Therefore, the present application provides a target ranging method, device and equipment fusing monocular camera and high-precision map, mainly aiming to solve the problems that: 1. the stability is easily affected by environmental light, weather and target appearance; 2. the accuracy of the calibration parameter is sensitive and prone to drift error due to vehicle vibration and temperature change; 3. the road topological structure is not effectively utilized, and the accuracy is difficult to guarantee in complex road conditions.

[0005] According to the first aspect of the present application, a target ranging method fusing monocular camera and high-precision map is provided, which comprises: obtaining vehicle current position information of a target vehicle in a world coordinate system, and a current frame image collected by a monocular camera on the target vehicle; performing 2D target detection processing on the current frame image, identifying at least one detection target, and determining a corresponding two-dimensional detection frame of each detection target in an image coordinate system; Based on the high-precision map data to which the vehicle's current location information belongs, global path planning and discretization are performed on the target vehicle and each of the detected targets respectively, and the road boundary information of the high-precision map data is combined to generate a set of road segments corresponding to each of the detected targets in the world coordinate system; The road segment set corresponding to each detection target is transformed to the image coordinate system, and the depth value of each detection target in the camera coordinate system is calculated in combination with the ground point of each detection target, wherein the ground point is the bottom midpoint of the two-dimensional detection box corresponding to the detection target; The three-dimensional ranging information of each detected target in the world coordinate system is determined by using the two-dimensional detection box corresponding to each detected target in the image coordinate system and the depth value of each detected target in the camera coordinate system.

[0006] According to a second aspect of this application, a target ranging device integrating a monocular camera and a high-precision map is provided, the device comprising: The acquisition module is used to acquire the current position information of the target vehicle in the world coordinate system, as well as the current frame image captured by the vehicle's onboard monocular camera. The target detection module is used to perform 2D target detection processing on the current frame image, identify at least one target, and determine the two-dimensional detection box corresponding to each target in the image coordinate system. The map data processing module is used to perform global path planning and discretization processing on the target vehicle and each of the detected targets based on the high-precision map data to which the current location information of the vehicle belongs, and to generate a set of road segments corresponding to each of the detected targets in the world coordinate system by combining the road boundary information of the high-precision map data. The calculation module is used to transform the set of road segments corresponding to each detected target to the image coordinate system, and calculate the depth value of each detected target in the camera coordinate system in combination with the ground point of each detected target, wherein the ground point is the bottom midpoint of the two-dimensional detection box corresponding to the detected target; The generation module is used to determine the three-dimensional ranging information of each detected target in the world coordinate system using the two-dimensional detection box corresponding to each detected target in the image coordinate system and the depth value of each detected target in the camera coordinate system.

[0007] According to a third aspect of this application, an apparatus is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the method described in any of the first aspects above.

[0008] By employing the above technical solutions, the technical solutions provided in the embodiments of this application have at least the following advantages: The application provides a target ranging method, device and equipment fusing a monocular camera and a high-precision map. The application obtains vehicle current position information of a target vehicle in a world coordinate system and a current frame image collected by a vehicle-mounted monocular camera of the target vehicle. Only a monocular camera is relied on, compared with a laser radar and a binocular camera scheme, the hardware cost is greatly reduced, and the application is more easily popularized in a civilian vehicle and a low-cost automatic driving scene. Then, 2D target detection processing is performed on the current frame image, at least one detection target is identified, and a corresponding two-dimensional detection frame of each detection target in an image coordinate system is determined. Subsequently, global path planning and discretization processing are respectively performed on the target vehicle and each detection target according to high-precision map data to which the vehicle current position information belongs, and a corresponding road segment set of each detection target in the world coordinate system is generated in combination with road boundary information of the high-precision map data. Then, the corresponding road segment set of each detection target is converted to the image coordinate system, and a depth value of each detection target in a camera coordinate system is calculated in combination with a grounding point of each detection target. The grounding point is a bottom midpoint of the two-dimensional detection frame corresponding to the detection target. Since the monocular itself is easy to be affected by light and texture in depth perception, the three-dimensional environment information of the high-precision map is used to compensate for the defects of insufficient depth estimation accuracy of the monocular camera, and the spatial constraint of the map provides an accurate reference for the coordinates, greatly improving the accuracy of three-dimensional positioning and ranging, and significantly improving the reliability of the system in a complex environment. Finally, three-dimensional ranging information of each detection target in the world coordinate system is determined by using the corresponding two-dimensional detection frame of each detection target in the image coordinate system and the depth value of each detection target in the camera coordinate system. By combining the high-precision map and using the road topological structure, the stability problem of the existing monocular ranging method that is easy to be affected by environmental light, weather and target appearance is effectively overcome, the ranging error caused by the drift of the calibration parameters due to the vibration and temperature change of the vehicle is reduced, and high ranging accuracy can be maintained in complex road conditions such as curves and slopes, greatly improving the stability and accuracy of monocular ranging.

[0009] The above description is only a summary of the technical scheme of the application. In order to more clearly understand the technical means of the application, the application can be implemented according to the content of the specification, and in order to make the above and other purposes, characteristics and advantages of the application more obvious and easy to understand, the following specific embodiments of the application are described. BRIEF DESCRIPTION OF DRAWINGS

[0010] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments. The accompanying drawings are included to provide a better understanding of the preferred embodiments, and are not considered limiting of the application. Moreover, like reference numerals denote same or similar components throughout the attached drawings. In the drawings: Figure 1A method flowchart for target ranging by fusing a monocular camera and a high-precision map is shown. Figure 2 Another method flowchart for target ranging by fusing a monocular camera and a high-precision map is shown. Figure 3 A road boundary line segment diagram is shown. Figure 4 A technical flowchart for target ranging by fusing monocular camera perception and a high-precision map is shown. Figure 5 A structure diagram for target ranging by fusing a monocular camera and a high-precision map is shown. Figure 6 A device structure diagram of an apparatus is shown. DETAILED DESCRIPTION

[0011] In the description of the present application, it should be understood that the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the indicated device or element must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application.

[0012] In addition, the terms "first", "second", "third", etc. are only used for description purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second", etc. can explicitly or implicitly include one or more of the features. In the description of the present application, the meaning of "multiple" is two or more, unless otherwise specifically limited.

[0013] In the present application, unless otherwise specifically defined and limited, the terms "mounting", "connecting", "connecting", "fixing", and the like should be understood broadly, for example, it can be fixedly connected, or detachably connected, or integrally connected; it can be mechanically connected, or electrically connected; it can be directly connected, or indirectly connected through an intermediate medium, or it can be the communication between two elements inside. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0014] Exemplary embodiments of the present application will be described herein below with reference to the accompanying drawings. While exemplary embodiments of the present application are shown in the drawings, it is understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that this application will be thoroughly and completely understood, and one skilled in the art will be able to convey the full scope of the present application to others skilled in the art.

[0015] 3D object detection, as one of the core technologies in the field of autonomous driving, its core task is to accurately identify and precisely locate various objects in three-dimensional space from various sensor data, such as vehicles, pedestrians and other obstacles, and output the key information such as the category, specific location, size, orientation, etc. of these objects. Compared with traditional 2D object detection technology, 3D object detection not only needs to identify the planar position of the object, but also must recover the depth information and spatial pose of the object, which puts higher requirements on the accuracy and robustness of the perception system. The application scenarios of 3D object detection technology are extremely wide, covering environmental perception of autonomous vehicles, robot navigation and obstacle avoidance, intelligent transportation systems and many other fields. At present, the main 3D object detection technology solutions mainly include the following: first, through the laser radar device to obtain the 3D point cloud data of the environment, using advanced point cloud processing algorithms (such as PointPillars, CenterPoint, etc.) to directly detect the position, size and orientation of the target; second, through binocular cameras for stereo matching, calculating the disparity map to obtain depth information, and combining 2D detection algorithms (such as YOLO3D) to realize 3D target detection; third, from monocular camera images, first obtain the two-dimensional information of the target through 2D detection algorithms (such as YOLO, etc.), and then estimate the depth and three-dimensional position of the target through geometric constraint method.

[0016] Invention patent CN120107927A proposes a heterogeneous collaborative perception 3D object detection method and device based on laser radar, which innovatively designs a two-stage heterogeneous collaborative perception training process based on laser radar sensors. The fusion network is updated independently without repeating the training of each backbone network, ensuring that the perception performance is not affected. However, this method still has some shortcomings in practical application: first, the computational complexity is high, the point cloud data generated by the laser radar is large, and the steps of multi-scale fusion, KL divergence calculation and adversarial training greatly increase the computational burden, therefore, expensive high-performance computing units need to be configured to meet the computing power requirements; second, this invention only uses laser radar sensors, and laser radar point cloud data mainly provides geometric information (such as distance, shape), but lacks rich semantic information such as color and texture. For small and irregularly shaped objects, laser radar point cloud is often sparse and lacks material and color information, which seriously affects the accuracy and reliability of detection.

[0017] Another invention patent CN118485726A proposes a garbage ranging and size calculation method and system based on monocular camera perception. This method uses a monocular camera to perceive garbage in front of the vehicle, trains a deep learning model based on YOLO, and combines the principle of small aperture imaging to calculate the actual distance and size of the camera to the garbage. However, this method also has some limitations in practical application: first, this method is very dependent on the accuracy of camera calibration. If the calibration board is placed improperly or the corner point detection is not accurate, it may cause cumulative system error and affect the final result. Second, the garbage ranging and size calculation method proposed by this patent is based on a monocular camera, and its core principle relies on the pinhole imaging model and the similar triangle calculation. However, it implicitly assumes that the ground is level, i.e., the relative height between the camera and the ground is fixed, and the ground is not inclined or undulating. This assumption often cannot be fully established in the actual road environment, which seriously affects the accuracy of target ranging and size calculation.

[0018] In the field of automatic driving, 3D target detection based on laser radar point cloud can accurately obtain 3D information of the target. However, due to the lack of semantic information such as color and texture of the target in point cloud data, the performance of laser radar in detecting some unconventional targets is not good. Traditional monocular camera 3D target detection methods face the fundamental problem of inaccurate depth estimation and cannot accurately obtain the real 3D information of the target in actual application relying on the ground level assumption.

[0019] To effectively solve the above problems, the application provides a target ranging method fusing a monocular camera and a high-precision map. The method first uses a traditional 2D detection technology to detect a target in an image to obtain the coordinates and size information of the target in an image coordinate system. Then, high-precision map information is introduced as a depth constraint, the self-vehicle driving trajectory obtained by real-time planning is discretized according to a fixed resolution, and the road boundary information in the high-precision map is fused to calculate a set of left boundary and right boundary point sets. Each set of left boundary and right boundary points can form a line segment, these line segments are arranged along the real-time trajectory line, the distance between the line segments is the resolution, and each line segment is not completely in the same horizontal plane, but reflects the actual ups and downs and inclination of the road, so that the 3D information of the target can be more accurately predicted. Then, the generated line segment set is projected onto the image plane, the 2D detected target is assigned accurate depth information through a geometric matching algorithm, and finally the accurate 3D coordinate calculation is realized, which significantly improves the accuracy and reliability of target detection. The execution subject of the application can be a target ranging system, which provides services for users relying on the computing power of a server. The server can be a standalone server, or a server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms, etc. basic cloud computing servers.

[0020] The embodiment of the application provides a target ranging method fusing a monocular camera and a high-precision map, as shown in the formula (I), the method comprises the following steps: Figure 1 101, obtaining the current position information of the target vehicle in the world coordinate system, and the current frame image collected by the vehicle-mounted monocular camera of the target vehicle.

[0021] In the embodiment of the application, the real-time position information of the target vehicle in the world coordinate system is accurately obtained through an advanced positioning system (such as a vehicle-mounted inertial navigation system, a global positioning system GPS, etc.), and at the same time, a high-quality image frame at the current time is efficiently collected by using a vehicle-mounted monocular camera. These data provide indispensable raw data support for subsequent key steps such as target detection, coordinate conversion and map fusion. The world coordinate system as a global unified space reference system ensures that the vehicle position information can be accurately matched with the high-precision map in the same coordinate system, thereby realizing highly accurate spatial correlation and providing a crucial global spatial reference for subsequent position compensation and three-dimensional ranging operations.

[0022] ​The image frame collected by the monocular camera not only contains rich visual semantic information (such as color, shape, texture, etc.) of the target, but also provides accurate spatial positioning reference for vehicle position information. The organic combination of the two realizes the complementation of visual semantics and spatial positioning information, lays a solid and reliable foundation for subsequent multi-link data processing. In addition, the monocular camera has a significant advantage in hardware cost, and the deployment method is flexible and diverse, while the vehicle positioning technology (such as the combination of GPS and inertial navigation) is already very mature and the cost is controllable. This data acquisition method not only ensures high performance, but also effectively reduces the overall implementation cost of the scheme, greatly promotes its large-scale application in civilian scenarios (such as auxiliary driving systems), and has a wide market prospect and social value.

[0023] 102. performing 2D target detection processing on the current frame image, identifying at least one detection target, and determining a corresponding two-dimensional detection frame of each detection target in the image coordinate system.

[0024] In the embodiment of the present application, the current frame image collected by the vehicle-mounted monocular camera is subjected to two-dimensional target detection operation, aiming to identify at least one detection target in the image, which may include but is not limited to vehicles, pedestrians, obstacles, etc. In the identification process, it is necessary to accurately determine the corresponding two-dimensional detection frame of each detection target in the image coordinate system, i.e. the rectangular area of the target in the image, which contains detailed information such as coordinate range and size. By using mature two-dimensional target detection algorithms such as YOLO series, Single Shot MultiBox Detector (SSD) and the like, the category of the target (such as vehicle, pedestrian, etc.) can be accurately identified, and its specific position in the image can be accurately located, thereby providing necessary semantic information and two-dimensional spatial basic data for the subsequent three-dimensional ranging step.

[0025] In order to meet the high requirements of automatic driving and other practical application scenarios on perception speed, the two-dimensional target detection algorithm has undergone a series of optimization processes, such as lightweight model design, hardware acceleration technology, etc., so that it can realize real-time running on the vehicle-mounted end. In addition, the coordinate and size information provided by the two-dimensional detection frame is the core input data for the subsequent coordinate conversion, position compensation and three-dimensional ranging whole process. If this step is missing, the spatial correlation analysis and three-dimensional information calculation of the target will not be able to proceed smoothly, thereby affecting the performance and accuracy of the whole system. Moreover, as a mature technology in the field of computer vision, the two-dimensional target detection has strong algorithm robustness, which can better adapt to various complex scenes such as light change and small target detection. At the same time, thanks to the rich resources of the open source ecology, this technology is easy to be engineered and further optimized, providing a solid technical support for the application in the field of automatic driving and the like.

[0026] 103. According to the high-precision map data to which the current position information of the vehicle belongs, global path planning and discretization processing are respectively performed on the target vehicle and each detected target, and road segment sets corresponding to each detected target in the world coordinate system are generated in combination with road boundary information of the high-precision map data.

[0027] In the embodiment of the present application, based on the high-precision map data corresponding to the current position of the vehicle (including centimeter-level road geometry information, road boundary information, etc.), global path planning is performed on the target vehicle (for example, an automatic driving vehicle, a sanitation operation vehicle) and each detected target (for example, garbage to be identified, an obstacle). This can be achieved by using a path planning algorithm such as A* algorithm to plan an optimal path, and discretizing the planned continuous path, that is, decomposing the continuous path into a series of discrete points according to a fixed resolution. At the same time, in combination with the road boundary information (for example, the spatial position of the lane line and the curb) of the high-precision map, road segment sets corresponding to each detected target are generated in the world coordinate system (a globally unified coordinate system). The road segment set is a set of line segments connected by discrete points of the road boundary.

[0028] The centimeter-level road information provided by the high-precision map lays a precise spatial foundation for path planning and generation of the segment set, solves the positioning and planning error problems caused by insufficient precision of ordinary maps, and is particularly suitable for scenarios with high requirements for spatial precision, such as 3D target perception of automatic driving and intelligent sanitation. Moreover, by discretizing the continuous path, the complex continuous geometry problem is decomposed into the calculation of discrete points, which can greatly reduce the engineering implementation difficulty of subsequent segment generation and point-line distance calculation, and improve the real-time performance of the algorithm.

[0029] In addition, the road segment set in the world coordinate system provides geometric constraints in the real world for the 2D detected target of the monocular camera. For example, the distance between the detected target and the segment set can be used to calculate the depth, which makes up for the defect of insufficient depth perception of monocular vision, and realizes the mapping of 3D information from the image to the real world. The integration of road boundary information enables the segment set to accurately reflect the geometric shape of the real road, such as the change of the curve and the lane width. Even in the scene of slope and uneven road surface, the segment set can still provide a stable spatial reference framework for the detected target, and improve the reliability of the algorithm in complex scenes.

[0030] 104. The road segment set corresponding to each detected target is converted to the image coordinate system, and the depth value of each detected target in the camera coordinate system is calculated in combination with the contact point of each detected target.

[0031] ​In this embodiment, the set of road segments in the world coordinate system generated based on high-precision maps and path planning is transformed into a set of line segments in the image coordinate system through precise projection transformation of camera intrinsic parameters (including key parameters such as focal length and principal point) and extrinsic parameters (including rotation matrix and translation vector), thereby establishing a close relationship between the real road geometry and the camera imaging plane. The grounding point of the detected target is defined as the bottom midpoint of its two-dimensional detection box (for example, in garbage detection or vehicle detection, this point is the center position of the bottom of the detection box). This grounding point simulates the real contact position between the target and the ground, becoming a key anchor point connecting image features and real-world geometry. In the image coordinate system, by accurately calculating the shortest distance from the grounding point of the detected target to the set of road segments, the closest line segment is found. This line segment corresponds to a discrete path point in the world coordinate system, and its depth information (calculated through the extrinsic parameter matrix) is the depth value of the detected target in the camera coordinate system. Each calculation is based on mature geometric projection principles or detection box feature analysis, featuring good real-time performance and easy engineering implementation, and can be adapted to application scenarios with extremely high requirements for real-time target perception, such as intelligent sanitation and autonomous driving.

[0032] Monocular cameras inherently lack depth information. This application successfully achieves depth estimation in monocular scenes by introducing road geometric constraints from high-precision maps, overcoming the limitations of monocular vision in depth perception and effectively compensating for the shortcoming of monocular vision's inability to directly acquire depth. This provides a solid foundation for subsequent tasks such as 3D target localization and size estimation. The road segment set is derived from a high-precision map (with centimeter-level accuracy), and its geometric information is extremely accurate. The grounding point serves as a stable geometric feature of the detection box (not easily affected by small changes in target pose). The geometric relationship between the two ensures that depth calculation maintains high accuracy even in complex scenes with varying slopes and uneven road surfaces. Furthermore, the discretized road segment set comprehensively covers the boundary morphology of real roads. Even if the target has a certain degree of occlusion or scale variation in the image, reliable depth information can be obtained through robust calculation of the distance between the grounding point and the road segment.

[0033] 105. Using the two-dimensional detection box corresponding to each detection target in the image coordinate system and the depth value of each detection target in the camera coordinate system, determine the three-dimensional ranging information of each detection target in the world coordinate system.

[0034] In the embodiments of the present application, based on the two-dimensional detection box in the image coordinate system (which contains the bounding box coordinates of the target and can accurately reflect the specific position and size information of the target on the imaging plane), combined with the depth value in the camera coordinate system, through the inverse projection transformation technology (which integrates the internal and external parameter information of the camera), the two-dimensional image features and depth information of the target are effectively converted into three-dimensional information in the world coordinate system. These three-dimensional information not only includes the spatial coordinates of the target, but also covers its actual physical dimensions such as width and height, that is, the acquisition of three-dimensional ranging information is realized. This process can realize the perception upgrade from two-dimensional to three-dimensional, because the traditional monocular camera can only output two-dimensional image information, and through the fusion of depth value and inverse projection transformation, the limitations of monocular vision in three-dimensional space perception are successfully broken through, so that the system can accurately obtain the three-dimensional position and size of the target in the real world, providing a reliable spatial basis for the subsequent decision-making process (such as garbage recycling path planning, obstacle avoidance, etc.).

[0035] The calculation of the depth value depends on the geometric constraints of the road segment set and the contact point, and through accurate calculation, its accuracy is ensured. The inverse projection transformation depends on the calibrated camera internal parameters (including focal length and principal point) and external parameters (including rotation matrix and translation vector), and the combination of the two ensures that the three-dimensional information obtained in the world coordinate system has high precision. Even in the scene where there is slope or uneven road surface, the reliability of the information can be maintained through accurate depth association. In addition, the world coordinate system as a global unified reference system, the generated three-dimensional ranging information is not limited by the local environment (such as camera pose, road surface form), so it can adapt to the global decision-making needs of various scenes such as intelligent environmental protection, automatic driving, etc., such as three-dimensional modeling of garbage distribution across regions, long-distance space obstacle avoidance between vehicles, etc.

[0036] Moreover, the technology of the present application not only can obtain the absolute distance between the target and the system (such as environmental protection vehicle, automatic driving vehicle), but also can accurately obtain the specific spatial coordinates and actual physical dimensions of the target in the world, which greatly helps the fine processing of the upper business, such as the development of classification and recycling strategies for garbage of different sizes, thereby improving the intelligent level and practicality of the whole system.

[0037] The embodiment of the present application provides a target ranging method fusing a monocular camera and a high-precision map. Compared with the prior art, the embodiment of the present application acquires vehicle current position information of a target vehicle in a world coordinate system and a current frame image collected by a vehicle-mounted monocular camera of the target vehicle, and only relies on the monocular camera, so that the hardware cost is greatly reduced compared with a laser radar and a binocular camera scheme, and the monocular camera is more easily popularized in a civilian vehicle and a low-cost automatic driving scene. Then, 2D target detection processing is performed on the current frame image, at least one detection target is identified, and a corresponding two-dimensional detection frame of each detection target in an image coordinate system is determined. Subsequently, global path planning and discretization processing are respectively performed on the target vehicle and each detection target according to high-precision map data to which the vehicle current position information belongs, and a corresponding road segment set of each detection target in the world coordinate system is generated in combination with road boundary information of the high-precision map data. Then, the corresponding road segment set of each detection target is converted to the image coordinate system, and a depth value of each detection target in a camera coordinate system is calculated in combination with a grounding point of each detection target. The grounding point is a bottom midpoint of the two-dimensional detection frame corresponding to the detection target. Since the monocular itself is easy to be affected by light and texture in depth perception, therefore, the three-dimensional environment information of the high-precision map is used to make up for the defects of insufficient depth estimation accuracy of the monocular camera, and the spatial constraint of the map provides accurate reference for the coordinates, so that the accuracy of three-dimensional positioning and ranging is greatly improved, and the reliability of the system in a complex environment is significantly improved. Finally, three-dimensional ranging information of each detection target in the world coordinate system is determined by using the corresponding two-dimensional detection frame of each detection target in the image coordinate system and the depth value of each detection target in the camera coordinate system. By combining the high-precision map and using the road topological structure, the stability problem of the existing monocular ranging method that is easy to be affected by environmental light, weather and target appearance is effectively overcome, meanwhile, the ranging error caused by the drift of calibration parameters due to vehicle vibration and temperature change is reduced, and high ranging accuracy can also be maintained in complex road conditions such as a curve and a slope, so that the stability and accuracy of monocular ranging are greatly improved.

[0038] Further, as a refinement and expansion of the above embodiment, in order to completely describe the specific implementation process of the embodiment, the embodiment of the present application provides another target ranging method fusing a monocular camera and a high-precision map, as shown in Figure 2 The method comprises the following steps. 201. Acquire vehicle current position information of a target vehicle in a world coordinate system and a current frame image collected by a vehicle-mounted monocular camera of the target vehicle.

[0039] In this embodiment, the unmanned sweeper successfully implemented a garbage recognition function. This function can accurately identify obstacles in front of the vehicle and precisely acquire the real coordinates, size, and category information of the target object and the vehicle. To achieve this goal, the unmanned sweeper needs to be equipped with a monocular camera with a field of view covering at least 40 meters of road surface in front, ensuring clear images without blind spots, and fixed on a dedicated bracket. Simultaneously, the camera's internal and external parameters need to be calibrated. Furthermore, an onboard inertial navigation system needs to be installed to achieve real-time positioning of the autonomous vehicle. Additionally, a GPU-equipped controller is required, responsible for image capture and data acquisition, preprocessing the acquired data and performing model inference, and combining the 2D detection results with a high-precision map to obtain 3D information through coordinate transformation. This allows the acquisition of the target vehicle's current position information in the world coordinate system, as well as the current frame image captured by the target vehicle's onboard monocular camera.

[0040] It is worth noting that, to significantly enhance perception capabilities in specific environments (such as at night or in adverse weather conditions), alternative camera types, such as infrared cameras, can be considered to replace traditional monocular cameras. Infrared cameras maintain good imaging performance even in low light or complex weather conditions, effectively enhancing the system's environmental perception capabilities. Furthermore, the performance of the perception system can be further optimized by introducing multiple sensors as auxiliary means, such as radar or ultrasonic sensors. Radar sensors can penetrate obstacles such as rain and fog, providing long-range detection capabilities; while ultrasonic sensors excel at accurate short-range ranging. Through the collaborative work of these sensors, not only can the accuracy of perception be significantly improved, but the robustness of the system can also be significantly enhanced, ensuring stable perception performance in various complex environments.

[0041] 202. Perform 2D target detection processing on the current frame image, identify at least one target, and determine the corresponding two-dimensional detection box for each target in the image coordinate system.

[0042] In this embodiment, a series of preprocessing operations are performed on the current frame image to obtain a standardized image. The current frame image is a standard RGB format image, and the preprocessing steps include, but are not limited to, image scaling and normalization operations to ensure that the image size and pixel values ​​meet the input requirements of subsequent deep learning models. Through these preprocessing steps, the quality of the image data can be effectively improved, making it more suitable for object detection tasks.

[0043] The advanced deep learning detector YOLOv5 is obtained, and the pre-processed standardized image is input into the YOLOv5 deep learning detector for target detection. The detector outputs multiple initial detection boxes and detailed information for each initial detection box, including bounding box coordinates, confidence, and class prediction vector. Among them, the confidence represents the predicted probability of the existence of the target in the initial detection box, and the class prediction vector contains the prediction probability of multiple target classes. YOLOv5 is selected as the basic detector mainly because of its efficiency and accuracy. After the input monocular camera collected RGB image is processed by the network, the original output is a multi-dimensional tensor. Through a series of post-processing steps, the final output information of each detection target includes: bounding box coordinates , which specifically represents the rectangular bounding box of the target in the image coordinate system, where represents the left lower corner coordinate, represents the right upper corner coordinate; confidence , which represents the probability that the prediction box contains a target object, is normalized by the sigmoid function. The confidence score integrates target prediction and positioning accuracy, and is an important indicator for measuring detection quality.

[0044] A confidence threshold is set to filter out initial detection boxes with confidence greater than or equal to the threshold from multiple initial detection boxes, and the initial detection boxes are used as target detection boxes, thereby obtaining multiple target detection boxes. This step helps to filter out detection results with low confidence and improves the accuracy of detection.

[0045] For each target detection box, among the multiple target class prediction probabilities contained in its class prediction vector, the largest target class prediction probability is selected, and the class label corresponding to the largest target class prediction probability is taken as the class label of the target detection box. Each class label clearly indicates the specific class of a detection target.

[0046] In multiple target detection boxes, target detection boxes with the same class label are grouped into a group, thereby obtaining at least one detection box group. According to the class label corresponding to each detection box group, the detection target corresponding to each detection box group is determined, and at least one detection target is finally determined.

[0047] The non-maximum suppression (NMS) method is used to remove repeated boxes for each detection target corresponding to the detection box group, to obtain a two-dimensional detection box corresponding to each detection target. In the non-maximum suppression process, the confidence score is used to filter out high-quality detection results to ensure that the final output detection box has high credibility. The center point of the detection box The accuracy of the center point coordinates plays a key role in subsequent geometric matching with map projection line segments, and directly affects the accuracy of depth assignment. Therefore, special attention should be paid to the positioning accuracy of the center point during the detection frame regression process to ensure the performance and reliability of the overall detection system.

[0048] It should be noted that advanced image segmentation models can also be used to replace traditional 2D object detection methods. Image segmentation models can more accurately identify and divide the boundaries of target objects by finely analyzing pixel information in the image, thereby achieving accurate segmentation of each region in the image. Compared to 2D object detection, image segmentation models exhibit higher accuracy and robustness in handling complex scenes and small targets, effectively improving overall detection effectiveness and system performance. Therefore, selecting image segmentation models as an alternative solution is to further enhance the comprehensive ability in the field of target recognition and image processing.

[0049] 203、For each detection target, based on high-precision map data, use algorithm to perform global path planning for the target vehicle and the detection target, obtaining a vehicle driving trajectory of the target vehicle in a world coordinate system.

[0050] In the embodiments of the present application, first, the high-precision map data to which the current position information of the vehicle belongs is extracted from the high-precision map database. Next, for each detection target, based on the extracted high-precision map data, use algorithm to perform global path planning for the target vehicle and the detection target. The algorithm is a heuristic search algorithm that can find an optimal path from the starting point to the ending point based on the current position information of the vehicle and the high-precision map data. This path not only needs to satisfy the feasibility of vehicle driving, but also needs to satisfy certain optimization objectives, such as shortest path length, shortest driving time, etc. Through the path planning of the algorithm, the vehicle driving trajectory of the target vehicle in the world coordinate system can be obtained, which will include various path points that the vehicle needs to pass through during driving, as well as driving direction and speed information at each path point. In this way, it can be ensured that the vehicle always advances along the optimal path during driving, thereby improving driving efficiency and reducing driving time.

[0051] 204、Discretize the vehicle driving trajectory according to a preset resolution to obtain a set of discretized path points.

[0052] In the embodiment of the present application, a preset resolution value is obtained, and the preset resolution directly determines the accuracy of the generated target three-dimensional information in subsequent prediction. Specifically, the smaller the preset resolution value, the more detailed the detail capture, and thus the higher the accuracy of the predicted target 3D information. Based on the preset resolution, the trajectory formed by the vehicle during driving is discretized, that is, the continuous driving trajectory is segmented and sampled according to the preset resolution, and finally a point set composed of multiple discretized path points is obtained. The discretized path point set provides basic data support for subsequent path planning and analysis, and the calculation formula is as follows Formula 1: Formula 1:

[0053] wherein, represents the i th discretized path point in the discretized path point set, , represents the total number of discretized path points, represents the starting point of the vehicle driving trajectory, represents the preset resolution, represents the trajectory direction vector of the i th discretized path point in the vehicle driving trajectory.

[0054] 205, according to the vehicle driving trajectory, the discretized path point set and the road boundary information, a road segment set of the detection target in the world coordinate system is generated.

[0055] In the embodiment of the present application, the road boundary information is extracted from the high-precision map data according to the vehicle driving trajectory, and the left boundary point set and the right boundary point set are calculated according to the vehicle driving trajectory, the discretized path point set and the road boundary information, and the calculation formula is as follows Formula 2: Formula 2:

[0056] wherein, represents the i th left boundary point in the left boundary point set, represents the i th right boundary point in the right boundary point set, represents the i th discretized path point in the discretized path point set, represents a unit vector perpendicular to the trajectory direction of the vehicle driving trajectory and pointing to the left side of the road at the i th discretized path point, represents a unit vector perpendicular to the trajectory direction of the vehicle driving trajectory and pointing to the right side of the road at the i th discretized path point, represents the road width in the road boundary information, and the left side of the road and the right side of the road are determined according to the road boundary information.

[0057] AsFigure 3 The illustrated road boundary line segment schematic diagram, for each specific discretization path point in the discretization path point set, first accurately extracts the left boundary point corresponding to the discretization path point in the left boundary point set, and then accurately extracts the right boundary point corresponding to the discretization path point in the right boundary point set. Subsequently, the left boundary point and the right boundary point corresponding to the extracted discretization path point are used to generate a complete road segment corresponding to the discretization path point. On this basis, by analyzing each discretization path point corresponding to the road segment, the complete road segment set in the world coordinate system is gradually constructed and generated It should be noted that these line segments are not in the same plane, but are distributed according to the actual road conditions. This fine processing process ensures accurate fusion and efficient use of road boundary information, laying a solid foundation for subsequent target ranging work.

[0058] 206、For each detection target, the left boundary point coordinates and the right boundary point coordinates of each road segment in the road segment set corresponding to the detection target are converted to the camera coordinate system by using the camera extrinsic coefficient, to obtain the left boundary point coordinates and the right boundary point coordinates of each road segment in the camera coordinate system.

[0059] In the embodiment of the present application, for each target to be detected, the system first performs coordinate system conversion operation by using the external parameter matrix of the camera, that is, the camera extrinsic coefficient. Specifically, this process involves extracting the coordinate values of the left boundary point and the coordinate values of the right boundary point of each road segment in the road segment set corresponding to the detection target. Subsequently, by applying the conversion formula of the camera extrinsic coefficient, these boundary point coordinates originally in the road coordinate system are mapped and converted one by one into the coordinate system of the camera itself. After such conversion processing, the new coordinate data of each road segment in the camera coordinate system can be successfully obtained, including the accurate coordinate position of the left boundary point of each road segment in the camera coordinate system and the accurate coordinate position of the right boundary point in the camera coordinate system. This step is crucial for subsequent image processing and target recognition, ensuring data consistency and accuracy.

[0060] 207、The left boundary point coordinates and the right boundary point coordinates of each road segment in the camera coordinate system are converted to the image coordinate system by using the camera intrinsic coefficient, the radial distortion coefficient and the tangential distortion coefficient, to obtain the left boundary point coordinates and the right boundary point coordinates of each road segment in the image coordinate system.

[0061] In the embodiments of the present application, for each road segment, the left boundary point coordinate and the right boundary point coordinate of the road segment in the camera coordinate system are converted into the image coordinate system by using the camera intrinsic coefficient, to obtain the initial left boundary point coordinate and the initial right boundary point coordinate of the road segment in the image coordinate system, and the calculation formula is as follows: Formula 3:

[0062] wherein, represents the initial left boundary point coordinate or the initial right boundary point coordinate of the road segment in the image coordinate system, represents the left boundary point coordinate or the right boundary point coordinate of the road segment in the camera coordinate system, represents the principal point pixel coordinate in the camera intrinsic coefficient, represents the pixel focal length of the vehicle-mounted monocular camera in the image X axis direction in the camera intrinsic coefficient, represents the pixel focal length of the vehicle-mounted monocular camera in the image Y axis direction in the camera intrinsic coefficient.

[0063] Then, the initial left boundary point coordinate and the initial right boundary point coordinate of the road segment in the image coordinate system are corrected by using the radial distortion coefficient and the tangential distortion coefficient, to obtain the left boundary point coordinate and the right boundary point coordinate of the road segment in the image coordinate system, and the calculation formula is as follows: Formula 4:

[0064] wherein, represents the left boundary point coordinate or the right boundary point coordinate of the road segment in the image coordinate system, represents the initial left boundary point coordinate or the initial right boundary point coordinate of the road segment in the image coordinate system, , , represents the radial distortion coefficient, , represents the tangential distortion coefficient, represents the principal point pixel coordinate in the camera intrinsic coefficient, represents the distance value from the initial left boundary point coordinate or the initial right boundary point coordinate of the road segment in the image coordinate system to the principal point pixel coordinate, represents the normalized pixel coordinate of the left boundary point or the right boundary point of the road segment in the image coordinate system.

[0065] It needs to be specially pointed out that the process of calibrating parameters is to use the chessboard calibration board to accurately calibrate the parameters of the monocular camera. The specific operation steps are as follows: first, place the chessboard calibration board at different positions and angles for multi-angle shooting, and ensure that the corner points of the chessboard in each photo are clear. Then, using an efficient corner detection algorithm, automatically identify and extract the corner point coordinates of the chessboard in each photo. By comparing and calculating these detected corner point coordinates with the actual corner point positions, the camera's intrinsic parameters and distortion coefficients are finally obtained, ensuring the accuracy and reliability of the camera imaging.

[0066] In addition, in order to calibrate the external parameters (external parameters) of the camera and the rear axle of the vehicle, the vehicle needs to be fixed on the car and driven to a professional calibration room for operation. During the calibration process, the chessboard calibration board is placed in front of the vehicle, ensuring that the calibration board is completely within the camera's field of view and that the ground is flat and smooth to ensure the accuracy of the measurement data. Next, measure the three-dimensional position relationship between the calibration board and the rear axle of the vehicle, including the longitudinal distance d, lateral offset l, and height difference h of the center of the calibration board to the rear axle, and other key parameters. By detecting the corner points of the calibration board with the camera, the accurate pose of the calibration board relative to the camera is calculated. According to the measurement relationship between the calibration board and the rear axle of the vehicle, the coordinate transformation relationship from the calibration board coordinate system (B) to the vehicle rear axle coordinate system (V) is established. Finally, through matrix operation, the external parameters (external parameters) between the camera (C) and the vehicle rear axle (V) are accurately obtained, providing a solid foundation for subsequent image processing and data analysis.

[0067] After completing the camera parameter calibration, the YOLOv5 target detector is used to perform efficient model inference on the RGB images collected by the monocular camera in real time. This detector can quickly identify and output two-dimensional detection boxes of different types of garbage in the image and their corresponding class information, thereby achieving accurate classification and positioning of garbage and providing strong technical support for subsequent garbage disposal and environmental protection work.

[0068] 208、Using the two-dimensional detection box corresponding to the detection target to calculate the grounding point coordinates of the detection target, and according to the grounding point coordinates of the detection target and the left boundary point coordinates and right boundary point coordinates of each road segment in the image coordinate system, determining the road segment in the road segment set that is closest to the grounding point of the detection target in the image coordinate system as the target road segment.

[0069] In the embodiments of the present application, for each road segment, firstly, the accurate position information of the road segment in the image coordinate system needs to be obtained, which includes the point coordinates of the left boundary and the point coordinates of the right boundary of the road segment. On this basis, the specific coordinate position of the grounding point of the detection target in the image coordinate system is further obtained. By using these coordinate data, the actual distance between the grounding point of the detection target and the current road segment is accurately calculated by a corresponding geometric calculation method. Through this step, a corresponding road distance value can be generated for each road segment. Subsequently, all the calculated road distance values are comprehensively compared, and the road distance value with the smallest value is selected from them, and the road segment corresponding to the smallest road distance value is determined as the final target road segment. This process ensures that the selection of the target road segment is based on the shortest distance principle, thereby improving the accuracy and reliability of the detection.

[0070] The distance between the two-dimensional detection target and the projection line segment is usually calculated by the distance between the grounding point of the target and the projection line segment. The grounding point of the target usually refers to the bottom center point of the target object in contact with the ground (assuming that the target object is located on the ground). For example, for pedestrians or vehicles, their bottom center points can be regarded as their grounding points. The calculation formula of the target frame grounding point is as follows: Formula 5:

[0071] wherein, represents the right boundary horizontal coordinate of the two-dimensional detection frame in the image coordinate system, represents the lower boundary vertical coordinate of the two-dimensional detection frame in the image coordinate system, represents the height of the two-dimensional detection frame. Then, the distance between the target and all discrete line segments is calculated by using the point-to-line distance formula, and the line segment closest to the target is found, and the calculation formula is as follows: Formula 6:

[0072] wherein, represents the distance value, represents the grounding point coordinates, represents the left boundary point coordinates, represents the right boundary point coordinates. Therefore, the depth Z of the straight line is the depth of the detection target in the camera coordinate system wherein, represents the distance between the i-th discrete line segment and the grounding point of the detection target, represents the number of the effective line segment closest to the target, represents the distance threshold value, which is used to screen the effective line segment, only the line segment with a distance less than the threshold value is regarded as a candidate, and the line segment that is too far or irrelevant is excluded, represents a minimum operation.

[0073] 209、determine the discrete path point corresponding to the target road segment in the discrete path point set as the target discrete path point, convert the coordinate of the target discrete path point from the world coordinate system to the camera coordinate system, and take the coordinate value of the target discrete path point in the Z-axis direction of the camera coordinate system as the depth value of the detection target in the camera coordinate system.

[0074] In the embodiment of the present application, the discrete path point set contains a plurality of discrete path points, each of which represents a specific position on the path. Next, in this discrete path point set, the discrete path point corresponding to the target road segment is accurately identified and determined. Then, the specific coordinate information of the target discrete path point is extracted from the discrete path point set. These coordinate information is initially based on the world coordinate system, in order to facilitate the camera to capture and analyze, the coordinate of the target discrete path point is converted from the world coordinate system to the camera coordinate system, and the conversion process ensures the unity of the coordinate system. After completing the conversion of the coordinate system, the coordinate value of the target discrete path point in the Z-axis direction of the camera coordinate system represents the depth information of the detection target relative to the camera lens. By obtaining this depth value, the specific position and distance of the detection target in the camera view can be further analyzed and understood, thereby providing important data support for subsequent image processing and target recognition.

[0075] 210、determine the three-dimensional ranging information of each detection target in the world coordinate system by using the two-dimensional detection frame corresponding to each detection target in the image coordinate system and the depth value of each detection target in the camera coordinate system.

[0076] In the embodiment of the present application, for each target object to be detected, the system first accurately uses the contact point coordinate information of the detection target in the image plane coordinate system according to the camera internal parameter matrix including the focal length, the principal point coordinate and other key internal parameter coefficients, and combines the depth distance value of the detection target in the camera coordinate system, to carefully calculate the three-dimensional space projection coordinates of the detection target in the camera coordinate system through a series of complex inverse projection transformation algorithms. This process not only ensures the accuracy of coordinate conversion, but also provides a solid foundation for subsequent three-dimensional reconstruction and target positioning. The calculation formula is as follows: Formula 7:

[0077] wherein, represents the three-dimensional projection coordinates of the detection target in the camera coordinate system, represents the contact point coordinates of the detection target in the image coordinate system, represents the depth value of the detection target in the camera coordinate system, represents the camera principal point pixel coordinates in the camera internal parameter coefficients, This represents the pixel focal length of the vehicle-mounted monocular camera in the X-axis direction of the image, representing the camera's intrinsic parameters. This represents the pixel focal length of the vehicle-mounted monocular camera in the Y-axis direction of the image, which is part of the camera's intrinsic parameters.

[0078] Next, the camera's extrinsic parameter matrix, i.e., the camera's extrinsic coefficients, is used to transform the three-dimensional projected coordinates of the detected target in the camera coordinate system. Specifically, this step transforms the three-dimensional projected coordinates of the detected target, originally located in the camera coordinate system, into the more widely used world coordinate system through precise mathematical transformation. This transformation process obtains the three-dimensional projected coordinates of the detected target in the world coordinate system, thus providing accurate spatial location information for subsequent positioning, navigation, or other related applications. The calculation formula is shown in Formula 8 below: Formula 8:

[0079] in, This represents the three-dimensional projected coordinates of the detected target in the world coordinate system. This represents the three-dimensional projected coordinates of the detected target in the camera coordinate system. Represents the camera's extrinsic coefficients Rotation matrix, Represents the camera's extrinsic coefficients Translation vector.

[0080] Subsequently, the bounding box coordinates of the two-dimensional detection box corresponding to the detected target in the image coordinate system are obtained. Using the camera intrinsic parameters, the bounding box coordinates of the two-dimensional detection box corresponding to the detected target in the image coordinate system, and the depth value of the detected target in the camera coordinate system, the width and height values ​​of the detected target in the world coordinate system are calculated, which can truly reflect the actual size of the detected target in the real world. The calculation formula is as follows: Formula 9: Formula 9:

[0081] in, This represents the width of the detected target in the world coordinate system. This represents the height value of the detected target in the world coordinate system. Represents the bounding box coordinates. Indicates the depth value. This represents the pixel focal length of the vehicle-mounted monocular camera in the X-axis direction of the image, representing the camera's intrinsic parameters. This represents the pixel focal length of the vehicle-mounted monocular camera in the Y-axis direction of the image, which is part of the camera's intrinsic parameters.

[0082] Finally, the system will take the width and height values of the detected target as the key parameters of the three-dimensional size information of the detected target in the world coordinate system. At the same time, the three-dimensional projection coordinates of the detected target in the world coordinate system are obtained, which can reflect the specific position of the target in space. In addition, combined with the three-dimensional size information of the detected target in the world coordinate system obtained before, the complete three-dimensional ranging information of the detected target in the world coordinate system is formed, which not only accurately describes the size of the target, but also accurately reflects its position relationship in space, providing comprehensive and accurate data support for subsequent analysis and processing.

[0083] From the above process, a technical flowchart of the target ranging technology of fusing monocular camera perception and high-precision map proposed by the embodiments of the present application is as follows: As shown in Figure 4 , first, for the input image data, i.e. the real-time collected RGB image by the monocular camera, the current advanced YOLOv5 target detection model is used for two-dimensional target detection processing. This step aims to identify various types of garbage existing in the image and output the corresponding two-dimensional detection frame and its corresponding category information. In order to obtain more accurate road boundary information, relevant data are extracted from the high-precision map, and these data are all based on the world coordinate system.

[0084] Next, the driving trajectory of the vehicle is obtained in real time by using the positioning information of the autonomous vehicle and the global path planning algorithm. It should be noted that although this trajectory is generated in the world coordinate system, it is based on the assumption of horizontal road surface, so it cannot truly reflect the complex road surface conditions. Especially in the face of scenarios with slope or uneven road surface, simply relying on this trajectory to predict the depth of the target will introduce large errors.

[0085] In order to further improve the accuracy of prediction, the global path is discretized, i.e. the planned driving trajectory is discretized at a fixed resolution. The higher the resolution, the more accurate the final predicted target depth. After discretization, these discrete trajectory points are fused with the road boundary information extracted from the high-precision map, so as to calculate the left and right boundary points corresponding to each trajectory point. By connecting these left and right boundary points into line segments and ensuring that these line segments are evenly distributed along the driving direction of the vehicle, a dense line segment set that truly reflects the road conditions can be obtained.

[0086] Subsequently, this line segment set is converted from the world coordinate system to the image coordinate system. On this basis, the distance between the contact point (i.e. the midpoint of the bottom of the detection frame) of each two-dimensional target and the line segment set is calculated, and the path discrete point coordinates corresponding to the line segment closest to the target are found. Combined with the extrinsic information, the depth value of the target can be calculated .

[0087] Finally, by using the matched depth estimation value, the two-dimensional target is reconstructed in three-dimensional coordinates through inverse projection transformation technology, and the three-dimensional result containing the target position and size information is finally output. This process not only improves the dimension and accuracy of target detection, but also provides more reliable data support for the navigation and decision-making of autonomous vehicles in complex environments.

[0088] The embodiment of the application provides a target ranging method fusing a monocular camera and a high-precision map. Compared with the prior art, the embodiment of the application only relies on a monocular camera by acquiring current position information of a target vehicle in a world coordinate system and a current frame image collected by a vehicle-mounted monocular camera of the target vehicle, so that the hardware cost is greatly reduced compared with a laser radar and a binocular camera scheme, and the method is more easily popularized in civilian vehicles and low-cost autonomous driving scenes. Then, 2D target detection processing is performed on the current frame image, at least one detected target is identified, and a corresponding two-dimensional detection frame of each detected target in an image coordinate system is determined. Subsequently, global path planning and discretization processing are respectively performed on the target vehicle and each detected target according to high-precision map data to which the current position information of the vehicle belongs, and a corresponding road segment set of each detected target in the world coordinate system is generated in combination with road boundary information of the high-precision map data. Then, the road segment set corresponding to each detected target is converted to the image coordinate system, and the depth value of each detected target in the camera coordinate system is calculated in combination with the grounding point of each detected target. The grounding point is the bottom midpoint of the two-dimensional detection frame corresponding to the detected target. Since the depth perception of the monocular itself is easily affected by light and texture, the three-dimensional environmental information of the high-precision map is used to compensate for the defects of insufficient depth estimation accuracy of the monocular camera, and the spatial constraint of the map provides accurate reference for the coordinates, thereby greatly improving the accuracy of three-dimensional positioning and ranging, and significantly improving the reliability of the system in complex environments. Finally, the three-dimensional ranging information of each detected target in the world coordinate system is determined by using the two-dimensional detection frame corresponding to each detected target in the image coordinate system and the depth value of each detected target in the camera coordinate system. By combining the road topology of the high-precision map, the stability problem of the existing monocular ranging method that is easily affected by environmental light, weather and target appearance is effectively overcome, the ranging error caused by the drift of the calibration parameters due to the vibration and temperature change of the vehicle is reduced, and high ranging accuracy can be maintained in complex road conditions such as curves and slopes, thereby greatly improving the stability and accuracy of monocular ranging.

[0089] Further, as Figure 1 a specific implementation of the method, the embodiment of the application provides a target ranging device fusing a monocular camera and a high-precision map, as shown in Figure 5 The device includes an acquisition module 301, a target detection module 302, a map data processing module 303, a calculation module 304 and a generation module 305.

[0090] The acquisition module 301 is configured to acquire vehicle current position information of a target vehicle in a world coordinate system and a current frame image collected by a monocular camera on the target vehicle; The target detection module 302 is configured to perform 2D target detection processing on the current frame image, identify at least one detection target, and determine a corresponding two-dimensional detection box of each detection target in an image coordinate system; The map data processing module 303 is configured to perform global path planning and discretization processing on the target vehicle and each detection target respectively according to high-precision map data to which the vehicle current position information belongs, and generate a corresponding road segment set of each detection target in the world coordinate system in combination with road boundary information of the high-precision map data; The calculation module 304 is configured to convert the road segment set corresponding to each detection target to the image coordinate system, and calculate a depth value of each detection target in a camera coordinate system in combination with a grounding point of each detection target, the grounding point being a bottom midpoint of the two-dimensional detection box corresponding to the detection target; The generation module 305 is configured to determine three-dimensional ranging information of each detection target in the world coordinate system by using the two-dimensional detection box corresponding to each detection target in the image coordinate system and the depth value of each detection target in the camera coordinate system.

[0091] In a specific application scenario, the acquisition module 301 is configured to perform preprocessing on the current frame image to obtain a standardized image, the current frame image is an RGB image, and the preprocessing includes scaling operation and normalization operation; a YOLOv5 deep learning detector is acquired, the standardized image is input into the YOLOv5 deep learning detector for target detection, a plurality of initial detection boxes, a bounding box coordinate, a confidence and a class prediction vector of each initial detection box are obtained, the confidence represents a target existence prediction probability of the initial detection box, and the class prediction vector includes a plurality of target class prediction probabilities; a confidence threshold is acquired, and initial detection boxes with a confidence greater than or equal to the confidence threshold are selected from the plurality of initial detection boxes as target detection boxes, and a plurality of target detection boxes are obtained; for each target detection box, a target class prediction probability with a maximum value is selected from a plurality of target class prediction probabilities included in the class prediction vector of the target detection box, a class label corresponding to the target class prediction probability with the maximum value is taken as a class label of the target detection box, each class label indicates a detection target; target detection boxes with the same class label are grouped into a group from the plurality of target detection boxes, at least one detection box group is obtained, a detection target corresponding to each detection box group is determined according to a class label corresponding to each detection box group, and at least one detection target is determined; a non-maximum suppression method is used to perform a repeated box removal operation on a detection box group corresponding to each detection target, and a two-dimensional detection box corresponding to each detection target is obtained.

[0092] In a specific application scenario, the map data processing module 303 is configured to acquire a high-precision map database, extract high-precision map data to which the current position information of the vehicle belongs in the high-precision map database; for each detection target, based on the high-precision map data, a global path planning algorithm is used to perform global path planning on the target vehicle and the detection target, and a vehicle driving trajectory of the target vehicle in a world coordinate system is obtained; a preset resolution is acquired, and the vehicle driving trajectory is discretized according to the preset resolution to obtain a discretized path point set, the discretized path point set includes a plurality of discretized path points,

[0093] wherein, represents an i th discretized path point in the discretized path point set, , represents a total number of discretized path points, represents a starting point of the vehicle driving trajectory, represents the preset resolution, ​a trajectory direction vector representing an i-th discretized path point in the vehicle driving trajectory; extracting road boundary information according to the vehicle driving trajectory in the high-precision map data, calculating a left boundary point set and a right boundary point set according to the vehicle driving trajectory, the discretized path point set and the road boundary information,

[0094] wherein, an i-th left boundary point in the left boundary point set, an i-th right boundary point in the right boundary point set, an i-th discretized path point in the discretized path point set, a unit vector perpendicular to a trajectory direction of the vehicle driving trajectory at the i-th discretized path point and pointing to a left side of a road, a unit vector perpendicular to a trajectory direction of the vehicle driving trajectory at the i-th discretized path point and pointing to a right side of the road, a road width in the road boundary information, the road left side and the road right side being determined according to the road boundary information; sequentially determining a left boundary point corresponding to each of the discretized path points in the left boundary point set and a right boundary point corresponding to each of the discretized path points in the right boundary point set, and generating a road segment corresponding to each of the discretized path points; and generating a road segment set of the detection target in a world coordinate system by sequentially using the road segments corresponding to the plurality of discretized path points.

[0095] In a specific application scenario, the calculation module 304 is configured to, for each detection target, obtain camera extrinsic parameter coefficients, convert left boundary point coordinates and right boundary point coordinates of each road segment in the road segment set corresponding to the detection target to a camera coordinate system by using the camera extrinsic parameter coefficients, to obtain left boundary point coordinates and right boundary point coordinates of each road segment in the camera coordinate system; obtain camera intrinsic parameter coefficients, radial distortion coefficients, and tangential distortion coefficients, convert the left boundary point coordinates and the right boundary point coordinates of each road segment in the camera coordinate system to an image coordinate system by using the camera intrinsic parameter coefficients, the radial distortion coefficients, and the tangential distortion coefficients, to obtain left boundary point coordinates and right boundary point coordinates of each road segment in the image coordinate system; calculate a grounding point coordinate of the detection target by using a two-dimensional detection frame corresponding to the detection target; determine, from a plurality of road segments in the road segment set, a road segment with the shortest distance to the grounding point of the detection target in the image coordinate system as a target road segment, according to the grounding point coordinate of the detection target and the left boundary point coordinates and the right boundary point coordinates of each road segment in the image coordinate system; obtain a set of discretized path points, determine a discretized path point corresponding to the target road segment in the set of discretized path points as a target discretized path point; and convert the target discretized path point coordinate from a world coordinate system to a camera coordinate system, and take a Z-axis direction coordinate value of the target discretized path point in the camera coordinate system as a depth value of the detection target in the camera coordinate system.

[0096] In a specific application scenario, the calculation module 304 is configured to, for the left boundary point coordinates and the right boundary point coordinates of each road segment in the camera coordinate system, convert the left boundary point coordinates and the right boundary point coordinates of the road segment in the camera coordinate system to an image coordinate system by using the camera intrinsic parameter coefficients, to obtain initial left boundary point coordinates and initial right boundary point coordinates of the road segment in the image coordinate system,

[0097] wherein, the initial left boundary point coordinates or the initial right boundary point coordinates of the road segment in the image coordinate system are represented by x0 and y0, the left boundary point coordinates or the right boundary point coordinates of the road segment in the camera coordinate system are represented by x and y, a camera principal point pixel coordinate in the camera intrinsic parameter coefficients is represented by x0, a pixel focal length of the vehicle-mounted monocular camera in the X-axis direction of an image in the camera intrinsic parameter coefficients is represented by fx, represents a pixel focal length of the vehicle-mounted monocular camera in the image Y-axis direction in the camera intrinsic parameter coefficients; and the initial left boundary point coordinates and the initial right boundary point coordinates of the road segment in the image coordinate system are corrected by using the radial distortion coefficient and the tangential distortion coefficient to obtain left boundary point coordinates and right boundary point coordinates of the road segment in the image coordinate system,

[0098] wherein, represents the left boundary point coordinates or the right boundary point coordinates of the road segment in the image coordinate system, represents the initial left boundary point coordinates or the initial right boundary point coordinates of the road segment in the image coordinate system, represents the radial distortion coefficient, represents the tangential distortion coefficient, represents a camera principal point pixel coordinate in the camera intrinsic parameter coefficients, represents a distance value from the initial left boundary point coordinates or the initial right boundary point coordinates of the road segment in the image coordinate system to the camera principal point pixel coordinate, represents the normalized pixel coordinates of the left boundary point or the right boundary point of the road segment in the image coordinate system.

[0099] In a specific application scenario, the calculation module 304 is configured to, for each road segment, calculate a distance between the touchdown point of the detection target and the road segment based on the left boundary point coordinates and the right boundary point coordinates of the road segment in the image coordinate system and the touchdown point coordinates of the detection target, to obtain a road distance value corresponding to the road segment; and take a road segment corresponding to a road distance value with a minimum value in the plurality of road distance values as the target road segment.

[0100] In a specific application scenario, the generation module 305 is configured to, for each detection target, perform inverse projection transformation on the touchdown point coordinates of the detection target in the image coordinate system and the depth value of the detection target in the camera coordinate system by using camera intrinsic parameter coefficients, to obtain three-dimensional projection coordinates of the detection target in the camera coordinate system,

[0101] wherein, represents the three-dimensional projection coordinates of the detection target in the camera coordinate system, represents the touchdown point coordinates of the detection target in the image coordinate system, represents the depth value of the detection target in the camera coordinate system, ​​​represents a camera principal point pixel coordinate in the camera intrinsic parameter coefficient, represents a pixel focal length of the vehicle-mounted monocular camera in the image X-axis direction in the camera intrinsic parameter coefficient, represents a pixel focal length of the vehicle-mounted monocular camera in the image Y-axis direction in the camera intrinsic parameter coefficient; the three-dimensional projection coordinates of the detection target in the camera coordinate system are converted from the camera coordinate system to the world coordinate system by using the camera extrinsic parameter coefficient, to obtain the three-dimensional projection coordinates of the detection target in the world coordinate system,

[0102] wherein, represents the three-dimensional projection coordinates of the detection target in the world coordinate system, represents the three-dimensional projection coordinates of the detection target in the camera coordinate system, represents a rotation matrix of the camera extrinsic parameter coefficient, represents a translation vector of the camera extrinsic parameter coefficient; the three-dimensional size information of the detection target in the world coordinate system is calculated according to the corresponding two-dimensional detection frame of the detection target in the image coordinate system and the depth value of the detection target in the camera coordinate system; the three-dimensional projection coordinates of the detection target in the world coordinate system and the three-dimensional size information of the detection target in the world coordinate system are taken as the three-dimensional ranging information of the detection target in the world coordinate system.

[0103] In a specific application scenario, the generation module 305 is configured to obtain the bounding box coordinates of the corresponding two-dimensional detection frame of the detection target in the image coordinate system; the camera intrinsic parameter coefficient, the bounding box coordinates of the corresponding two-dimensional detection frame of the detection target in the image coordinate system, and the depth value of the detection target in the camera coordinate system are used for calculation, to obtain the width value and the height value of the detection target in the world coordinate system,

[0104] wherein, represents the width value of the detection target in the world coordinate system, represents the height value of the detection target in the world coordinate system, represents the bounding box coordinates, represents the depth value, represents a pixel focal length of the vehicle-mounted monocular camera in the image X-axis direction in the camera intrinsic parameter coefficient, represents a pixel focal length of the vehicle-mounted monocular camera in the image Y-axis direction in the camera intrinsic parameter coefficient; the width value and the height value of the detection target are taken as the three-dimensional size information of the detection target in the world coordinate system.

[0105] Compared with the prior art, the embodiment of the application acquires vehicle current position information of a target vehicle in a world coordinate system and a current frame image collected by a vehicle-mounted monocular camera of the target vehicle, and only relies on the monocular camera, so that the hardware cost is greatly reduced compared with a laser radar and a binocular camera scheme, and the monocular camera is more easily popularized in a civilian vehicle and a low-cost automatic driving scene. Then, 2D target detection processing is performed on the current frame image, at least one detection target is identified, and a corresponding two-dimensional detection frame of each detection target in an image coordinate system is determined. Subsequently, global path planning and discretization processing are respectively performed on the target vehicle and each detection target according to high-precision map data to which the vehicle current position information belongs, and a corresponding road segment set of each detection target in the world coordinate system is generated in combination with road boundary information of the high-precision map data. Then, the corresponding road segment set of each detection target is converted to the image coordinate system, and a depth value of each detection target in a camera coordinate system is calculated in combination with a grounding point of each detection target, the grounding point being a bottom midpoint of the two-dimensional detection frame corresponding to the detection target. Since monocular depth perception is easily affected by light and texture, three-dimensional environmental information of the high-precision map is used to compensate for the defects of insufficient depth estimation accuracy of the monocular camera, and the spatial constraint of the map provides an accurate reference for the coordinates, thereby greatly improving the accuracy of three-dimensional positioning and ranging and significantly improving the reliability of the system in a complex environment. Finally, three-dimensional ranging information of each detection target in the world coordinate system is determined by using the corresponding two-dimensional detection frame of each detection target in the image coordinate system and the depth value of each detection target in the camera coordinate system. By combining the road topology structure of the high-precision map, the stability problem of the existing monocular ranging method that is easily affected by environmental light, weather and target appearance is effectively overcome, ranging errors caused by calibration parameter drift due to vehicle vibration and temperature change are reduced, and high ranging accuracy can be maintained in complex road conditions such as curves and slopes, thereby greatly improving the stability and accuracy of monocular ranging.

[0106] It should be noted that other corresponding descriptions of the functions of the target ranging device provided by the embodiment of the application and involved in the fusion of the monocular camera and the high-precision map can be referred to Figure 1 and Figures 2 to 4 , and will not be described here.

[0107] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the application are all information and data authorized by the user or authorized by all parties.

[0108] Any technical features in the above embodiments can be combined, and for the sake of brevity, not all possible combinations are described, but it is understood that any combination of the technical features is within the scope of the present disclosure.

[0109] The above embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed, but it should not be understood as a limitation on the patent scope of the present application. It should be noted that, for those skilled in the art, some modifications and improvements can be made without departing from the concept of the present application, and these are within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

[0110] In the example embodiment, referring to Figure 6 An apparatus is also provided, which comprises a bus, a processor, a memory, and a communication interface, and can further comprise an input / output interface and a display apparatus, wherein the communication between the various functional units can be completed through the bus. The memory stores a computer program, and the processor is configured to execute the program stored in the memory to implement the target ranging method of fusing a monocular camera and a high-precision map in the above embodiments.

[0111] A computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the target ranging method of fusing a monocular camera and a high-precision map.

[0112] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by hardware, or by means of software and a necessary general hardware platform. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.), and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various implementation scenarios of the present application.

[0113] Those skilled in the art can understand that the drawings are only a schematic diagram of a preferred implementation scenario, and the modules or processes in the drawings are not necessarily required for implementing the present application.

[0114] Those skilled in the art can understand that the modules in the apparatus in the implementation scenario can be distributed in the apparatus in the implementation scenario according to the description of the implementation scenario, or can be changed and located in one or more apparatuses different from the implementation scenario. The modules in the above implementation scenario can be combined into one module, or can be further split into a plurality of sub-modules.

[0115] The above application serial number is only for description, and does not represent the advantages and disadvantages of the implementation scene.

[0116] The above disclosure is only a few specific implementation scenarios of the application, but the application is not limited thereto, and any changes that can be thought of by those skilled in the art should fall within the protection scope of the application.

Claims

1. A target ranging method integrating a monocular camera and a high-precision map, characterized in that, include: Obtain the current position information of the target vehicle in the world coordinate system, as well as the current frame image captured by the vehicle's onboard monocular camera; Perform 2D target detection processing on the current frame image to identify at least one target and determine the corresponding two-dimensional detection box for each target in the image coordinate system; Based on the high-precision map data to which the vehicle's current location information belongs, global path planning and discretization are performed on the target vehicle and each of the detected targets respectively, and the road boundary information of the high-precision map data is combined to generate a set of road segments corresponding to each of the detected targets in the world coordinate system; The road segment set corresponding to each detection target is transformed to the image coordinate system, and the depth value of each detection target in the camera coordinate system is calculated in combination with the ground point of each detection target, wherein the ground point is the bottom midpoint of the two-dimensional detection box corresponding to the detection target; The three-dimensional ranging information of each detected target in the world coordinate system is determined by using the two-dimensional detection box corresponding to each detected target in the image coordinate system and the depth value of each detected target in the camera coordinate system.

2. The method according to claim 1, characterized in that, The step of performing 2D target detection processing on the current frame image to identify at least one target and determine the corresponding two-dimensional detection box for each target in the image coordinate system includes: The current frame image is preprocessed to obtain a standardized image. The current frame image is an RGB image. The preprocessing includes scaling and normalization operations. A YOLOv5 deep learning detector is obtained. The standardized image is input into the YOLOv5 deep learning detector for target detection, resulting in multiple initial detection boxes, and the bounding box coordinates, confidence score, and class prediction vector of each initial detection box. The confidence score represents the predicted probability of the target in the initial detection box, and the class prediction vector includes multiple target class prediction probabilities. Obtain a confidence threshold, and select initial detection boxes with a confidence level greater than or equal to the confidence threshold from the plurality of initial detection boxes as target detection boxes to obtain a plurality of target detection boxes; For each target detection box, the target category prediction probability with the largest value is selected from the multiple target category prediction probabilities included in the category prediction vector of the target detection box, and the category label corresponding to the target category prediction probability with the largest value is used as the category label of the target detection box, and each category label indicates a detected target; Target detection boxes with the same category label are selected from the plurality of target detection boxes and grouped into a group to obtain at least one detection box group. The detection target corresponding to each detection box group is determined according to the category label corresponding to each detection box group, and the at least one detection target is determined. The non-maximum suppression method is used to perform duplicate box removal operation on the detection boxes corresponding to each detection target group to obtain the two-dimensional detection boxes corresponding to each detection target.

3. The method according to claim 1, characterized in that, The step involves performing global path planning and discretization processing on the target vehicle and each detected target based on the high-precision map data to which the vehicle's current location information belongs, and generating a set of road segments corresponding to each detected target in the world coordinate system by combining the road boundary information of the high-precision map data. Obtain a high-precision map database, and extract the high-precision map data to which the vehicle's current location information belongs from the high-precision map database; For each of the detected targets, based on the high-precision map data, the following is adopted: The algorithm performs global path planning on the target vehicle and the detected target to obtain the vehicle's trajectory in the world coordinate system; A preset resolution is obtained, and the vehicle's trajectory is discretized according to the preset resolution to obtain a discretized path point set, which includes multiple discretized path points. in, This represents the i-th discretized path point in the set of discretized path points. , This represents the total number of discretized path points. This indicates the starting point of the vehicle's driving trajectory. This refers to the preset resolution. This represents the trajectory direction vector of the i-th discretized path point in the vehicle's driving trajectory; Based on the vehicle's trajectory, road boundary information is extracted from the high-precision map data. Then, based on the vehicle's trajectory, the discretized path point set, and the road boundary information, the left and right boundary point sets are calculated. in, This represents the i-th left boundary point in the set of left boundary points. This represents the i-th right boundary point in the set of right boundary points. This represents the i-th discretized path point in the set of discretized path points. Let represent the unit vector at the i-th discretized path point that is perpendicular to the vehicle's trajectory and points to the left side of the road. Let represent the unit vector at the i-th discretized path point that is perpendicular to the vehicle's trajectory and points to the right side of the road. The road width is indicated in the road boundary information, and the left and right sides of the road are determined based on the road boundary information; The left boundary point corresponding to each discretized path point in the left boundary point set and the right boundary point corresponding to each discretized path point are determined sequentially to generate the road segment corresponding to each discretized path point. The road segment set of the detected target in the world coordinate system is generated sequentially using the road segments corresponding to the multiple discretized path points.

4. The method according to claim 1, characterized in that, The step of transforming the set of road segments corresponding to each detected target to the image coordinate system and calculating the depth value of each detected target in the camera coordinate system based on the ground point of each detected target includes: For each of the detected targets, the camera extrinsic coefficients are obtained. The coordinates of the left and right boundary points of each road segment in the set of road segments corresponding to the detected target are transformed to the camera coordinate system using the camera extrinsic coefficients, so as to obtain the coordinates of the left and right boundary points of each road segment in the camera coordinate system. The camera intrinsic coefficients, radial distortion coefficients, and tangential distortion coefficients are obtained. The coordinates of the left and right boundary points of each road segment in the camera coordinate system are transformed to the image coordinate system using the camera intrinsic coefficients, the radial distortion coefficients, and the tangential distortion coefficients, so as to obtain the coordinates of the left and right boundary points of each road segment in the image coordinate system. The coordinates of the grounding point of the target are calculated using the two-dimensional detection frame corresponding to the target. Based on the grounding point coordinates of the detected target and the left and right boundary point coordinates of each road segment in the image coordinate system, the road segment with the shortest distance to the grounding point of the detected target in the image coordinate system is determined from the multiple road segments in the road segment set, and is taken as the target road segment. Obtain a set of discretized path points, and determine the discretized path points corresponding to the target road segment in the set of discretized path points as target discretized path points. The coordinates of the target discretized path point are obtained from the discretized path point set. The coordinates of the target discretized path point are transformed from the world coordinate system to the camera coordinate system. The Z-axis coordinate value of the target discretized path point in the camera coordinate system is used as the depth value of the detected target in the camera coordinate system.

5. The method according to claim 4, characterized in that, The process of transforming the coordinates of the left and right boundary points of each road segment in the camera coordinate system to the image coordinate system using the camera intrinsic parameters, the radial distortion coefficient, and the tangential distortion coefficient, to obtain the coordinates of the left and right boundary points of each road segment in the image coordinate system, includes: For each road segment, the coordinates of its left and right boundary points in the camera coordinate system are transformed to the image coordinate system using the camera intrinsic parameters, thus obtaining the initial left and right boundary point coordinates of the road segment in the image coordinate system. in, This represents the initial left boundary point coordinates or the initial right boundary point coordinates of the road segment in the image coordinate system. This indicates the coordinates of the left or right boundary point of the road segment in the camera coordinate system. This represents the camera principal point pixel coordinates in the camera intrinsic parameter coefficients. This represents the pixel focal length of the vehicle-mounted monocular camera in the X-axis direction of the image, which is one of the camera intrinsic parameter coefficients. This represents the pixel focal length of the vehicle-mounted monocular camera in the Y-axis direction of the image, which is one of the camera intrinsic parameter coefficients. The initial left and right boundary point coordinates of the road segment in the image coordinate system are corrected using the radial and tangential distortion coefficients to obtain the left and right boundary point coordinates of the road segment in the image coordinate system. in, This indicates the coordinates of the left or right boundary point of the road segment in the image coordinate system. This represents the initial left boundary point coordinates or the initial right boundary point coordinates of the road segment in the image coordinate system. , , This represents the radial distortion coefficient. , This represents the tangential distortion coefficient. This represents the camera principal point pixel coordinates in the camera intrinsic parameter coefficients. This represents the distance from the initial left or right boundary point coordinates of the road segment in the image coordinate system to the pixel coordinates of the camera's principal point. This represents the normalized pixel coordinates of the left or right boundary point of the road segment in the image coordinate system.

6. The method according to claim 4, characterized in that, The step of determining the road segment with the shortest distance in the image coordinate system to the grounding point of the detected target from multiple road segments in the road segment set, based on the grounding point coordinates of the detected target and the left and right boundary point coordinates of each road segment in the image coordinate system, as the target road segment, includes: For each road segment, based on the coordinates of the left and right boundary points of the road segment in the image coordinate system, and the coordinates of the grounding point of the detected target, the distance between the grounding point of the detected target and the road segment is calculated to obtain the road distance value corresponding to the road segment. The road segment corresponding to the smallest road distance value among the multiple road distance values ​​is taken as the target road segment.

7. The method according to claim 1, characterized in that, The step of determining the three-dimensional ranging information of each detected target in the world coordinate system using the two-dimensional detection box corresponding to each detected target in the image coordinate system and the depth value of each detected target in the camera coordinate system includes: For each detected target, based on the camera intrinsic parameters, an inverse projection transformation is performed using the ground point coordinates of the detected target in the image coordinate system and the depth value of the detected target in the camera coordinate system to obtain the three-dimensional projected coordinates of the detected target in the camera coordinate system. in, This represents the three-dimensional projected coordinates of the detected target in the camera coordinate system. This represents the coordinates of the grounding point of the detected target in the image coordinate system. This represents the depth value of the detected target in the camera coordinate system. This represents the camera principal point pixel coordinates in the camera intrinsic parameter coefficients. This represents the pixel focal length of the vehicle-mounted monocular camera in the X-axis direction of the image, which is one of the camera intrinsic parameter coefficients. This represents the pixel focal length of the vehicle-mounted monocular camera in the Y-axis direction of the image, which is one of the camera intrinsic parameter coefficients. The 3D projected coordinates of the detected target in the camera coordinate system are transformed to the world coordinate system using camera extrinsic coefficients, thus obtaining the 3D projected coordinates of the detected target in the world coordinate system. in, This represents the three-dimensional projected coordinates of the detected target in the world coordinate system. This represents the three-dimensional projected coordinates of the detected target in the camera coordinate system. Represents the camera extrinsic coefficients Rotation matrix, Represents the camera extrinsic coefficients Translation vector; The three-dimensional size information of the detected target in the world coordinate system is calculated based on the two-dimensional detection box corresponding to the detected target in the image coordinate system and the depth value of the detected target in the camera coordinate system. The three-dimensional projection coordinates of the detected target in the world coordinate system and the three-dimensional size information of the detected target in the world coordinate system are used as the three-dimensional ranging information of the detected target in the world coordinate system.

8. The method according to claim 7, characterized in that, The step of calculating the three-dimensional size information of the detected target in the world coordinate system based on the two-dimensional detection box corresponding to the detected target in the image coordinate system and the depth value of the detected target in the camera coordinate system includes: Obtain the bounding box coordinates of the two-dimensional detection box corresponding to the detected target in the image coordinate system; The width and height of the detected target in the world coordinate system are calculated using the camera intrinsic parameters, the bounding box coordinates of the two-dimensional detection box corresponding to the detected target in the image coordinate system, and the depth value of the detected target in the camera coordinate system. in, This represents the width of the detected target in the world coordinate system. This represents the height value of the detected target in the world coordinate system. Indicates the coordinates of the bounding box. This represents the depth value. This represents the pixel focal length of the vehicle-mounted monocular camera in the X-axis direction of the image, which is one of the camera intrinsic parameter coefficients. This represents the pixel focal length of the vehicle-mounted monocular camera in the Y-axis direction of the image, which is one of the camera intrinsic parameter coefficients. The width and height values ​​of the detected target are used as the three-dimensional size information of the detected target in the world coordinate system.

9. A target ranging device integrating a monocular camera and a high-precision map, applied to the target ranging method integrating monocular camera perception and high-precision map as described in claim 1, characterized in that... include: The acquisition module is used to acquire the current position information of the target vehicle in the world coordinate system, as well as the current frame image captured by the vehicle's onboard monocular camera. The target detection module is used to perform 2D target detection processing on the current frame image, identify at least one target, and determine the two-dimensional detection box corresponding to each target in the image coordinate system. The map data processing module is used to perform global path planning and discretization processing on the target vehicle and each of the detected targets based on the high-precision map data to which the current location information of the vehicle belongs, and to generate a set of road segments corresponding to each of the detected targets in the world coordinate system by combining the road boundary information of the high-precision map data. The calculation module is used to transform the set of road segments corresponding to each detected target to the image coordinate system, and calculate the depth value of each detected target in the camera coordinate system in combination with the ground point of each detected target, wherein the ground point is the bottom midpoint of the two-dimensional detection box corresponding to the detected target; The generation module is used to determine the three-dimensional ranging information of each detected target in the world coordinate system using the two-dimensional detection box corresponding to each detected target in the image coordinate system and the depth value of each detected target in the camera coordinate system.

10. An apparatus comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Garbage distance measurement and size calculation method and system based on monocular camera perception

    CN118485726A

  • Heterogeneous cooperative sensing 3D target detection method and device based on laser radar

    CN120107927A