External parameter adjusting method and device of vehicle camera and vehicle machine equipment

By constructing a coordinate index structure in vehicle data, pixel coordinates are directly determined from road environment targets and projection errors are calculated, solving the problem of low reliability of camera extrinsic parameter adjustment. Stable extrinsic parameter adjustment is achieved in scenarios without specific calibration, improving the perception accuracy and safety of autonomous driving systems.

CN120876618APending Publication Date: 2025-10-31WHITE RHINO ZHIDA (BEIJING) TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511059108.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In the absence of specific calibration scenarios in existing technologies, the reliability of camera extrinsic parameter adjustment calibration is low, which threatens the perception accuracy and safety of autonomous driving systems.

Method used

By directly determining the target pixel coordinates of road environment targets from vehicle data, a coordinate index structure is constructed. This index structure is then used to query the two-dimensional coordinates of the candidate camera extrinsic projections and calculate the projection error, thereby filtering out accurate target camera extrinsic parameters and avoiding reliance on specific calibration scenarios.

Benefits of technology

It achieves stable and reliable adjustment of camera extrinsic parameters without the need for specific calibration scenarios, thereby improving the perception accuracy and safety of autonomous driving systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876618A_ABST
    Figure CN120876618A_ABST
Patent Text Reader

Abstract

The invention relates to an external parameter adjusting method and device of a vehicle camera and vehicle equipment. The method comprises the following steps: reading vehicle landing data, and extracting three-dimensional target detection data and camera image data of a static or low-speed moving road environment target from the vehicle landing data; obtaining target pixel coordinates of a road environment target in the camera image data; constructing a coordinate index structure according to the target pixel coordinates; determining a plurality of candidate camera external parameters and two-dimensional projection coordinates of the candidate camera external parameters to the sampling points of the plurality of three-dimensional target detection data; querying the coordinate index structure, and determining a coordinate index result of the two-dimensional projection coordinates; and screening target camera external parameters of the vehicle camera from the plurality of candidate camera external parameters according to a projection error obtained by a coordinate index result. Therefore, statistical optimization can be carried out by using mass driving data, so that the camera external parameters with high reliability can be obtained by adjusting and calibrating without depending on a specific calibration scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of autonomous driving technology, and in particular to a method, apparatus and vehicle-mounted equipment for adjusting the extrinsic parameters of a vehicle camera. Background Technology

[0002] With the development of vehicle intelligence, autonomous driving, as one of the important functions of vehicles, has gained increasing popularity among users. Autonomous driving systems rely on the fusion of multiple sensors, such as cameras and LiDAR, to perceive the environment. Accurate fusion of camera and LiDAR data depends on accurate extrinsic parameters (i.e., the rotation matrix and translation vector between the camera coordinate system and the vehicle or LiDAR coordinate system). Deviations in extrinsic parameters can lead to incorrect projection of 3D object detection boxes and distortion of multi-sensor fusion results, seriously threatening the perception accuracy and safety of autonomous driving systems.

[0003] Currently, the mainstream methods for extrinsic parameter calibration include: the calibration board method, which requires the use of a specific pattern calibration board in a dedicated site, and combines multi-angle images with LiDAR point cloud computing to calculate extrinsic parameters; and the scene feature method, which uses natural features such as building edges and road markings to establish the correspondence between images and point clouds to solve for extrinsic parameters.

[0004] However, during implementation, the applicant discovered that the relevant technology suffers from low reliability in adjusting camera extrinsic parameters when there is a lack of specific calibration scenarios. Summary of the Invention

[0005] Based on this, the purpose of this application is to at least solve one of the above-mentioned technical defects, in particular the technical defect of poor reliability of calibration adjustment of camera extrinsic parameters in the prior art due to the lack of specific calibration. This application provides a method, device and vehicle equipment for adjusting the extrinsic parameters of a vehicle camera.

[0006] In a first aspect, this application provides a method for adjusting the extrinsic parameters of a vehicle camera, the method comprising:

[0007] Read vehicle landing data, and extract three-dimensional target detection data and camera image data of stationary or low-speed moving road environment targets from the vehicle landing data; wherein, the low-speed movement is movement below a preset speed threshold;

[0008] Obtain the target pixel coordinates of road environment targets in camera image data; and construct a coordinate index structure based on the target pixel coordinates;

[0009] Determine the extrinsic parameters of multiple candidate cameras, as well as the two-dimensional projection coordinates of multiple sampling points of the 3D target detection data based on the candidate camera extrinsic parameters; the candidate camera extrinsic parameters are obtained by adjusting the current camera extrinsic parameters.

[0010] The coordinate index structure is queried to determine the coordinate index result of the two-dimensional projected coordinates;

[0011] Based on the projection error obtained from the coordinate index results, the target camera extrinsic parameters of the vehicle camera are selected from multiple candidate camera extrinsic parameters.

[0012] In one embodiment, obtaining the target pixel coordinates corresponding to the road environment target extracted from camera image data includes:

[0013] Obtain the target pixel mask after semantic segmentation of camera image data;

[0014] Morphological dilation is performed on the mask region corresponding to the road environment target to obtain the dilated pixel mask;

[0015] Traverse the dilated pixel mask, extract the pixel coordinates of the mask region corresponding to the road environment target, and obtain the target pixel coordinates.

[0016] In one embodiment, morphological dilation is performed on the mask region corresponding to the road environment target to obtain a dilated pixel mask, including:

[0017] Obtain the correspondence between the preset distance threshold and the expansion range;

[0018] Based on the correspondence and the distance between the vehicle and the road environment target, the mask area corresponding to the road environment target is morphologically dilated.

[0019] In one embodiment, the coordinate index structure includes a coordinate hash table;

[0020] Based on the target pixel coordinates, construct a coordinate index structure, including:

[0021] The target pixel coordinates are used as the index data of the coordinate hash table, and the preset query matching identifier is used as the target value data of the coordinate hash table.

[0022] Construct a coordinate hash table based on the index data and target value data.

[0023] In one embodiment, the method further includes:

[0024] If the target value data corresponding to the two-dimensional projected coordinates is found in the coordinate hash table, the coordinate index result indicates that the two-dimensional projected coordinates have been matched.

[0025] If no target value data corresponding to the two-dimensional projected coordinates is found in the coordinate hash table, the coordinate index result indicates that the two-dimensional projected coordinates did not match.

[0026] The projection error is obtained by determining the hit rate based on the number of matched two-dimensional projection coordinates and the number of sampling points.

[0027] In one embodiment, based on the projection error obtained from the coordinate indexing result, the target camera extrinsic parameters of the vehicle camera are selected from multiple candidate camera extrinsic parameters, including:

[0028] Based on the projection errors of candidate camera extrinsic parameters for multiple road environment targets, the cumulative projection error corresponding to the candidate camera extrinsic parameters is obtained;

[0029] The candidate camera extrinsic parameters with the smallest cumulative projection error are used as the target camera extrinsic parameters.

[0030] In one embodiment, the candidate camera extrinsic parameters include a candidate extrinsic parameter combination, which includes multiple rotation axis parameters and multiple translation direction parameters.

[0031] The method also includes:

[0032] Determine the cumulative projection error corresponding to the candidate extrinsic parameter combinations used for coarse search adjustment;

[0033] The candidate extrinsic combination with the smallest cumulative projection error after coarse search adjustment is taken as the target extrinsic combination for coarse search adjustment.

[0034] Determine the cumulative projection error corresponding to the candidate extrinsic parameter combinations used for fine search adjustment; wherein, the candidate extrinsic parameter combinations for fine search adjustment are obtained by adjusting the target extrinsic parameter combinations for coarse search adjustment;

[0035] The candidate extrinsic parameter combination with the smallest cumulative projection error after fine-search adjustment is used as the target extrinsic parameter combination for the vehicle camera.

[0036] In one embodiment, the method further includes:

[0037] Based on a preset coarse search range, the current camera extrinsic parameters are adjusted with a preset coarse search step size to determine multiple candidate extrinsic parameter combinations for coarse search adjustment;

[0038] Based on a preset fine search range and a preset fine search step size, the target extrinsic parameter combination for coarse search adjustment is adjusted to determine multiple candidate extrinsic parameter combinations for fine search adjustment.

[0039] Secondly, this application provides an external parameter adjustment device for a vehicle camera, the device comprising:

[0040] The vehicle data reading module is used to read vehicle landing data and extract three-dimensional target detection data and camera image data of stationary or low-speed moving road environment targets from the vehicle landing data; wherein, the low-speed movement is movement below a preset speed threshold.

[0041] The coordinate index building module is used to obtain the target pixel coordinates corresponding to road environment targets in camera image data; and to build a coordinate index structure based on the target pixel coordinates.

[0042] The projection coordinate determination module is used to determine the extrinsic parameters of multiple candidate cameras, as well as the two-dimensional projection coordinates of multiple sampling points of the 3D target detection data based on the extrinsic parameters of the candidate cameras; the extrinsic parameters of the candidate cameras are obtained by adjusting the extrinsic parameters of the current camera.

[0043] The coordinate index result module is used to query the coordinate index structure and determine the coordinate index result of the two-dimensional projected coordinates;

[0044] The target extrinsic parameter filtering module is used to filter the target camera extrinsic parameters of the vehicle camera from multiple candidate camera extrinsic parameters based on the projection error obtained from the coordinate index results.

[0045] Thirdly, this application provides an in-vehicle infotainment device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method.

[0046] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:

[0047] The vehicle camera extrinsic parameter adjustment method, apparatus, and vehicle-mounted equipment provided in this application directly determine the target pixel coordinates of road environment targets from vehicle data. This allows the construction of a coordinate index structure. This index structure can be used to query the two-dimensional coordinates projected using candidate camera extrinsic parameters. The projection error of the candidate camera extrinsic parameter is calculated based on the query results. Furthermore, the accurate target camera extrinsic parameter can be selected from multiple candidate camera extrinsic parameters based on the projection error. Therefore, compared to traditional technologies, this application, by constructing a coordinate index structure to select the optimal camera extrinsic parameter and establishing a stable index relationship using stationary or low-speed moving targets, can perform statistical optimization using massive amounts of driving data. This allows for stable and highly reliable camera extrinsic parameters to be obtained without relying on specific calibration scenarios. Ultimately, this can improve the perception accuracy and safety of autonomous driving systems. Attached Figure Description

[0048] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0049] Figure 1 A flowchart illustrating a method for adjusting the extrinsic parameters of a vehicle camera, provided in an embodiment of this application;

[0050] Figure 2A flowchart illustrating a step for obtaining target pixel coordinates provided in an embodiment of this application;

[0051] Figure 3 A flowchart illustrating the steps for determining the extrinsic parameters of a target camera, provided as an embodiment of this application;

[0052] Figure 4 This application provides a flowchart illustrating a specific implementation of a method for adjusting the extrinsic parameters of a vehicle camera.

[0053] Figure 5 This is a schematic diagram of the structure of an external parameter adjustment device for a vehicle camera provided in an embodiment of this application;

[0054] Figure 6 This is a schematic diagram of the internal structure of a vehicle-mounted device provided in an embodiment of this application. Detailed Implementation

[0055] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0056] Autonomous driving technology is a crucial development direction for the current automotive industry. Autonomous driving systems rely on various sensors (such as cameras, LiDAR, and millimeter-wave radar) to perceive the surrounding environment. Among these, the fusion of cameras and LiDAR is one of the key technologies for achieving accurate environmental perception.

[0057] Camera extrinsic parameters refer to the transformation relationship between the camera coordinate system and the vehicle coordinate system (or LiDAR coordinate system), including rotation matrices and translation vectors. Accurate extrinsic parameter calibration is a prerequisite for achieving sensor data fusion. Deviations in extrinsic parameters can lead to problems such as: inaccurate projection positions of 3D object detection boxes on the image; deviations in multi-sensor fusion results; and impact on the perception accuracy and safety of autonomous driving systems.

[0058] Currently, the mainstream methods for calibrating camera extrinsic parameters include:

[0059] (1) Calibration board-based method: Using a calibration board with a checkerboard or other specific pattern, the calibration board is photographed from multiple angles within the calibration site. The camera extrinsic parameters are calculated by detecting the corner points of the calibration board and combining them with the lidar point cloud data. This method requires: a dedicated calibration site; a precisely manufactured calibration board; manual operation to park the vehicle in the designated location; and professional technicians to perform the calibration operation.

[0060] (2) Scene feature-based method: This method uses natural features in the environment (such as building edges, road markings, etc.) for calibration. This method calculates extrinsic parameters by extracting corresponding features from images and point clouds and establishing matching relationships. However, it has the following problems: high requirements for the quality of scene features; complex feature extraction and matching algorithms; poor performance in scenes with sparse features; and projection error calculation usually involves complex geometric operations.

[0061] In general, the relevant technologies require dedicated calibration sites, calibration boards and other equipment, or have high computational complexity and low efficiency when processing large-scale data. As a result, the relevant technologies have the problem of low reliability of camera extrinsic parameter adjustment calibration when there is a lack of specific calibration scenarios.

[0062] Based on this, this application provides a method, apparatus, and vehicle-mounted device for adjusting the extrinsic parameters of a vehicle camera. By directly determining the target pixel coordinates of road environment targets from vehicle data, a coordinate index structure can be constructed. This coordinate index structure can be used to query the two-dimensional coordinates projected using candidate camera extrinsic parameters. By querying the index results, the projection error of the candidate camera extrinsic parameters can be calculated. Furthermore, by using the projection error of the candidate camera extrinsic parameters, the accurate target camera extrinsic parameters can be selected from multiple candidate camera extrinsic parameters. Stable and reliable camera extrinsic parameters can be obtained by adjusting the calibration without relying on a specific calibration scene.

[0063] In one exemplary embodiment, Figure 1 This is a flowchart illustrating a method for adjusting the extrinsic parameters of a vehicle camera, as provided in an embodiment of this application. Figure 1 As shown, a method for adjusting the extrinsic parameters of a vehicle camera is provided. This method is illustrated using a terminal as an example. It is understood that this method can also be applied to a server, or to a system including both a terminal and a server, and is implemented through interaction between the terminal and the server. The terminal can be, but is not limited to, an in-vehicle infotainment system or a vehicle-mounted terminal. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. In this embodiment, the method includes the following steps S101 to S105: Wherein:

[0064] S101. Read the vehicle landing data, and extract three-dimensional target detection data and camera image data of stationary or low-speed moving road environment targets from the vehicle landing data; wherein, the low-speed movement is movement below a preset speed threshold.

[0065] The vehicle landing data refers to the spatiotemporally aligned raw dataset synchronously collected and stored on a local storage device (such as an SSD) by onboard sensor systems (cameras, LiDAR, millimeter-wave radar, etc.) during vehicle operation. The 3D target detection data can be target 3D attribute data generated through LiDAR point cloud segmentation or vision-radar fusion algorithms. Camera image data refers to raw image frames captured in real-time by vehicle cameras. Vehicle cameras refer to cameras installed on the vehicle, which can be used in Advanced Driver Assistance Systems (ADAS) or autonomous driving systems to perceive the surrounding environment; for example, there are forward-looking, rear-looking, and surround-view cameras.

[0066] Road environment targets refer to pedestrians, cyclists, or specific objects that exist in the vehicle's driving environment and represent the real world, such as streetlights, traffic signs, lane markings, etc., which are stationary or below the moving target.

[0067] For example, the vehicle-mounted device can read spatiotemporally aligned data packets from the vehicle's solid-state storage to obtain vehicle landing data, and can use vehicle perception algorithms, such as YOLO, SSD and other target detection models, to process point cloud data from LiDAR, millimeter-wave radar and other sources as well as raw image data collected by cameras in real time to identify road environment targets in the images, such as identifying road environment targets in the images through semantic segmentation.

[0068] Optionally, during normal vehicle operation, the data processed and stored in the vehicle's perception system in real time may include: camera image data (including timestamps), LiDAR point cloud data (including timestamps), 3D object detection results (including 3D bounding box information for pedestrians), semantic segmentation results (including pixel-level masks for pedestrians), and vehicle pose information. The vehicle's in-vehicle equipment can directly read this stored data without re-detecting and segmenting, significantly improving processing efficiency. Furthermore, the collected data is time-synchronized to ensure time alignment between the image and point cloud data.

[0069] In practical applications, by using stationary targets and / or low-speed moving targets below a preset speed threshold as road environment targets, real-time calibration and adjustment of camera extrinsic parameters can be performed in dynamic driving scenarios. For example, pedestrians with low moving speeds can be selected as calibration targets to reduce the impact of sensor synchronization errors and improve calibration accuracy.

[0070] S102. Obtain the target pixel coordinates of the road environment target in the camera image data; and construct a coordinate index structure based on the target pixel coordinates.

[0071] The target pixel coordinates refer to the set of pixel positions of the road environment target contour region extracted from the camera image through image processing technology. Specifically, it can be achieved by semantic segmentation combined with morphological dilation processing, which is used to establish the distribution characteristics of the road environment target in the image space.

[0072] The coordinate index structure refers to the data organization form used for fast retrieval of pixel coordinates. Specifically, a hash table structure can be used to map pixel coordinates into key-value pairs, and a query operation with O(1) time complexity can be achieved through a hash function to improve the efficiency of projected coordinate matching.

[0073] For example, the position of a stationary or slow-moving target in the image data can be directly obtained from camera image data captured while the vehicle is in motion; that is, the target pixel coordinates. The vehicle-mounted device can organize all the extracted target pixel coordinates into a spatial index data structure, such as a hash table, KD-Tree, grid index, or quadtree.

[0074] Optionally, a hash table structure can be built to store these coordinates, forming an indexing mechanism for fast lookups.

[0075] S103. Determine multiple candidate camera extrinsic parameters and the two-dimensional projection coordinates of multiple sampling points of the three-dimensional target detection data based on the candidate camera extrinsic parameters; the candidate camera extrinsic parameters are obtained by adjusting the current camera extrinsic parameters.

[0076] Here, candidate camera extrinsics can refer to the combination of rotation matrices and translation vectors generated through parameter space sampling. Specifically, they can be generated by superimposing small perturbations in different directions on the current extrinsics to explore a better extrinsic solution space. Current camera extrinsics can refer to the camera extrinsic parameters before this optimization adjustment.

[0077] Sampling points refer to a predefined set of points in a three-dimensional coordinate system, such as a vehicle coordinate system or a ground coordinate system. For example, sampling points can be uniformly sampled on six surfaces formed by the boundary of a three-dimensional coordinate system. Each surface can have a preset number of sampling points. As an example, each surface can sample 10×10 points, for a total of 600 sampling points. Sampling points can include the edges and interiors of the bounding box in the three-dimensional coordinate system, used to comprehensively represent the projection area of ​​the 3D bounding box. Two-dimensional projected coordinates refer to the theoretical positions on the two-dimensional image plane obtained by using the camera's intrinsic parameters and a candidate extrinsic parameters of a camera, through perspective projection transformation of the sampling points in the three-dimensional world coordinate system.

[0078] For example, the vehicle-mounted device can generate N sets of candidate camera extrinsic parameters by applying multiple sets of parameter perturbations, such as random offsets of rotation angle ±0.5° and translation amount ±0.1m, based on the current camera extrinsic parameters, for example, the initial calibration value or the extrinsic parameter value optimized in the previous cycle.

[0079] As an example, the rotation angle adjustment range may include ±2 degrees for each axis (roll, pitch, yaw); the translation adjustment range may include ±10 centimeters for each direction (x, y, z).

[0080] As another example, the rotation angle adjustment range may include ±0.2 degrees; the translation adjustment range may include ±2 centimeters.

[0081] S104. Query the coordinate index structure to determine the coordinate index result of the two-dimensional projected coordinates.

[0082] The coordinate index result can refer to the result returned when a two-dimensional projected coordinate is queried using the constructed coordinate index structure; it can also refer to the target pixel coordinates that are closest to the projected coordinates, or the information of that neighboring point.

[0083] For example, the vehicle-mounted device can perform a query on the index structure constructed in S102 for the two-dimensional projected coordinates corresponding to each candidate extrinsic parameter, and return a query result indicating whether the corresponding target pixel coordinates exist. Furthermore, the projection error of the candidate camera extrinsic parameters can be calculated by checking whether the index query returns the number of corresponding target pixel coordinates.

[0084] Optionally, during the parameter optimization stage, multiple sets of candidate camera extrinsic parameters can be generated, and the preset 3D sampling points can be projected onto the image plane through each candidate extrinsic parameter. A hash table can then be used to quickly determine whether the projected points fall within the target area.

[0085] S105. Based on the projection error obtained from the coordinate index results, the target camera extrinsic parameters of the vehicle camera are selected from multiple candidate camera extrinsic parameters.

[0086] The projection error refers to the mismatch between the candidate extrinsic parameters, which project the 3D sampling points onto the image plane, and the actual pixel distribution. Specifically, it can be calculated by statistically analyzing the hit rate of the projected coordinates in the coordinate index structure, reflecting the accuracy of the extrinsic parameter calibration. The target camera extrinsic parameters refer to the set of extrinsic parameters selected from numerous candidate camera extrinsic parameters that minimizes the projection error or satisfies specific optimization conditions. These can include rotation and translation parameters.

[0087] For example, for each candidate camera extrinsic parameter, the vehicle-mounted device can calculate the number of times the two-dimensional projection coordinates of all its sampling points match the target pixel coordinates, and use this to calculate the projection error of the candidate camera extrinsic parameter; and according to the projection error of the candidate camera extrinsic parameter, the target camera extrinsic parameter of the vehicle camera can be selected from multiple adjusted candidate camera extrinsic parameters.

[0088] For example, the candidate camera extrinsic parameter with the smallest projection error can be used as the target camera extrinsic parameter, thereby improving the accuracy of extrinsic parameter adjustment for vehicle cameras.

[0089] Optionally, by statistically analyzing the projection hit rate of each candidate camera's extrinsic parameters, the extrinsic parameter with the highest hit rate can be selected as the optimization result, ensuring that the projection result after adjusting the extrinsic parameters has the highest probability of matching the actual distribution of the detected targets.

[0090] In this embodiment, slow-moving dynamic targets and / or static targets continuously detected during driving are directly used as calibration references. By establishing a pixel coordinate index and parameter space search mechanism, complex geometric calculations can be simplified into low-complexity table lookup operations, significantly improving computational efficiency. Simultaneously, this embodiment eliminates dependence on dedicated equipment and overcomes the limitation of environmental feature stability, thus allowing for stable and reliable camera extrinsic parameters to be obtained through calibration without relying on specific calibration scenarios. Therefore, this embodiment utilizes the characteristics of the targets detected by the vehicle itself to construct a dynamically updatable coordinate index structure. Combined with a multi-candidate parameter projection matching index, it can perform statistical optimization using massive amounts of driving data. This satisfies the requirement of obtaining stable and highly reliable camera extrinsic parameters through calibration without relying on specific calibration scenarios, improving the accuracy of vehicle camera extrinsic parameter adjustment and calibration.

[0091] In one exemplary embodiment, such as Figure 2 As shown, Figure 2 This is a flowchart illustrating a step for obtaining target pixel coordinates provided in an embodiment of this application. Figure 1 Based on this, the steps of the vehicle camera extrinsic parameter adjustment method are explained by example. In step S101, the target pixel coordinates corresponding to the road environment target extracted from the camera image data are obtained, including steps S201 to S203, wherein:

[0092] S201. Obtain the target pixel mask after semantic segmentation of camera image data.

[0093] S202. Morphological dilation is performed on the mask region corresponding to the road environment target to obtain the dilated pixel mask.

[0094] S203. Traverse the dilated pixel mask, extract the pixel coordinates of the mask area corresponding to the road environment target, and obtain the target pixel coordinates.

[0095] The target pixel mask can be a two-dimensional binary matrix, including 0 and 1 values. For example, a pixel with a value of 1 can belong to the road environment target (stationary / low-speed target) region. The dilated pixel mask can be the target pixel mask after morphological dilation of the road environment target, which can cover the original target area and the extended boundary area.

[0096] For example, the vehicle-mounted device can generate binary region labels by performing pixel-level classification of camera images using semantic segmentation algorithms. For instance, a fully convolutional neural network model can be used to distinguish road environment targets from background areas. The vehicle-mounted device can achieve morphological dilation by performing image processing operations on the target region's boundaries using structuring elements. For instance, circular or rectangular kernels can be used for convolution operations to eliminate jagged edges generated by semantic segmentation and expand the effective area coverage. The dilation operation can tolerate certain segmentation errors and extrinsic parameter deviations. The expanded region labels obtained after morphological dilation processing can then be used to obtain a dilated pixel mask. The vehicle-mounted device can traverse each pixel of the dilated mask; if the pixel value = 1, its coordinates can be recorded into the target coordinate set to determine the target pixel coordinates.

[0097] In practical applications, after acquiring camera image data, pixel regions containing road environment targets can be identified using a semantic segmentation model to generate an initial binary mask. A morphological dilation algorithm is then used to expand the boundaries of the mask region, eliminating missing edge pixels caused by insufficient segmentation accuracy. The dilated mask region is then scanned pixel-by-pixel, and the coordinates of all pixels belonging to the road environment targets are aggregated to form a target coordinate set used for extrinsic parameter calculations. This target coordinate set includes the target pixel coordinates.

[0098] Optionally, using pedestrians as the target in the road environment, the vehicle-mounted device can read the pedestrian mask output by the semantic segmentation network from the data dropped onto the disc. The mask can be a binary image, with the pedestrian region pixel value being 1 and the background being 0. Furthermore, a morphological dilation operation can be performed on the original pedestrian mask. The dilation operation tolerates a certain amount of segmentation error and extrinsic parameter deviation, thereby improving the accuracy of camera extrinsic parameter adjustment.

[0099] In this embodiment, by introducing morphological dilation processing, the coverage of the mask area corresponding to the road environment target is effectively expanded, ensuring that all effective pixels of the actual target edge are included, avoiding the problem of coordinate omission caused by segmentation error, thereby improving the accuracy of camera extrinsic parameter adjustment.

[0100] In an exemplary embodiment, step S202 involves morphological dilation of the mask region corresponding to the road environment target to obtain a dilated pixel mask. This process may include, for example:

[0101] Obtain the correspondence between the preset distance threshold and the expansion range.

[0102] Based on the correspondence and the distance between the vehicle and the road environment target, the mask area corresponding to the road environment target is morphologically dilated.

[0103] The distance threshold can refer to a pre-defined physical distance range, such as close distance (<20m), medium distance (20-40m), and long distance (>40m). The dilation range refers to the size of the structuring element in the morphological dilation operation, which determines the pixel width by which the target mask area expands outward. The distance between the vehicle and road environment targets can refer to the three-dimensional straight-line distance calculated using onboard sensors (millimeter-wave radar, lidar) or monocular ranging algorithms.

[0104] For example, the vehicle-mounted device can determine the size range of structural elements in the morphological dilation operation based on the physical distance between the vehicle and the detected target. This can be achieved, for instance, by using a pre-established mapping table or piecewise function, such as dividing the distance into multiple intervals and assigning a different dilation kernel size to each interval. When a stationary or low-speed target is detected while the vehicle is moving, the distance between the vehicle and the road environment target can be obtained using LiDAR or visual ranging methods. A pre-defined correspondence can be queried based on this distance; for example, a smaller dilation kernel size can be selected when the distance is close to avoid over-dilation, while a larger dilation kernel size can be selected when the distance is far to compensate for the positioning deviation of the target pixel coordinates. Furthermore, the selected dilation kernel size can be used to perform morphological dilation processing on the mask area. For example, a 5×5 rectangular kernel can be used to dilate near-distance targets, and a 9×9 rectangular kernel can be used to dilate far-distance targets, thereby generating a dilated pixel mask that can cover the actual physical boundaries of the target.

[0105] Optionally, the correspondence between the preset distance threshold and the expansion range can be as follows:

[0106] Close range (<20m): the size of the expansion kernel is 5×5 pixels;

[0107] Mid-range (20-40m): The size of the expansion kernel is 7×7 pixels;

[0108] Long distance (>40m): The size of the expansion kernel is 9×9 pixels.

[0109] In this embodiment, by dynamically adjusting the expansion range, the expansion degree of the mask area is matched with the target distance, which avoids pixel coordinate redundancy for near-range targets and improves the coordinate coverage integrity for far-range targets. This improves the accuracy of extracting target pixel coordinates.

[0110] In an exemplary embodiment, the coordinate index structure includes a coordinate hash table; in step S102, constructing the coordinate index structure based on the target pixel coordinates includes:

[0111] The target pixel coordinates are used as the index data of the coordinate hash table, and the preset query matching identifier is used as the target value data of the coordinate hash table.

[0112] Construct a coordinate hash table based on the index data and target value data.

[0113] In this context, the coordinate hash table refers to a hash structure that stores data using pixel coordinates as keys. For example, it can be implemented using open addressing or chaining, mapping two-dimensional coordinates to one-dimensional storage space through a hash function. The index data refers to the set of target pixel coordinate values, which can be implemented using integer-processed coordinate values, such as rounding down floating-point coordinates. The target value data refers to identifiers used to indicate whether a coordinate exists. For example, it can be implemented using boolean values ​​or enumeration values, such as using 1 to indicate the existence of a matching coordinate and 0 to indicate its non-existence.

[0114] For example, when constructing the coordinate hash table, the target pixel coordinates can be normalized, such as converting the coordinate values ​​into integer form. A hash function can be used to calculate the storage location corresponding to each coordinate, and a preset query matching identifier can be written to that location. For example, when the target pixel coordinates are (100, 200), they can be converted into hash keys and stored in the hash table, while the target value data is set to 1. During the query phase, the two-dimensional projection coordinates corresponding to the candidate camera extrinsic parameters are calculated using the same hash function. If a matching key value can be retrieved in the hash table, the two-dimensional projection coordinates are considered valid; otherwise, they are considered invalid.

[0115] Optionally, after morphological dilation, the vehicle-mounted device can traverse the dilated mask and extract all pixel coordinates with a value of 1. These pixel coordinates with a value of 1 can be used to construct a coordinate hash table, where the key is a coordinate tuple (x, y) and the value can be True. This can be implemented using a Python dictionary or C++'s unordered_set, ensuring a lookup complexity of O(1).

[0116] In this embodiment, a hash table is used to implement coordinate indexing, reducing the query complexity to a constant level. Ideally, the hash table can complete the query in only O(1) time, thereby improving the efficiency of projected coordinate matching. This can reduce the consumption of computing resources and improve the speed of external parameter filtering. In this way, stable and reliable camera external parameters can be obtained by adjusting the calibration without relying on a specific calibration scenario.

[0117] In one exemplary embodiment, the method further includes:

[0118] If the target value data corresponding to the two-dimensional projected coordinates is found in the coordinate hash table, the coordinate index result indicates that the two-dimensional projected coordinates have been matched.

[0119] If no target value data corresponding to the two-dimensional projected coordinates is found in the coordinate hash table, the coordinate index result indicates that the two-dimensional projected coordinates did not match.

[0120] The projection error is obtained by determining the hit rate based on the number of matched two-dimensional projection coordinates and the number of sampling points.

[0121] The target value data can refer to verification identifiers pre-stored in a hash table. For example, these can be implemented using Boolean values ​​or fixed numerical codes, used to mark whether a coordinate point belongs to a valid detection target. A hit match means that the projected coordinates and the detection target coordinates have a corresponding relationship in the hash table. For example, this can be achieved through hash key-value comparison, used to verify the projection accuracy of candidate extrinsic parameters. A miss match means that the projected coordinates are not within the range of the detection target coordinates, specifically manifested as an invalid identifier in the hash table query result, used to exclude erroneous projection points. The hit rate can refer to the ratio of the number of valid projection points to the total number of sampled points. For example, it can be calculated by statistically analyzing the ratio of the number of matches to the total number of samples, used to quantify the projection error of candidate extrinsic parameters.

[0122] For example, after constructing the coordinate hash table, the vehicle-mounted device can input the two-dimensional projected coordinates of the sampling points of the candidate camera extrinsic parameters into the hash table for batch queries. When the two-dimensional projected coordinates exist in the hash table, they are recorded as valid matching points; when the projected coordinates do not exist, they are marked as invalid projection points. By statistically analyzing the proportion of valid matching points to all sampling points, the hit rate is calculated as an inverse indicator of the projection error. For example, when the projection error of a candidate extrinsic parameter is low, its corresponding hit rate will be significantly higher than that of other candidate parameters. Therefore, by comparing the hit rate values ​​of different candidate extrinsic parameters, the optimal combination of extrinsic parameters can be quickly selected.

[0123] Optionally, the vehicle-mounted equipment can use the current camera intrinsic parameters K and extrinsic parameters [R|t] to project each sampling point P_3d onto the image plane, as shown in the following expression (1):

[0124] P_2d=K×[R|t]×P_3d (1);

[0125] Where P_2d represents the projected image coordinates (u, v), i.e., the two-dimensional projected coordinates;

[0126] For each projection point P_2d(u, v), it can be processed as follows:

[0127] The floating-point coordinates can be rounded to integer coordinates (u', v'); the hash table can be queried using an algorithm: hit = hash_table.get((u', v'), False); and the number of hits can be counted: if hit is True, then N_in += 1 in the algorithm.

[0128] The projection error can be calculated using the following algorithm:

[0129] Hit rate = N_in / N_total; Projection error = 1 - Hit rate; where N_total is the number of sampling points and N_in is the number of hits.

[0130] Therefore, this hash table-based method simplifies the complex point determination within a polygon to a simple table lookup operation, reducing the time complexity from O(n) to O(1), and significantly improving efficiency when processing massive amounts of data.

[0131] In this embodiment, the coordinate lookup complexity is reduced to a constant level by using a hash table structure, thereby improving the projection verification efficiency of large-scale sampling points by two orders of magnitude. Furthermore, the error calculation method based on hit rate avoids the computational burden of point-by-point distance calculation in traditional methods, achieving stable and highly reliable camera extrinsic parameters without relying on specific calibration scenarios.

[0132] Through the above technical solution, this application can complete the projection error evaluation of thousands of sampling points in milliseconds, effectively improving the calculation efficiency of projection error, thereby enabling stable and highly reliable camera extrinsic parameters to be obtained by adjusting the calibration without relying on a specific calibration scenario.

[0133] In an exemplary embodiment, in step S105, the target camera extrinsic parameters of the vehicle camera are selected from multiple candidate camera extrinsic parameters based on the projection error obtained from the coordinate index result. This may include, for example, the following:

[0134] Based on the projection errors of the candidate camera extrinsic parameters onto multiple road environment targets, the cumulative projection error corresponding to the candidate camera extrinsic parameters is obtained.

[0135] The candidate camera extrinsic parameters with the smallest cumulative projection error are used as the target camera extrinsic parameters.

[0136] The cumulative projection error refers to the total error value obtained by summing the projection errors of the same candidate extrinsic parameter on multiple road environment targets. For example, it can be achieved by weighted summation of the hit rate error of each detected target. This error reflects the overall matching degree of the candidate extrinsic parameter on all detected targets.

[0137] For example, during vehicle movement, after multiple stationary or low-speed targets are detected and their pixel coordinates are extracted, each candidate camera extrinsic parameter projects the sampling points in three-dimensional space onto a two-dimensional image plane to form projected coordinates. The projection error of each candidate camera extrinsic parameter can be calculated by statistically analyzing the hit rate of the two-dimensional projected coordinates corresponding to all detected targets in a coordinate hash table. Furthermore, the projection errors of the same candidate camera extrinsic parameter on different detected targets can be accumulated to form a cumulative projection error characterizing the global matching degree. Thus, the parameter with the smallest cumulative error among multiple candidate camera extrinsic parameters can be selected as the optimal solution, thereby ensuring that the extrinsic parameter adjustment results have optimal projection consistency across multiple targets.

[0138] Optionally, the vehicle-mounted device can accumulate the projection errors of all pedestrian samples (typically >10,000) and select the candidate camera extrinsic parameters with the smallest total error as the optimal camera extrinsic parameters. It can also record the error improvement rate before and after optimization to evaluate the optimization effect.

[0139] In this embodiment, by accumulating the projection errors of multiple detection targets, the influence of outliers from a single target can be effectively eliminated. The multi-target cumulative error evaluation mechanism enables the extrinsic parameter calibration results to have higher stability and reliability in complex road environments. Simultaneously, the selection method based on cumulative error can improve the accuracy of extrinsic parameter adjustment.

[0140] In one exemplary embodiment, such as Figure 3 As shown, Figure 3 This application provides a flowchart illustrating the steps for determining the extrinsic parameters of a target camera, which can be implemented in an embodiment of the present application. Figure 1 Based on this, the method for adjusting the extrinsic parameters of a vehicle camera further includes the following steps, which are illustrated by example: The candidate camera extrinsic parameters include a combination of candidate extrinsic parameters, which in turn includes multiple rotation axis parameters and multiple translation direction parameters; the method for adjusting the extrinsic parameters of a vehicle camera may also specifically include:

[0141] S301. Determine the cumulative projection error corresponding to the candidate extrinsic parameter combinations used for coarse search adjustment;

[0142] S302. The candidate extrinsic parameter combination with the smallest cumulative projection error after coarse search adjustment is taken as the target extrinsic parameter combination for coarse search adjustment.

[0143] S303. Determine the cumulative projection error corresponding to the candidate extrinsic parameter combinations used for fine search adjustment; wherein, the candidate extrinsic parameter combinations for fine search adjustment are obtained by adjusting the target extrinsic parameter combinations for coarse search adjustment;

[0144] S304. The candidate extrinsic parameter combination with the smallest cumulative projection error after fine search and adjustment is taken as the target extrinsic parameter combination for the vehicle camera.

[0145] The candidate extrinsic parameter combination refers to a set of rotation and translation parameters between the camera coordinate system and the vehicle coordinate system. For example, rotation parameters can be represented by Euler angles, quaternions, or axis angles, and translation parameters can be represented by three-dimensional vectors. This combination is used to characterize the spatial pose relationship of the camera on the vehicle, and extrinsic parameter deviations can be corrected by adjusting the parameters.

[0146] The coarse search adjustment refers to preliminary screening within a large parameter space. For example, it can be achieved by discretizing the rotation and translation parameters using a preset coarse search step size. This step quickly narrows the parameter search range by using a larger step size, avoiding the computational burden of global traversal.

[0147] Fine-tuning the search can refer to performing a localized, refined search based on the coarse search results. For example, it can be achieved by using a smaller fine-tuning step size to fine-tune the parameters in the neighborhood of the target extrinsic parameter combination. This step improves the precision of parameter adjustment by using a smaller step size, ensuring the accuracy of the extrinsic parameter calibration results.

[0148] For example, the vehicle-mounted device can generate candidate extrinsic parameter combinations based on a coarse search range, such as traversing within a range of ±2 degrees for rotation parameters and ±10 cm for translation parameters with step sizes of 0.2 degrees and 2 cm, respectively. After calculating the cumulative projection error of each candidate extrinsic parameter combination to the road environment target, the combination with the smallest error is selected as the coarse adjustment result. Further, using the target extrinsic parameter combination obtained from this coarse search adjustment as the center, within a smaller range, such as ±0.2 degrees for rotation parameters and ±2 cm for translation parameters, fine-tuned candidate extrinsic parameter combinations are generated with step sizes of 0.2 degrees and 2 cm, respectively. The candidate extrinsic parameter combination with the smallest error is then selected again as the final target extrinsic parameter combination. In this way, through a staged adjustment strategy, the high computational cost of global search is avoided while ensuring the accuracy of parameter adjustment.

[0149] In this embodiment, a multi-level precision search strategy reduces computational complexity by orders of magnitude while maintaining calibration accuracy. This allows for stable and highly reliable camera extrinsic parameters to be obtained without relying on specific calibration scenarios. The coarse search phase quickly eliminates significantly deviating parameter combinations, while the fine search phase refines the candidate parameters. This enables the extrinsic parameter adjustment process to be completed within milliseconds while controlling projection errors within pixel-level precision, thereby improving the accuracy of real-time extrinsic parameter adjustment.

[0150] In one exemplary embodiment, the method further includes:

[0151] Based on a preset coarse search range, the current camera extrinsic parameters are adjusted with a preset coarse search step size to determine multiple candidate extrinsic parameter combinations for coarse search adjustment;

[0152] Based on a preset fine search range and a preset fine search step size, the target extrinsic parameter combination for coarse search adjustment is adjusted to determine multiple candidate extrinsic parameter combinations for fine search adjustment.

[0153] The coarse search range refers to the initial search interval for camera extrinsic parameters in terms of rotation axis and translation direction parameters. For example, it can be achieved using pre-defined angle and displacement ranges, such as ±2 degrees for rotation and ±10 centimeters for translation, which can be used to quickly filter out potential optimal solutions within a large range. The coarse search step size refers to the interval of parameter adjustment, such as setting the angle step size to 0.2 degrees and the displacement step size to 2 centimeters, reducing the computational load by using a larger step size.

[0154] The fine search range refers to a parameter interval that is further narrowed based on the target extrinsic parameter combination obtained from the coarse search. For example, the angle range is narrowed to ±0.2 degrees and the displacement range is narrowed to ±2 centimeters, which is used for fine adjustment within a local range. The fine search step size uses a smaller interval, for example, an angle step size of 0.2 degrees and a displacement step size of 2 centimeters, to improve the accuracy of parameter adjustment.

[0155] For example, a coarse search range can cover the possible range of extrinsic parameter deviations, and candidate extrinsic parameter combinations can be quickly generated using a larger step size, avoiding the computational burden of traversing the entire parameter space. Furthermore, based on the target extrinsic parameter combination with the smallest cumulative projection error in the coarse search results, fine-tuning can be performed within the fine search range with smaller step sizes. For example, by gradually reducing the adjustment range of rotation angle and translation, the optimal extrinsic parameter combination can be gradually approximated, thereby achieving a balance between accuracy and efficiency with limited computational resources.

[0156] Optionally, the vehicle-mounted equipment can improve search efficiency while ensuring accuracy through a multi-level precision search strategy. For example, this can be achieved in the following ways:

[0157] The parameter search space definition for the first-level coarse search can include: rotation angle adjustment range: ±2 degrees for each axis (roll, pitch, yaw); translation adjustment range: ±10 cm for each direction (x, y, z). The search step size can include: rotation step size: 0.2 degrees; translation step size: 2 cm; approximately 21×21×21×11×11×11 can be generated as candidate extrinsic parameter combinations.

[0158] The second-level fine search can define the fine search range based on the optimal parameters found in the first level, including: rotation angle fine search range: ±0.2 degrees (i.e., the step size of the previous level); translation fine search range: ±2 centimeters (i.e., the step size of the previous level); the fine search step size can include: rotation step size 0.02 degrees; translation step size 0.2 centimeters; in this way, this sub-step fine search strategy can make more precise parameter adjustments near the optimal solution.

[0159] Optionally, in the operation of multi-level precision extrinsic optimization, the parameter space can be divided into multiple subspaces, and the projection error of each candidate extrinsic parameter combination can be calculated in parallel using multi-threading or GPU. Each thread processes a set of candidate extrinsic parameters independently without interfering with each other.

[0160] In this embodiment, a phased adjustment strategy is adopted. First, a coarse search is used to quickly narrow down the parameter range, and then a fine search is used to achieve local optimization. This significantly reduces the consumption of computational resources and avoids the problems of insufficient accuracy or low efficiency caused by a single step size. In this way, rapid calibration of camera extrinsic parameters can be achieved during vehicle dynamic driving. Through the synergistic effect of coarse and fine search, the computational latency is effectively reduced while ensuring calibration accuracy. Thus, stable and highly reliable camera extrinsic parameters can be obtained by adjusting the calibration without relying on a specific calibration scenario.

[0161] In one exemplary embodiment, Figure 4 This application provides a flowchart illustrating a specific implementation of a method for adjusting the extrinsic parameters of a vehicle camera. This specific implementation can be used to explain the detailed steps of the method for adjusting the extrinsic parameters of a vehicle camera. Figure 4 As shown, the method may include, for example:

[0162] S401, Read vehicle-end platen data and perform data preprocessing; including reading 3D detection box data, reading semantic segmentation mask, and time synchronization alignment;

[0163] S402, Hash table construction may include: reading pedestrian segmentation mask, adaptive expansion processing, extracting pedestrian region coordinates, and constructing coordinate hash table;

[0164] S403, Projection error calculation, including: dense sampling of 3D bounding box (600 points), projection onto 2D using the current extrinsic parameters, hash table lookup (O(1) complexity), and statistical hit rate;

[0165] S404, Level 1 coarse search, including: rotation angle adjustment range: ±2 degrees for each axis (roll, pitch, yaw); translation adjustment range: ±10 cm for each direction (x, y, z). The search step size can include: rotation step size: 0.2 degrees; translation step size: 2 cm; all candidate extrinsic parameter combinations are computed in parallel.

[0166] S405. Select the optimal parameters, including: sum up all projection errors and find the candidate combination of extrinsic parameters with the minimum total error;

[0167] S406, Second-level fine search, including: Rotation angle fine search range: ±0.2 degrees (i.e., the step size of the previous level); Translation fine search range: ±2 centimeters (i.e., the step size of the previous level); Fine search step size may include: Rotation step size 0.02 degrees; Translation step size 0.2 centimeters;

[0168] S407. Determine the final optimization result, including: outputting the optimized extrinsic parameter matrix and recording the error improvement rate;

[0169] S408. Determine whether the error improvement is significant. For example, you can determine whether it is significant by comparing the error improvement rate before and after optimization with the improvement rate threshold.

[0170] If so, S409, update the saved results of camera extrinsic parameters;

[0171] If not, S410, maintain the original camera extrinsic parameter recording log.

[0172] It is understood that the implementation of this embodiment can be explained through the above embodiments, and will not be repeated here.

[0173] In this embodiment, the coordinates of pedestrian regions in the camera image are constructed as a hash table, and the complex projection error measurement calculation is abstracted into an O(1) lookup operation, thereby greatly improving the efficiency of processing massive amounts of data; a multi-level optimization strategy of coarse search and fine search is adopted, and based on the optimal solution found by coarse search, a precise search is performed in the vicinity using a sub-step size, which takes into account both search efficiency and optimization accuracy; pedestrians with low moving speed can be selected as calibration targets, which effectively reduces the impact of sensor synchronization error on calibration accuracy; the detection and segmentation results of the vehicle-mounted disk can be directly used, without the need for dedicated calibration equipment and site, realizing fully automated external parameter fine-tuning; global optimization can be performed by collecting a large number of pedestrian samples (>10,000 frames), which improves the stability and reliability of calibration results.

[0174] Alternatively, a KD-tree (K-Dimensional Tree) can be used: suitable for nearest neighbor queries in high-dimensional spaces; query complexity O(log n), construction complexity O(n log n); suitable for scenarios that require querying the nearest point rather than an exact match; or a bitmap method: representing the entire image as a bitmap, with pedestrian areas marked as 1; fixed memory usage, fast query speed; suitable for scenarios where the image resolution is not particularly high, replacing the coordinate hash table query index method.

[0175] Alternatively, a genetic algorithm can be used: it has strong global search capabilities and is less prone to getting trapped in local optima; it can adaptively adjust the search direction and step size; however, the computation time is relatively long. Gradient descent can be used: if the projection error function is differentiable, gradient information can be used to accelerate convergence; however, it requires calculating the gradient of the error with respect to the parameters; it is sensitive to initial values ​​and may get trapped in local optima; these methods can replace the multi-level search implementation methods mentioned above.

[0176] Optionally, a multi-target fusion approach can be used, simultaneously employing pedestrians, cyclists, and stationary targets (streetlights, signs); different weights can be assigned based on target type and movement speed; this yields richer constraint information. Static features such as lane lines can also be used: completely eliminating synchronization errors caused by motion; this approach has the advantage of performing well in scenarios such as highways.

[0177] In this application, the time complexity of projection error calculation is reduced from O(n) to O(1) by using a hash table lookup method, and the processing time for a single sample is reduced from milliseconds to microseconds, resulting in an overall processing efficiency improvement of more than 100 times, achieving a revolutionary improvement in computational efficiency. No dedicated calibration site needs to be built, and no calibration boards need to be purchased or maintained, reducing the calibration cost per vehicle by more than 90%. Traditional methods require 500,000 to 1 million yuan to build a calibration site, while this application only requires software development costs, offering a cost advantage. Through a multi-level precision search strategy and massive data statistics, the external parameter calibration accuracy is improved by more than 30%. The precision search step size can reach 0.02 degrees and 0.2 centimeters, far exceeding the accuracy of traditional methods, demonstrating improved accuracy. It directly uses existing detection and segmentation results from the vehicle end, requiring no manual intervention, supporting automatic calibration and continuous optimization of batch vehicles, achieving a fully automated effect. It can work in various road environments, unaffected by weather, lighting, or other conditions, making it particularly suitable for the external parameter maintenance of large-scale fleets, demonstrating strong adaptability.

[0178] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0179] The following describes the extrinsic parameter adjustment device for a vehicle camera provided in the embodiments of this application. The extrinsic parameter adjustment device for a vehicle camera has the same inventive concept as the extrinsic parameter adjustment method for a vehicle camera described above. The solution to the problem provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more embodiments of the extrinsic parameter adjustment device for a vehicle camera provided below can be referred to the limitations of the extrinsic parameter adjustment method for a vehicle camera described above. The extrinsic parameter adjustment device for a vehicle camera described below and the extrinsic parameter adjustment method for a vehicle camera described above can be referred to each other, and will not be repeated here.

[0180] In one exemplary embodiment, Figure 5 This is a schematic diagram of the structure of a vehicle camera extrinsic parameter adjustment device provided in an embodiment of this application, as shown below. Figure 5 As shown, the extrinsic parameter adjustment device 50 for the vehicle camera includes: a vehicle data reading module 510, a coordinate index construction module 520, a projection coordinate determination module 530, a coordinate index result module 540, and a target extrinsic parameter screening module 550, wherein:

[0181] The vehicle data reading module 510 is used to read vehicle landing data and extract three-dimensional target detection data and camera image data of stationary or low-speed moving road environment targets from the vehicle landing data; wherein, the low-speed movement is movement below a preset speed threshold.

[0182] The coordinate index construction module 520 is used to obtain the target pixel coordinates corresponding to the road environment target in the camera image data; and to construct the coordinate index structure based on the target pixel coordinates.

[0183] The projection coordinate determination module 530 is used to determine multiple candidate camera extrinsic parameters and the two-dimensional projection coordinates of multiple sampling points of the three-dimensional target detection data based on the candidate camera extrinsic parameters; the candidate camera extrinsic parameters are obtained by adjusting the current camera extrinsic parameters.

[0184] The coordinate index result module 540 is used to query the coordinate index structure and determine the coordinate index result of the two-dimensional projected coordinates.

[0185] The target extrinsic parameter screening module 550 is used to filter the target camera extrinsic parameters of the vehicle camera from multiple candidate camera extrinsic parameters based on the projection error obtained from the coordinate index results.

[0186] In an exemplary embodiment, the pixel coordinate acquisition module 510 is used to acquire the target pixel mask obtained after semantic segmentation of camera image data; perform morphological dilation on the mask region corresponding to the road environment target to obtain the dilated pixel mask; traverse the dilated pixel mask to extract the pixel coordinates of the mask region corresponding to the road environment target to obtain the target pixel coordinates.

[0187] In an exemplary embodiment, the pixel coordinate acquisition module 510 is used to acquire the correspondence between a preset distance threshold and the dilation range; and to perform morphological dilation on the mask region corresponding to the road environment target according to the correspondence and the distance between the vehicle and the road environment target.

[0188] In an exemplary embodiment, the coordinate index structure includes a coordinate hash table; the coordinate index building module 520 is used to use the target pixel coordinates as the index data of the coordinate hash table and the preset query matching identifier as the target value data of the coordinate hash table; and to build the coordinate hash table according to the index data and the target value data.

[0189] In an exemplary embodiment, the apparatus further includes a projection error determination module. The projection error determination module is configured to: if a target value corresponding to the two-dimensional projection coordinates is found in the coordinate hash table, the coordinate index result indicates a two-dimensional projection coordinate match; if no target value corresponding to the two-dimensional projection coordinates is found in the coordinate hash table, the coordinate index result indicates a two-dimensional projection coordinate mismatch; and obtain the projection error based on a hit rate determined between the number of two-dimensional projection coordinate matches and the number of sampling points.

[0190] In an exemplary embodiment, the target extrinsic parameter screening module 550 is used to obtain the cumulative projection error corresponding to the candidate camera extrinsic parameters based on the projection error of the candidate camera extrinsic parameters for multiple road environment targets; and to use the candidate camera extrinsic parameter with the smallest cumulative projection error as the target camera extrinsic parameter.

[0191] In an exemplary embodiment, the candidate camera extrinsic parameters include candidate extrinsic parameter combinations, which include multiple rotation axis parameters and multiple translation direction parameters; the device also includes a multi-level search and adjustment module. The multi-level search and adjustment module is used to determine the cumulative projection error corresponding to the candidate extrinsic parameter combinations used for coarse search and adjustment; to select the candidate extrinsic parameter combination with the smallest cumulative projection error in the coarse search and adjustment as the target extrinsic parameter combination for the coarse search and adjustment; and to determine the cumulative projection error corresponding to the candidate extrinsic parameter combinations used for fine search and adjustment; wherein the candidate extrinsic parameter combinations for fine search and adjustment are adjusted based on the target extrinsic parameter combinations for coarse search and adjustment; and to select the candidate extrinsic parameter combination with the smallest cumulative projection error in the fine search and adjustment as the target extrinsic parameter combination for the vehicle camera.

[0192] In an exemplary embodiment, the level search adjustment module is used to adjust the current camera extrinsic parameters based on a preset coarse search range and a preset coarse search step size to determine multiple candidate extrinsic parameter combinations for coarse search adjustment; and to adjust the target extrinsic parameter combination for coarse search adjustment based on a preset fine search range and a preset fine search step size to determine multiple candidate extrinsic parameter combinations for fine search adjustment.

[0193] In one exemplary embodiment, this application also provides a computer-readable storage medium storing a computer program that, when executed by one or more processors, causes the one or more processors to perform the steps of any of the vehicle camera extrinsic parameter adjustment methods described in the above embodiments.

[0194] In one exemplary embodiment, this application also provides a vehicle-mounted device that stores a computer program. When computer-readable instructions are executed by one or more processors, the one or more processors cause the one or more processors to perform the steps of any of the vehicle camera extrinsic parameter adjustment methods described in the above embodiments.

[0195] In one exemplary embodiment, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the extrinsic parameter adjustment method for a vehicle camera as described in any of the above embodiments.

[0196] Indicatively, such as Figure 6 As shown, Figure 6 This is a schematic diagram of the internal structure of a vehicle-mounted device 600 provided in an embodiment of this application. The vehicle-mounted device 600 can be provided as a server. (Refer to...) Figure 6 The vehicle infotainment device 600 includes a processing component 602, which further includes one or more processors, and memory resources represented by memory 601 for storing instructions, such as applications, that can be executed by the processing component 602. The applications stored in memory 601 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 602 is configured to execute instructions to perform the extrinsic parameter adjustment method for the vehicle camera in any of the above embodiments.

[0197] The vehicle infotainment device 600 may also include a power supply component 603 configured to perform power management of the vehicle infotainment device 600, a wired or wireless network interface 604 configured to connect the vehicle infotainment device 600 to a network, and an input / output (I / O) interface 605. The vehicle infotainment device 600 can operate on an operating system stored in memory 601, such as Windows Server™, Mac OS X™, Unix™, Linux™, Free BSD™, or similar.

[0198] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the vehicle-mounted equipment to which the present application is applied. Specific vehicle-mounted equipment may include more or fewer components than those shown in the figure, or may combine certain components, or may have different component arrangements.

[0199] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0200] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0201] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

[0202] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for adjusting the extrinsic parameters of a vehicle camera, characterized in that, The method includes: Read vehicle landing data, and extract three-dimensional target detection data and camera image data of stationary or low-speed moving road environment targets from the vehicle landing data; wherein, the low-speed movement is movement below a preset speed threshold; Obtain the target pixel coordinates of the road environment target in the camera image data; and construct a coordinate index structure based on the target pixel coordinates; Multiple candidate camera extrinsic parameters are determined, as well as the two-dimensional projection coordinates of the candidate camera extrinsic parameters onto multiple sampling points of the three-dimensional target detection data; the candidate camera extrinsic parameters are obtained by adjusting the current camera extrinsic parameters. The coordinate index structure is queried to determine the coordinate index result of the two-dimensional projected coordinates; Based on the projection error obtained from the coordinate index result, the target camera extrinsic parameters of the vehicle camera are selected from multiple candidate camera extrinsic parameters.

2. The method according to claim 1, characterized in that, The step of obtaining the target pixel coordinates of the road environment target in the camera image data includes: Obtain the target pixel mask obtained after semantic segmentation of the camera image data; Morphological dilation is performed on the mask region corresponding to the road environment target to obtain a dilated pixel mask; Traverse the dilated pixel mask, extract the pixel coordinates of the mask region corresponding to the road environment target, and obtain the target pixel coordinates.

3. The method according to claim 2, characterized in that, The step of morphologically dilating the mask region corresponding to the road environment target to obtain a dilated pixel mask includes: Obtain the correspondence between the preset distance threshold and the expansion range; Based on the aforementioned correspondence and the distance between the vehicle and the road environment target, the mask region corresponding to the road environment target is morphologically dilated.

4. The method according to claim 2, characterized in that, The coordinate index structure includes a coordinate hash table; The step of constructing a coordinate index structure based on the target pixel coordinates includes: The target pixel coordinates are used as the index data of the coordinate hash table, and the preset query matching identifier is used as the target value data of the coordinate hash table. The coordinate hash table is constructed based on the index data and the target value data.

5. The method according to claim 4, characterized in that, The method further includes: If the target value data corresponding to the two-dimensional projected coordinates is found in the coordinate hash table, then the coordinate index result indicates that the two-dimensional projected coordinates have a match. If no target value data corresponding to the two-dimensional projected coordinates is found in the coordinate hash table, the coordinate index result indicates that the two-dimensional projected coordinates did not match. The projection error is obtained based on the hit rate determined by the number of matching two-dimensional projection coordinates and the number of sampling points.

6. The method according to claim 1, characterized in that, The step of selecting the target camera extrinsic parameters of the vehicle camera from a plurality of candidate camera extrinsic parameters based on the projection error obtained from the coordinate index result includes: Based on the projection errors of the candidate camera extrinsic parameters for multiple road environment targets, the cumulative projection error corresponding to the candidate camera extrinsic parameters is obtained; The candidate camera extrinsic parameters with the smallest cumulative projection error are used as the target camera extrinsic parameters.

7. The method according to claim 6, characterized in that, The candidate camera extrinsic parameters include a candidate extrinsic parameter combination, which includes multiple rotation axis parameters and multiple translation direction parameters. The method further includes: Determine the cumulative projection error corresponding to the candidate extrinsic parameter combinations used for coarse search adjustment; The candidate extrinsic combination with the smallest cumulative projection error after coarse search adjustment is taken as the target extrinsic combination for coarse search adjustment. Determine the cumulative projection error corresponding to the candidate extrinsic parameter combinations used for fine search adjustment; wherein, the candidate extrinsic parameter combinations for fine search adjustment are obtained by adjusting the target extrinsic parameter combinations for coarse search adjustment; The candidate extrinsic parameter combination with the smallest cumulative projection error after fine-search adjustment is used as the target extrinsic parameter combination for the vehicle camera.

8. The method according to claim 7, characterized in that, The method further includes: Based on a preset coarse search range, the current camera extrinsic parameters are adjusted with a preset coarse search step size to determine multiple candidate extrinsic parameter combinations for coarse search adjustment; Based on a preset fine search range and a preset fine search step size, the target extrinsic parameter combination for coarse search adjustment is adjusted to determine multiple candidate extrinsic parameter combinations for fine search adjustment.

9. A device for adjusting the external parameters of a vehicle camera, characterized in that, The device includes: The vehicle data reading module is used to read vehicle landing data and extract three-dimensional target detection data and camera image data of stationary or low-speed moving road environment targets from the vehicle landing data; wherein, the low-speed movement is movement below a preset speed threshold. A coordinate index construction module is used to obtain the target pixel coordinates of the road environment target in the camera image data; and to construct a coordinate index structure based on the target pixel coordinates; The projection coordinate determination module is used to determine multiple candidate camera extrinsic parameters, and the two-dimensional projection coordinates of the candidate camera extrinsic parameters on multiple sampling points of the three-dimensional target detection data; the candidate camera extrinsic parameters are obtained by adjusting the current camera extrinsic parameters; The coordinate index result module is used to query the coordinate index structure and determine the coordinate index result of the two-dimensional projected coordinates; The target extrinsic parameter screening module is used to select the target camera extrinsic parameters of the vehicle camera from a plurality of candidate camera extrinsic parameters based on the projection error obtained from the coordinate index result.

10. A vehicle infotainment system, comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Camera state judgment method based on road surface covering and Hash comparison

    CN121414847A

  • A camera state judgment method based on road surface masking and hash comparison

    CN121414847B