Vehicle camera external parameter optimization method and vehicle

By filtering and extracting various types of feature points, and combining epipolar geometry constraints and lane line physical world constraints, the camera extrinsic parameters are optimized, solving the problem of insufficient accuracy and robustness in vehicle camera calibration in existing technologies, and achieving calibration with higher accuracy and stronger robustness.

CN120997309APending Publication Date: 2025-11-21ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511098693.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing online road calibration methods mainly rely on single lane line physical feature constraints or keyframe matching, resulting in low accuracy and weak robustness of vehicle camera extrinsic parameter calibration, making them susceptible to the influence of dynamic obstacles and environmental changes.

Method used

By acquiring image sequences from multiple cameras, filtering valid frame images, and extracting matching feature point pairs from the same camera at different times, matching feature point pairs from different cameras at the same time, and discrete feature points of lane lines, the camera extrinsic parameters are optimized using various types of feature points. Combined with epipolar geometric constraints and lane line physical world constraints, loss is iteratively optimized.

Benefits of technology

It improves the calibration accuracy and robustness of camera extrinsic parameters, avoids angle drift and parameter estimation jitter, and enhances calibration adaptability and accuracy in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997309A_ABST
    Figure CN120997309A_ABST
Patent Text Reader

Abstract

The invention provides a vehicle camera external parameter optimization method and a vehicle, and relates to the technical field of calibration. The method comprises the following steps: in an external parameter on-line calibration process of at least two cameras of a vehicle, acquiring an acquisition image sequence of each camera in the at least two cameras, and screening effective frame images from the acquisition image sequence of each camera to obtain an effective frame image sequence of each camera; extracting effective features based on the effective frame image sequence of each camera; wherein the effective features comprise first matching feature point pairs of effective frame images of the same camera at different moments, second matching feature point pairs of effective frame images of different cameras at the same moment, and discrete feature points of lane lines; and optimizing external parameters of each camera based on the first matching feature point pair, the second matching feature point pair and the lane line discrete feature points corresponding to the effective frame image sequence. The method and the device are used for improving the precision and robustness of online road calibration of the vehicle camera external parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of calibration technology, specifically to a method for optimizing the extrinsic parameters of a vehicle camera and a vehicle. Background Technology

[0002] With the increasing prevalence of autonomous driving, the demand for vehicle sensor calibration is growing. Calibration is used for spatial synchronization between vehicle sensors and the vehicle body, and it is a fundamental key technology for autonomous driving. As a critical sensor for autonomous driving, ensuring the accuracy of vehicle camera calibration is crucial for the success of autonomous driving. Summary of the Invention

[0003] In view of this, embodiments of the present disclosure provide a method for optimizing vehicle camera extrinsics and a vehicle, so as to improve the accuracy and robustness of vehicle camera extrinsics in online road calibration.

[0004] In a first aspect, this disclosure provides a method for optimizing the extrinsic parameters of a vehicle camera, including:

[0005] During the online calibration of the extrinsic parameters of at least two cameras on a vehicle, the image sequence acquired by each of the at least two cameras is obtained, and valid frame images are filtered from the image sequence acquired by each camera to obtain a valid frame image sequence for each camera.

[0006] Based on the effective frame image sequence of each camera, effective features are extracted; wherein, the effective features include first matching feature point pairs of effective frame images of the same camera at different times, second matching feature point pairs of effective frame images of different cameras at the same time, and lane line discrete feature points;

[0007] Based on the first matching feature point pair, the second matching feature point pair, and the discrete feature points of the lane line corresponding to the effective frame image sequence, the extrinsic parameters of each camera are optimized.

[0008] Secondly, this disclosure provides an electronic device, including:

[0009] At least one processor; and

[0010] A memory communicatively connected to the at least one processor; wherein,

[0011] The memory stores at least one computer program that can be executed by the at least one processor, the at least one computer program being executed by the at least one processor to enable the at least one processor to perform the vehicle camera extrinsic optimization method as described in the first aspect.

[0012] Thirdly, this disclosure provides a computer program product, which includes a computer program that, when run in a processor, implements the vehicle camera extrinsic parameter optimization method described in the first aspect.

[0013] Fourthly, this disclosure provides a vehicle configured to perform the vehicle camera extrinsic parameter optimization method described in the first aspect.

[0014] The embodiments provided in this disclosure, during the online calibration of vehicle camera extrinsic parameters, extract the effective frame image sequence of each camera, and extract effective features based on the effective frame image sequence of each camera. The effective features include a first matching feature point pair of effective frame images from the same camera at different times, a second matching feature point pair of effective frame images from different cameras at the same time, and discrete lane line feature points. Based on the first matching feature point pair, the second matching feature point pair, and the discrete lane line feature points corresponding to the effective frame image sequence, the extrinsic parameters of each camera are optimized. This allows for the comprehensive use of matching feature point pairs of shared viewing areas from multiple consecutive frames of the same camera, matching feature point pairs of shared viewing areas from different cameras at the same time, and lane line feature points during the process of coarse-to-fine calibration of camera extrinsic parameters. This achieves fine-tuning of camera extrinsic parameters, integrates multiple types of feature points to optimize camera extrinsic parameters, avoids the defects of single-type feature point optimization, and improves the accuracy and robustness of camera extrinsic parameter optimization. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0016] Figure 1 The diagram shown is a schematic flowchart of the vehicle camera extrinsic parameter optimization method in an embodiment of this disclosure.

[0017] Figure 2 The figure shown is a schematic diagram of the feature point positions of a single camera and a single frame in an embodiment of this disclosure;

[0018] Figure 3 The diagram shown is a schematic representation of the feature point positions of multiple cameras simultaneously in multiple frames in an embodiment of this disclosure.

[0019] Figure 4 The diagram shown is a schematic representation of the process of online optimization of vehicle camera extrinsic parameters in an embodiment of this disclosure.

[0020] Figure 5 The diagram shown is a schematic diagram of the vehicle camera extrinsic parameter optimization device in an embodiment of this disclosure.

[0021] Figure 6 The diagram shown is a structural schematic of an electronic device in an embodiment of this disclosure. Detailed Implementation

[0022] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0023] Where there is no conflict, the various embodiments of this disclosure and the features thereof in the embodiments may be combined with each other.

[0024] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.

[0025] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Words such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.

[0026] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.

[0027] Overview

[0028] There are two main methods for calibrating vehicle cameras:

[0029] First, static calibration, which determines the camera's extrinsic parameters using single-frame images in a calibration room or fixed location;

[0030] Second, dynamic calibration, which can be completed using multiple frames of images in environments that do not require prior planning (such as roads).

[0031] Online road calibration is a type of dynamic calibration. Most existing online road calibration methods rely on vehicle planar motion for fine-tuning optimization, primarily using lane line physical feature constraints or keyframe-matched feature points for epipolar constraint optimization. These methods are simplistic and prone to angle drift; surrounding dynamic obstacles also significantly impact the calibration results. This leads to low accuracy and weak robustness in current online road calibration methods.

[0032] Exemplary methods

[0033] The vehicle camera extrinsic parameter optimization method provided in this disclosure can be applied to the vehicle's main controller, or to other terminals or servers that can communicate with the vehicle and / or sensors installed on the vehicle.

[0034] The vehicle camera extrinsic parameter optimization method provided in this disclosure embodiment, such as Figure 1 As shown, the main steps include:

[0035] Step 101: During the online calibration of the extrinsic parameters of at least two cameras of the vehicle, the image sequence of each of the at least two cameras is acquired, and valid frame images are filtered from the image sequence of each camera to obtain a valid frame image sequence of each camera.

[0036] In some embodiments, after the online calibration process of at least two cameras on the vehicle is triggered, vehicle body parameters, initial extrinsic values ​​of each camera, and intrinsic parameters of each camera are acquired. Based on the acquired vehicle body parameters, initial extrinsic values ​​of the cameras, and intrinsic parameters of the cameras, the error of the camera extrinsic parameters in the acquired images is evaluated, and the camera extrinsic parameters are optimized based on the evaluation results.

[0037] In some embodiments, the step of filtering valid frame images from the acquired image sequence of each camera to obtain a valid frame image sequence for each camera includes: performing quality filtering on the acquired image sequence of each camera; filtering valid frame images from the quality-filtered acquired image sequence of each camera based on a valid frame window, wherein the valid frame images include turning valid frame images and straight-line valid frame images; and determining a valid frame image sequence for each camera based on the valid frame images within the valid frame window.

[0038] It includes both turning and straight-line effective frames, enabling the inclusion of the complete motion cycle, preventing possible angle drift and parameter estimation jitter, and effectively solving the problem that large disturbances in the initial extrinsic parameters cause nonlinear optimization to fall into local optima.

[0039] The purpose of quality filtering is mainly to remove low-quality data, such as removing blurry frames with image entropy <5 bits or low texture frame feature points <30 from the acquired image sequence.

[0040] The effective frame window includes both turning and straight-line movement, thus including complete motion constraints and preventing angle error drift (e.g., angle error > 0.5 degrees) that may occur during the extrinsic parameter calibration and optimization process.

[0041] In some embodiments, the ratio of the number of valid turning frame images to the number of valid straight-line frame images in the valid frame window is a preset value; and / or, the number of valid frame images in the valid frame window is greater than a lower frame count threshold and less than an upper frame count threshold.

[0042] In an exemplary embodiment, the effective frame window includes both a turning effective frame image and a straight effective frame image, with the ratio of the turning effective frame image to the straight effective frame image being 1:2.

[0043] In the exemplary embodiment, the effective frame window size is set to 6 frames to a maximum of 30 frames. If the window is too small (less than 6 frames), it may result in insufficient historical data, which may not be able to suppress single-frame noise and increase the possibility of parameter estimation jitter. If the window is too large (more than 30 frames), it may consume too much memory and computing resources, and an edge mechanism needs to be triggered to eliminate old data in order to balance computing efficiency. The edge mechanism refers to adding or removing 1-2 frames each time to avoid sudden changes in the calibration optimization parameters caused by sudden window changes.

[0044] In some embodiments, the step of filtering valid frame images from the acquired image sequence after quality filtering for each camera based on a valid frame window includes: performing the following processing for each camera respectively:

[0045] Determine whether the rotation angle of two adjacent frames of the acquired images in the effective frame window is greater than a first angle threshold. If the first angle threshold is met and the number of matching point pairs of the two adjacent frames of the acquired images is greater than a quantity threshold, then the two adjacent frames of the acquired images are determined to be effective frames of the rotation.

[0046] Determine whether the rotation angle of two adjacent acquired images in the effective frame window is less than a second angle threshold. If the conditions are met and the translation distance corresponding to the two adjacent acquired images is greater than a distance threshold, and the number of matching point pairs of the two adjacent acquired images is greater than a quantity threshold, then the two adjacent acquired images are determined to be straight effective frame images; the first angle threshold is greater than the second angle threshold.

[0047] In the exemplary embodiment, all images other than the straight-line valid frame image and the turning valid frame image are considered invalid frame images and are not used in the extrinsic parameter optimization process.

[0048] In an exemplary embodiment, if the rotation angle between two adjacent acquired frames is greater than 5 degrees, it is determined that the vehicle is currently in a steering motion, and the two adjacent acquired frames are steering frame images; it is checked whether the number of matching point pairs between the two acquired frames is greater than 35, and if it is greater, the two acquired frames are determined to be valid steering frame images.

[0049] In an exemplary embodiment, if the rotation angle between two adjacent acquired frames is less than 0.5 degrees and the translation distance is greater than 1 meter, it is considered that the current movement is straight, and the two adjacent acquired frames are straight frame images; check whether the number of matching point pairs between the two adjacent acquired frames is greater than 35, and if it is greater, it is determined to be a valid straight frame image.

[0050] Step 102: Extract effective features based on the effective frame image sequence of each camera; wherein, the effective features include first matching feature point pairs of effective frame images of the same camera at different times, second matching feature point pairs of effective frame images of different cameras at the same time, and lane line discrete feature points.

[0051] In some embodiments, a valid feature point window is used to extract valid features from the valid frame image sequence of each camera. For the first and second matching feature point pairs, reprojection errors exceeding a set threshold (e.g., >3 pixels) are eliminated using the fundamental matrix or epipolar geometry. Simultaneously, feature points in the valid feature point window that were not subsequently successfully tracked are discarded.

[0052] In some embodiments, extracting effective features based on the effective frame image sequence of each camera includes: extracting lane line discrete points from the straight effective frame images of the effective frame image sequence of each camera; and extracting a first matching feature point pair of effective frame images at different times and a second matching feature point pair of effective frame images of different cameras at the same time from each effective frame image of the effective frame image sequence of each camera.

[0053] In some embodiments, each feature point in the effective features satisfies the following: the matching error of the feature point in two consecutive effective frame images is less than a preset error; the motion estimation deviation of the feature point is within the deviation range; and the predicted position change of the feature point is within a preset range.

[0054] In the exemplary embodiment, the matching error of the feature points in the valid features between two consecutive valid frames is <2.5 pixels; the deviation between the feature points and the system's estimated motion model is within 2 sigma, and the spatial position change is <5%. Feature points that do not simultaneously meet both of the above conditions are invalid.

[0055] Step 103: Optimize the extrinsic parameters of each camera based on the first matching feature point pair, the second matching feature point pair, and the discrete feature points of the lane lines corresponding to the effective frame image sequence.

[0056] In some embodiments, optimizing the extrinsic parameters of each camera based on the first matching feature point pair, the second matching feature point pair, and the discrete lane line feature points corresponding to the effective frame image sequence includes: performing the following loss iteration process for each first matching feature point pair, each second matching feature point pair, and each discrete lane line feature point in the effective frame image sequence:

[0057] Based on the discrete feature points of the lane lines and the physical world constraints of the lane lines, a first error term generated by the extrinsic parameters of the corresponding camera is determined; based on the first matching feature point pairs and the epipolar geometry constraints, a second error term generated by the extrinsic parameters of the corresponding camera is determined; based on the second matching feature point pairs and the epipolar geometry constraints, a third error term generated by the extrinsic parameters of the corresponding camera is determined; loss iteration is performed based on the first error term, the second error term, and the third error term and their respective weights.

[0058] The extrinsic parameters of the camera are optimized for the next loss iteration process until the loss converges.

[0059] During the loss iteration process, feature points that can be continuously tracked among the effective features are continuously calibrated, that is, the corresponding error terms are continued to be calculated in the next loss iteration process; feature points that cannot be continuously tracked are discarded.

[0060] The initial value of the camera's extrinsic parameters can be the previous calibration value or a coarse calibration value calibrated using any extrinsic parameter calibration method; this is called the coarse calibration value.

[0061] For example, nonlinear least squares (LM) can be used for loss iteration.

[0062] In the example embodiment, the first error term includes the geometric error term of discrete feature points of lane lines in a single effective frame image captured by a single camera, referred to as the single-camera single-frame geometric error term; and the geometric error term of discrete feature points of lane lines in effective frame images captured by multiple cameras at the same time, referred to as the multi-camera simultaneous multi-frame geometric error term.

[0063] The physical world constraints for lane lines include that lane lines are parallel, perpendicular, and equidistant in the BEV diagram.

[0064] Based on the physical constraints of lane lines, the lane lines in the BEV image should be parallel, perpendicular, and equidistant. For a single effective frame image (BEV image) from a front-view (or rear-view) camera, the lane lines are fitted from discrete points. The absolute value of the difference in the x-coordinates of the intersection points of adjacent lane lines and the horizontal line is calculated. Based on this absolute value, the geometric error term for a single camera and single frame is determined. The horizontal line is a horizontal reference line in the image or BEV coordinate system, typically a straight line parallel to the horizontal axis (x-axis) of the image, used to intersect the lane lines to obtain the x-coordinates of the intersection points.

[0065] Based on the physical constraints of lane lines, the consistency of lane lines across multiple perspectives should be satisfied, and the seam at the same physical point should be zero. For any two valid frame images (BEV images) acquired at the same time from the front-view, rear-view, left-view, and right-view cameras, the lane lines are obtained by fitting discrete feature points of the lane lines. The x-coordinates of the intersection points of the same lane lines on the same horizontal line in different valid frame images are obtained, and the absolute value of the difference between the x-coordinates of the two intersection points is calculated. Based on the absolute value of the difference, the geometric error term of multiple frames at the same time of multiple cameras is determined.

[0066] Understandably, an excessively large or small pitch angle in the camera's extrinsic parameters can cause parallel lane lines to appear inward or outward; an incorrect yaw angle will cause the lane lines to tilt to one side, forming a parallelogram shape in the world coordinate system; and unequal roll angles will result in unequal widths for the left and right lanes. The first error term above reflects the accuracy of the camera's extrinsic parameter calibration based on this principle.

[0067] In an exemplary embodiment, the epipolar geometric constraint requires that the product of the normal vector of a plane and any vector on that plane should be close to 0. The second error term optimizes the camera extrinsic parameters by minimizing the projection error of the first matching feature point pair. The epipolar geometric constraint restricts the search range of the first matching point pair to the epipolar line. The second error term quantifies the geometric consistency of the matching point pair. If the first matching feature point pair satisfies the epipolar constraint, the second error term is essentially zero; otherwise, the error increases.

[0068] Using the image acquisition times j and k corresponding to each feature point in the first matching feature point pair, the rotation and translation of the vehicle body between the two acquisition times are obtained. The rotation and translation of the vehicle body are superimposed on the camera extrinsic parameters. Using epipolar geometry constraints, the error function of the second error term of the first matching feature point pair is constructed.

[0069] In the exemplary embodiment, the third error term optimizes the camera extrinsic parameters by minimizing the projection error of the second matching feature point pair. After superimposing the extrinsic parameters of the two cameras, an error function for the third error term of the second matching feature point pair is constructed using epipolar geometry constraints.

[0070] In some embodiments, the method further includes: taking at least one of the effective frame images as a target effective frame image group, determining a first quality evaluation value of the first matching feature point pair, a second quality evaluation value of the second matching feature point pair, and a third quality evaluation value of the lane line discrete feature point in the target effective frame image group; and adjusting the weights of the first matching feature point, the second matching feature point pair, and the lane line discrete feature point based on the first quality evaluation value, the second quality evaluation value, and the third quality evaluation value.

[0071] The first quality evaluation value indicates the quality of the first matching feature point pair, the second quality evaluation value indicates the quality of the second matching feature point pair, and the third quality evaluation value indicates the quality of the discrete feature points of the lane line. The quality evaluation indicators include the number of similar feature points or feature point pairs. For example, the more feature points or feature point pairs there are, the higher the quality, and the higher the corresponding quality evaluation value.

[0072] In some embodiments, adjusting the weights of the first matching feature point, the second matching feature point pair, and the lane line discrete feature point based on the first quality evaluation value, the second quality evaluation value, and the third quality evaluation value includes:

[0073] When the third quality evaluation value is higher than the first quality evaluation value and the second quality evaluation value, the first weight learning factor is increased; or, when the third quality evaluation value is lower than the first quality evaluation value and the second quality evaluation value, the first weight learning factor is decreased; or, when the third quality evaluation value is lower than the quality threshold, the first weight learning factor is set to 0; the first weight learning factor is within the range of [0,1].

[0074] When the first quality evaluation value is higher than the second quality evaluation value and the third quality evaluation value, the second weight learning factor is increased; or, when the first quality evaluation value is lower than the second quality evaluation value and the third quality evaluation value, the second weight learning factor is decreased; or, when the first quality evaluation value is lower than the quality threshold, the second weight learning factor is set to 0; the second weight learning factor is within the range of [0,1].

[0075] Based on the first weight learning factor, a first weight value corresponding to the discrete feature points of the lane line is determined; a first intermediate value obtained by subtracting the first weight learning factor from 1 is determined, and a second weight value corresponding to the first matching feature point pair is determined based on the product of the second weight learning factor and the first intermediate value; a second intermediate value obtained by subtracting the second weight learning factor from 1 is determined, and a third weight value corresponding to the second matching feature point pair is determined based on the product of the first intermediate value and the second intermediate value; the sum of the first weight value, the second weight value, and the third weight value is equal to 1.

[0076] In an exemplary embodiment, loss iteration is performed based on the first error term, the second error term, and the third error term, along with their respective weights, including:

[0077] The loss function is determined as follows: the first loss term is obtained by multiplying the first error term by the first weight value; the second loss term is obtained by multiplying the second error term by the second weight value; and the third loss term is obtained by multiplying the third error term by the third weight value. The sum of the first, second, and third loss terms is used as the loss function. The loss function is expressed as:

[0078]

[0079] Where α represents the first weight learning factor, β represents the second weight learning factor, and their values ​​range from [0,1]; E Lane E represents the first error term; Point_1 E represents the second error term. point_2 This indicates the third error term.

[0080] When substituting a single feature point, E Lane E Point_1 and E point_2 Individual losses are represented by E, which is accumulated over the feature points of the entire frame image. Lane E Point_1 and E point_2 When , is the sample loss or sample cost function. It can be the sum of individual losses of all valid features across the entire frame, followed by optimization of camera extrinsic parameters and minimization of the sample loss. Alternatively, it can be the sum of individual losses of all valid features across multiple frames, followed by optimization of camera extrinsic parameters and minimization of the loss.

[0081] The loss term included in the loss function will vary depending on the values ​​of the first and second weight learning factors, as shown in Table 1.

[0082] Table 1

[0083]

[0084] In this table, when both α and β are 0, the loss function only includes the loss terms corresponding to the two frames taken simultaneously by the two cameras. This represents the third loss term for the second matching feature point pairs in the effective frame images from different cameras at the same time, i.e., (1-β)*(1-α)*E. point_2 .

[0085] When α is 1 and β is 0, the loss function includes loss terms for single-camera single-frame and multi-camera multi-frame scenarios, indicating that the loss function only includes the first loss term corresponding to the discrete feature points of the lane line, i.e., α*E. Lane Single-camera single-frame refers to the feature points of a single effective frame image from a single camera, while multi-camera multi-frame refers to the feature points of multiple effective frame images from multiple cameras.

[0086] When α takes values ​​in the range (0,1) and β is 0, the loss function includes both the first and third loss terms, i.e., α*E. Lane +(1-β)*(1-α)*E point_2 Single-camera single-frame refers to the feature points of a single effective frame image from a single camera; multi-camera multi-frame refers to the feature points of multiple effective frame images from multiple cameras; and two-camera simultaneous two-frame refers to the feature points of two effective frame images captured by two cameras at the same time.

[0087] When α is 0 and β is 1, the loss function includes only the second loss term, i.e., β*(1-α)*E. Point_1 A single-camera two-frame representation refers to the feature points of two consecutive valid frames from a single camera.

[0088] When α is 1 and β is 1, the loss function includes only the first loss term, i.e., α*E. Lane Single-camera single-frame refers to the feature points of a single effective frame image from a single camera, while multi-camera multi-frame refers to the feature points of multiple effective frame images from multiple cameras.

[0089] When α takes values ​​in the range (0,1) and β is 1, the loss function includes both the first and second loss terms, i.e., α*E. Lane +β*(1-α)*E Point_1 Single-camera single-frame refers to the feature points of a single effective frame image from a single camera; multi-camera multi-frame refers to the feature points of multiple effective frames from multiple cameras; and single-camera two-frame refers to the feature points of two consecutive effective frames from a single camera.

[0090] When α takes the value 0 and β is in the range (0,1), the loss function includes a second loss term and a third loss term, i.e., β*(1-α)*E. Point_1 +(1-β)*(1-α)*E point_2 ;

[0091] When α is 1 and β is in the range (0,1), the loss function includes only the first loss term, i.e., α*E.Lane Single-camera single-frame refers to the feature points of a single effective frame image from a single camera, while multi-camera multi-frame refers to the feature points of multiple effective frame images from multiple cameras.

[0092] When α takes values ​​within (0,1) and β takes values ​​within (0,1), the loss function includes a first loss term, a second loss term, and a third loss term, i.e., α*E. Lane +β*(1-α)*E Point_1 +(1-β)*(1-α)*E point_2 Single-camera single-frame refers to the feature points of a single effective frame image from a single camera; multi-camera multi-frame refers to the feature points of multiple effective frame images from multiple cameras; single-camera multi-frame refers to the feature points of multiple effective frame images from a single camera; two-camera simultaneous two-frame refers to two effective frame images captured by two cameras at the same time.

[0093] In an exemplary embodiment, the second error term E in the loss function is determined using a single camera and a single frame. Point_1 The formula is expressed as follows:

[0094]

[0095] in, To normalize the matching point pairs between two valid frame images acquired by the same camera at times j and k; Let be the rotation matrix and translation vector of the camera extrinsic parameters; For the rotation and translation of the vehicle body between time j and time k; Ω i This is the preset identity matrix.

[0096] In the exemplary embodiment, the loss function utilizes the third error term E determined by two simultaneous frames captured by two cameras. Point_2 The formula is expressed as follows:

[0097]

[0098] in, To normalize the matching point pairs between two valid frame images acquired simultaneously by two adjacent cameras; Let be the rotation matrix and translation vector of the first camera extrinsic parameter; Let be the rotation matrix and translation vector of the second camera extrinsic parameters.

[0099] In the exemplary embodiment, the geometric error term determined using a single camera and a single frame in the loss function is summed with the geometric error term determined using multiple cameras and multiple frames simultaneously to obtain the first error term E. Lane The formula is expressed as follows:

[0100]

[0101] e i =E S_Lane +E M_Lane ;

[0102]

[0103] E M_Lane =|Fl-Lf|+|Fr-Rf|+|Lb-Bl|+|Rb-Br|;

[0104] Among them, E S_Lane E represents the geometric error term for a single camera and single frame. M_Lane For multi-camera simultaneous multi-frame geometric error terms; (Ui, Uj, Di, Dj) are the x-coordinates of the intersection points of the front-view or rear-view BEV lane lines and horizontal lines; (Fl, Fr) are the x-coordinates of the intersection points of the front-view BEV lane lines and horizontal lines; (Bl, Br) are the x-coordinates of the intersection points of the rear-view BEV lane lines and horizontal lines; (Lf, Lb) are the x-coordinates of the intersection points of the left-view BEV lane lines and horizontal lines; (Rf, Rb) are the x-coordinates of the intersection points of the right-view BEV lane lines and horizontal lines; (N l N r () represents the number of qualified lane lines on the left and right sides of the vehicle. Figure 2 The image shows a schematic diagram illustrating the positions of feature points (Ui, Uj, Di, Dj) in a single frame from a single camera. Figure 3 The diagram shows the positions of feature points (Fl,Fr), (Bl,Br), (Lf,Lb), and (Rf,Rb) across multiple frames captured simultaneously by multiple cameras.

[0105] In one exemplary embodiment, such as Figure 4The diagram illustrates the process of online optimization of vehicle camera extrinsic parameters. It involves online image acquisition, filtering of valid frames, distortion removal and segmentation of the original image to extract feature points, and extraction of effective features through a multi-view sliding window control of the effective feature point window. Effective features are extracted in three parts: the first part consists of first-matching feature point pairs from different valid frames of the same camera at adjacent time points; the second part consists of second-matching feature point pairs from valid frames of different cameras at the same time point where they share a common viewing area; and the third part consists of discrete lane line points. For the first and second matching feature point pairs, individual losses are constructed using epipolar geometry constraints. For the discrete lane line points, the existence of a valid lane line is determined by fitting the discrete lane line points. If no valid lane line exists, new valid frames are filtered; if a valid lane line exists, individual losses are constructed based on the physical world constraints of the lane line. Based on the individual losses of the first, second, and discrete lane line feature point pairs, a sample cost function (loss function) is constructed. The first and second weight learning factors are calculated to optimize the weights of each error term in the loss function. The camera extrinsic parameters are nonlinearly adjusted to optimize the loss function until the residuals converge, and the optimized camera extrinsic parameter values ​​are output.

[0106] The embodiments provided in this disclosure, during the online calibration of vehicle camera extrinsic parameters, extract the effective frame image sequence of each camera, and extract effective features based on the effective frame image sequence of each camera. The effective features include a first matching feature point pair of effective frame images from the same camera at different times, a second matching feature point pair of effective frame images from different cameras at the same time, and discrete lane line feature points. Based on the first matching feature point pair, the second matching feature point pair, and the discrete lane line feature points corresponding to the effective frame image sequence, the extrinsic parameters of each camera are optimized. This allows for the comprehensive use of matching feature point pairs of shared viewing areas from multiple consecutive frames of the same camera, matching feature point pairs of shared viewing areas from different cameras at the same time, and lane line feature points during the process of coarse-to-fine calibration of camera extrinsic parameters. This achieves fine-tuning of camera extrinsic parameters, integrates multiple types of feature points to optimize camera extrinsic parameters, avoids the defects of single-type feature point optimization, and improves the accuracy and robustness of camera extrinsic parameter optimization.

[0107] The camera extrinsic parameter optimization methods disclosed in this embodiment are abundant (reprojection, stitching, polarization, etc.), the loss function has complete joint optimization, the effective frame window covers the complete motion cycle (such as vehicle turning) and the sample number is sufficient to effectively suppress single-frame noise, prevent possible angle drift and parameter estimation jitter, effectively solve the defect of nonlinear optimization methods getting trapped in local optima due to large perturbations in the initial extrinsic parameters, and the generated camera extrinsic parameters have high accuracy and quantifiable geometric errors, thereby improving the performance and safety of intelligent driving technology. The extrinsic parameter calibration parameters are highly robust, do not require manual markers, and do not require pre-setting of common viewpoints, which can overcome possible failures in local road texture loss or motion blur scenes with many surrounding dynamic obstacles, and have good environmental adaptability.

[0108] It is understood that the various method embodiments mentioned above in this disclosure can be combined with each other to form combined embodiments without violating the principle and logic. Due to space limitations, this disclosure will not elaborate further. Those skilled in the art will understand that in the above methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic, and the execution order between steps is not limited to implementation according to step number.

[0109] Exemplary device

[0110] Figure 5 This is a block diagram of a vehicle camera extrinsic parameter optimization device provided in an embodiment of this disclosure. The vehicle camera extrinsic parameter optimization device mainly includes:

[0111] The first processing module 501 is used to acquire the image sequence of each of the at least two cameras during the online calibration of the extrinsic parameters of at least two cameras of the vehicle, and to filter valid frame images from the acquired image sequence of each camera to obtain a valid frame image sequence of each camera.

[0112] The second processing module 502 is used to extract effective features based on the effective frame image sequence of each camera; wherein, the effective features include first matching feature point pairs of effective frame images of the same camera at different times, second matching feature point pairs of effective frame images of different cameras at the same time, and lane line discrete feature points.

[0113] The third processing module 503 is used to optimize the extrinsic parameters of each camera based on the first matching feature point pair, the second matching feature point pair, and the discrete feature points of the lane line corresponding to the effective frame image sequence.

[0114] Exemplary electronic devices

[0115] Figure 6 This is a block diagram of an electronic device provided in an embodiment of the present disclosure.

[0116] This disclosure provides an electronic device comprising: at least one processor 601; at least one memory 602; and one or more I / O interfaces 603 connected between the processor 601 and the memory 602; wherein the memory 602 stores one or more computer programs executable by the at least one processor 601, the one or more computer programs being executed by the at least one processor 601 to enable the at least one processor 601 to execute the above-described vehicle camera extrinsic parameter optimization method.

[0117] The modules in the aforementioned electronic devices can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0118] Exemplary computer program products and storage media

[0119] This disclosure also provides a computer program product, including a computer program that, when run in a processor, implements the above-described vehicle camera extrinsic parameter optimization method.

[0120] The computer program may be stored on a readable storage medium of a computer device or in the cloud; the processor of the computer device reads the computer program from the readable storage medium or in the cloud.

[0121] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically manifested as a computer storage medium; in another optional embodiment, the computer program product is specifically manifested as a software product, such as a software development kit (SDK), etc.

[0122] Those skilled in the art will understand that all or some of the steps, systems, and apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software can be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).

[0123] As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable program instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, it is known to those skilled in the art that communication media typically contain computer-readable program instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0124] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0125] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0126] The computer program product described herein can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0127] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0128] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0129] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0130] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0131] The above description is merely a preferred embodiment of this disclosure and is not intended to limit this disclosure. Any modifications or equivalent substitutions made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for optimizing extrinsic parameters of a vehicle camera, characterized in that, include: During the online calibration of the extrinsic parameters of at least two cameras on a vehicle, the image sequence acquired by each of the at least two cameras is obtained, and valid frame images are filtered from the image sequence acquired by each camera to obtain a valid frame image sequence for each camera. Based on the effective frame image sequence of each camera, effective features are extracted; wherein, the effective features include first matching feature point pairs of effective frame images of the same camera at different times, second matching feature point pairs of effective frame images of different cameras at the same time, and lane line discrete feature points; Based on the first matching feature point pair, the second matching feature point pair, and the discrete feature points of the lane line corresponding to the effective frame image sequence, the extrinsic parameters of each camera are optimized.

2. The method according to claim 1, characterized in that, The step of filtering valid frame images from the acquired image sequence of each camera to obtain a valid frame image sequence for each camera includes: The acquired image sequence of each camera is quality filtered, and valid frame images are selected from the quality-filtered acquired image sequence of each camera based on the valid frame window. The valid frame images include turning valid frame images and straight-line valid frame images. Based on the valid frame images within the valid frame window, the valid frame image sequence for each camera is determined.

3. The method according to claim 2, characterized in that, The ratio of the number of valid turning frame images to the number of valid straight-line frame images in the valid frame window is a preset value; and / or, the number of valid frame images in the valid frame window is greater than the lower limit frame number threshold and less than the upper limit frame number threshold.

4. The method according to claim 2, characterized in that, The step of filtering valid frame images from the acquired image sequence after quality filtering for each camera based on the valid frame window includes: For each camera, perform the following processing: Determine whether the rotation angle of two adjacent frames of the acquired images in the effective frame window is greater than a first angle threshold. If the first angle threshold is met and the number of matching point pairs of the two adjacent frames of the acquired images is greater than a quantity threshold, then the two adjacent frames of the acquired images are determined to be effective frames of the rotation. Determine whether the rotation angle of two adjacent acquired images in the effective frame window is less than a second angle threshold. If the conditions are met and the translation distance corresponding to the two adjacent acquired images is greater than a distance threshold, and the number of matching point pairs of the two adjacent acquired images is greater than a quantity threshold, then the two adjacent acquired images are determined to be straight effective frame images; the first angle threshold is greater than the second angle threshold.

5. The method according to claim 2, characterized in that, The extraction of effective features based on the effective frame image sequence of each camera includes: From the straight valid frame images of the valid frame image sequence of each camera, extract discrete points of the lane lines; From each valid frame image sequence of each camera, extract the first matching feature point pair of valid frame images at different times, and the second matching feature point pair of valid frame images of different cameras at the same time.

6. The method according to claim 5, characterized in that, Each feature point in the effective features satisfies the following: the matching error of the feature point in two consecutive effective frame images is less than a preset error; the motion estimation deviation of the feature point is within the deviation range; and the predicted position change of the feature point is within the preset range.

7. The method according to claim 1, characterized in that, The optimization of the extrinsic parameters of each camera based on the first matching feature point pair, the second matching feature point pair, and the discrete feature points of the lane lines corresponding to the effective frame image sequence includes: For each first matching feature point pair, each second matching feature point pair, and each lane line discrete feature point in the effective frame image sequence, the following loss iteration process is performed: Based on the discrete feature points of the lane lines and the physical world constraints of the lane lines, a first error term generated by the extrinsic parameters of the corresponding camera is determined; based on the first matching feature point pairs and the epipolar geometry constraints, a second error term generated by the extrinsic parameters of the corresponding camera is determined; based on the second matching feature point pairs and the epipolar geometry constraints, a third error term generated by the extrinsic parameters of the corresponding camera is determined; loss iteration is performed based on the first error term, the second error term, and the third error term and their respective weights. The extrinsic parameters of the camera are optimized for the next loss iteration process until the loss converges.

8. The method according to claim 7, characterized in that, The method further includes: Take at least one of the effective frame images as the target effective frame image group, and determine the first quality evaluation value of the first matching feature point pair, the second quality evaluation value of the second matching feature point pair, and the third quality evaluation value of the lane line discrete feature point in the target effective frame image group. Based on the first quality evaluation value, the second quality evaluation value, and the third quality evaluation value, the weights of the first matching feature point, the second matching feature point pair, and the discrete feature point of the lane line are adjusted.

9. The method according to claim 8, characterized in that, The step of adjusting the weights of the first matching feature point, the second matching feature point pair, and the lane line discrete feature point based on the first quality evaluation value, the second quality evaluation value, and the third quality evaluation value includes: When the third quality evaluation value is higher than the first quality evaluation value and the second quality evaluation value, the first weight learning factor is increased; or, when the third quality evaluation value is lower than the first quality evaluation value and the second quality evaluation value, the first weight learning factor is decreased; or, when the third quality evaluation value is lower than the quality threshold, the first weight learning factor is set to 0; the first weight learning factor is within the range of [0,1]. When the first quality evaluation value is higher than the second quality evaluation value and the third quality evaluation value, the second weight learning factor is increased; or, when the first quality evaluation value is lower than the second quality evaluation value and the third quality evaluation value, the second weight learning factor is decreased; or, when the first quality evaluation value is lower than the quality threshold, the second weight learning factor is set to 0; the second weight learning factor is within the range of [0,1]. Based on the first weight learning factor, a first weight value corresponding to the discrete feature points of the lane line is determined; a first intermediate value obtained by subtracting the first weight learning factor from 1 is determined, and a second weight value corresponding to the first matching feature point pair is determined based on the product of the second weight learning factor and the first intermediate value; a second intermediate value obtained by subtracting the second weight learning factor from 1 is determined, and a third weight value corresponding to the second matching feature point pair is determined based on the product of the first intermediate value and the second intermediate value; the sum of the first weight value, the second weight value, and the third weight value is equal to 1.

10. A vehicle, characterized in that, The vehicle is configured to perform the vehicle camera extrinsic parameter optimization method as described in any one of claims 1-9.