A Method and System for Self-Calibrating the Extrinsic Parameters of a Surround-View Fisheye Camera
By establishing the mapping relationship between the camera and BEV image, extracting and back-projecting texture points, and adopting adaptive threshold binarization and offset random search strategies, the problem of insufficient external parameter calibration accuracy of the circumferential fisheye camera is solved, and high-precision and real-time calibration of the autonomous driving system is achieved.
Patent Information
- Application Number
- CN202311023017.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-14
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2043-08-14
AI Technical Summary
The prior art is difficult to effectively calibrate the external parameters of the circumferential fisheye camera in a short period of time, resulting in insufficient calibration accuracy of the autonomous driving system, affecting the algorithm results and accuracy.
By establishing the mapping relationship between the camera source image and the BEV image, multiple camera images are projected to the same BEV space, texture points of the BEV viewing angle are extracted and back-projected to the camera perspective viewing angle, the adaptive threshold binarization method is used to optimize the photometric difference, and the offset random search strategy is used to optimize the position.
The position relationship of multiple cameras in the surround view system is realized in a short time, and a seamless 360-degree BEV image is generated, which improves the calibration accuracy and real-timeness of the autonomous driving system.
Smart Images

Figure CN117036502B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of autonomous driving, and in particular, to a method and system for self-calibrating the extrinsic parameters of a surround-view fisheye camera. Background Art
[0002] Currently, autonomous driving is one of the hot topics of concern in the industrial and academic fields. The autonomous driving system relies on various sensors to perceive the environment in order to achieve downstream tasks such as decision-making, planning, and control. Before using the sensors, it is necessary to calibrate their internal and external parameters first, which lays a solid foundation for subsequent mapping, positioning, perception, and control. Calibration is the core part and prerequisite for the stable operation of the autonomous driving system. The calibration accuracy will affect the upper limit accuracy of the sensor's use, and ultimately affect the implementation effect of each function of autonomous driving.
[0003] The surround-view system of an autonomous driving vehicle is an important part of the advanced driver assistance system. The surround-view system usually consists of four fisheye cameras distributed around the front, rear, left, and right of the vehicle body. The field of view (Fov) of the fisheye camera is usually greater than 180 degrees, providing a wider field of view compared to ordinary cameras, and providing richer semantic information for downstream tasks of autonomous driving. The surround-view system composed of fisheye cameras is widely used in scenarios such as automatic parking, parking space detection, and lane line detection. In such application scenarios, the calibration accuracy of the camera will directly affect the results and accuracy of the algorithm.
[0004] Traditional multi-camera online calibration algorithms are usually for ordinary cameras without fisheye distortion, and mostly based on feature point extraction and matching algorithms or algorithms using lane line geometric constraints for calibration. Such methods are difficult to apply to fisheye cameras because fisheye images have strong distortion. For feature point-based methods, even after distortion removal processing, it will still have a great impact on the extraction and matching accuracy of feature points. In addition, the extraction of lane lines and the use of multi-view geometric constraints related to lane lines will also have a greater impact on the algorithm accuracy. On the other hand, whether it is a feature point-based or lane line-based method, such algorithms have strong assumptions about the scene and require strong texture and semantic information. In addition, other multi-camera calibration algorithms such as odometer-based methods require a long time to achieve algorithm convergence. Therefore, there is currently a lack of a surround-view fisheye camera calibration scheme that combines effectiveness and real-time performance. Summary of the Invention
[0005] The embodiments of the present application provide a method and system for self-calibrating the extrinsic parameters of a surround-view fisheye camera, which can calibrate the pose relationship of multiple cameras in the surround-view system and generate a seamless 360-degree BEV image in a very short time for multi-scenario applications of autonomous / assisted driving.
[0006] To solve the above technical problems, in a first aspect, an embodiment of the present application provides a method for self-calibrating the extrinsic parameters of a surround-view fisheye camera, including the following steps: First, establish a mapping relationship between the camera source image and the BEV image; the BEV is the bird's-eye view perspective; then, based on the mapping relationship, project multiple camera images onto the same BEV space to obtain BEV images respectively corresponding to the camera images one by one; and extract the region of interest based on the BEV image corresponding to the camera image; Next, based on the BEV image corresponding to the camera image, extract the texture points from the BEV perspective, and back-project the texture points from the BEV perspective to convert the BEV perspective into the camera perspective to obtain the texture points from the camera perspective; Finally, use the adaptive threshold binary method to process the BEV image and the camera image respectively, and perform gray-scale extraction to optimize the photometric difference between the texture points from the BEV perspective and the texture points from the camera perspective.
[0007] In some exemplary embodiments, the mapping relationship is shown in formula (1):
[0008] (1)
[0009] Where, represents the texture points extracted from the BEV perspective, represents the depth of the texture points back-projected into the camera space C j ; represents the pose from the BEV camera to the camera C j ; , respectively represent the intrinsic parameters of the camera and the BEV camera.
[0010] In some exemplary embodiments, based on the BEV image corresponding to the camera image, extracting the texture points from the BEV perspective includes: defining the image gray-scale gradient; when the image gray-scale gradient meets the preset conditions, using the road surface semantic segmentation mask to extract the texture points from the BEV perspective; the image gray-scale gradient is shown in formula (2):
[0011] (2)
[0012] Where, represents the gray-scale value at the gray-scale image (x, y); the preset conditions are: =2; meets formula (3):
[0013] , (3)
[0014] Among them, mask is the road surface semantic segmentation mask; if the pixel point with coordinates (x, y) meets the preset condition, the pixel point with coordinates (x, y) is a texture point.
[0015] In some exemplary embodiments, the adaptive threshold binaryzation method is used to process the BEV image and the camera image respectively, and gray-scale extraction is performed. The expressions are shown in formulas (4) and (5) respectively:
[0016] (4)
[0017] (5)
[0018] Among them, adb ( ) represents the adaptive threshold binaryzation operation on the image, and i and j respectively represent the adjacent camera indices in the multi-camera system.
[0019] In some exemplary embodiments, the random search method is used to optimize the photometric difference between the texture points in the BEV view and the texture points in the camera view.
[0020] In some exemplary embodiments, the random search method includes: in each round of random search, by calculating the photometric loss and determining whether the sum of the photometric losses of adjacent cameras is less than the photometric loss calculated by the optimal pose in the current round; if so, update the optimal pose in the current round; if not, do nothing.
[0021] In some exemplary embodiments, in each round of random search, calculating the photometric loss includes: in the first round of random search, calculating the photometric loss through the initial optimization pose; the initial optimization pose is obtained by multiplying the pose correction by the initial pose in the current round; in each round of random search in the second and third round optimization stages, calculating the photometric loss through the temporary optimization pose; the temporary optimization pose is obtained by multiplying the pose correction by the optimal pose in the current round.
[0022] In some exemplary embodiments, the expression of the initial optimization pose is shown in formula (6):
[0023] (6)
[0024] Among them, represents the initial optimization pose, represents the pose correction, represents the initial pose in this stage; C i 、C j represent adjacent cameras, k represents the k-th round of random search in the first stage; calculating the photometric loss through the initial optimization pose is shown in formula (7):
[0025] (7)
[0026] Among them, is a texture pixel in the common view of the BEV images of two adjacent cameras.
[0027] In some exemplary embodiments, in the first-round random search, the expression of the photometric loss of the BEV image generated by the optimal pose is shown in formula (8):
[0028] (8)
[0029] If the sum of the photometric losses of adjacent cameras is less than the BEV image generated by the optimal pose in the first round, update the optimal pose in the first round, as shown in formula (9):
[0030] (9)
[0031] In each round of random search in the second-round and third-round optimization phases, if the sum of the photometric losses of adjacent cameras in the current round is less than the BEV image generated by the optimal pose then update the optimal pose ; among them, is the pose generated by the k-th round of random search in the second-round and third-round optimization phases.
[0032] In a second aspect, the embodiments of the present application further provide a self-calibration system for the extrinsic parameters of an omnidirectional fisheye camera, including: a sparse projection module, a BEV view texture extraction module, and a hierarchical optimization module connected in sequence; the sparse projection module is used to establish a mapping relationship between the camera source image and the BEV image; the BEV is the bird's-eye view perspective; based on the mapping relationship, project multiple camera images onto the same BEV space to obtain BEV images respectively corresponding to the camera images one by one; and based on the BEV images corresponding to the camera images, extract the region of interest (subsequently extract the texture in this region); the BEV view texture extraction module is used to extract the texture points in the BEV view according to the BEV images corresponding to the camera images, and back-project the texture points in the BEV view to convert the BEV view to the camera perspective view to obtain the texture points in the camera view; the hierarchical optimization module is used to process the BEV image and the camera image respectively by using the adaptive threshold binary method, and perform gray-scale extraction, and optimize the photometric difference between the texture points in the BEV view and the texture points in the camera view.
[0033] The technical solutions provided by the embodiments of the present application have at least the following advantages:
[0034] An embodiment of the present application provides a method and system for self-calibrating the extrinsic parameters of an omnidirectional fisheye camera. The method includes the following steps: First, establish the mapping relationship between the camera source image and the BEV image; BEV is the bird's-eye view perspective; Then, based on the mapping relationship, project multiple camera images into the same BEV space to obtain BEV images respectively corresponding to the camera images one by one; and extract the region of interest based on the BEV image corresponding to the camera image; Next, extract the texture points from the BEV perspective based on the BEV image corresponding to the camera image, and back-project the texture points from the BEV perspective to convert the BEV perspective into the camera perspective to obtain the texture points from the camera perspective; Finally, use the adaptive threshold binaryzation method to process the BEV image and the camera image respectively, and perform gray-scale extraction to optimize the photometric difference between the texture points from the BEV perspective and the texture points from the camera perspective.
[0035] The present application provides a method and system for self-calibrating the extrinsic parameters of an omnidirectional fisheye camera. By proposing a brand-new multi-camera online calibration algorithm for the omnidirectional system, it can calibrate the pose relationship of the multi-cameras in the omnidirectional system and generate a seamless 360-degree BEV image in a very short time for multi-scenario applications of autonomous / assisted driving. Compared with traditional methods, the present invention innovatively introduces an offset random search method to solve the problem of estimating the pose based on the direct method of minimizing the photometric error, and can still calibrate the effective multi-camera poses of the omnidirectional system when the pose perturbation is too large. In addition, the problem of exposure difference in the multi-camera system is introduced by the adaptive threshold binaryzation result, and the semantic segmentation method from the BEV perspective is introduced to solve the problem of high interference in the optimization problem based on the traditional IPM algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not constitute limitations on the embodiments. Unless otherwise stated, the figures in the drawings do not constitute a proportional limitation.
[0037] Figure 1 It is a schematic flowchart of a method for self-calibrating the extrinsic parameters of an omnidirectional fisheye camera provided by an embodiment of the present application;
[0038] Figure 2 It is a schematic structural diagram of a system for self-calibrating the extrinsic parameters of an omnidirectional fisheye camera provided by an embodiment of the present application;
[0039] Figure 3 It is an example diagram of the conversion from the camera perspective to the BEV perspective provided by an embodiment of the present application;
[0040] Figure 4 It is an example diagram of texture point extraction provided by an embodiment of the present application;
[0041] Figure 5This is an example diagram of a visualization of a semantic segmentation mask of a road segment provided in an embodiment of the present application;
[0042] Figure 6 An example diagram of texture extraction based on adaptive threshold binarization provided in an embodiment of the present application;
[0043] Figure 7 An example diagram of back-projection of BEV texture points back to the camera perspective provided in an embodiment of the present application;
[0044] Figure 8 A schematic diagram of the structure of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0045] As can be seen from the background technology, the prior art lacks a surround-view fisheye camera calibration method that is both effective and real-time.
[0046] The surround view system can achieve real-time seamless 360-degree BEV image generation on board through the precisely calibrated camera intrinsics and extrinsics. On the one hand, it can directly provide the driver with scene information around the vehicle body to reduce accidents caused by visual blind spots. On the other hand, it can provide strong support for downstream perception, decision-making, and planning tasks. The intrinsic parameters of the camera are usually realized through the Zhang Zhengyou camera calibration algorithm. The intrinsic calibration of the camera is usually completed on the production line and remains constant. Usually, recalibration is not required. However, the extrinsic parameters of the camera often change due to collisions or over time, so the camera needs to be recalibrated from time to time.
[0047] The existing surround-view fisheye camera calibration scheme adopts a sparse direct method, which uses the IPM algorithm to convert the perspective image to BEV (bird's eye view), extract the common view area mask of adjacent cameras in BEV, and use the common view area mask and the BEV image of a camera to extract texture points. The texture points are back-projected back to the image planes of the camera and the adjacent cameras, and the corresponding pixel coordinates in different camera image planes corresponding to the same BEV texture point can be obtained. The pixel photometric error is calculated, and the photometric errors of all texture points are summed and used as optimization items. The camera pose is optimized using a nonlinear optimization method. This type of algorithm based on the sparse direct method to solve the surround-view fisheye camera calibration can avoid the impact of image distortion on the calibration results, and can be both real-time and effective.
[0048] The biggest disadvantage of the existing surround-view fisheye camera calibration technology based on sparse direct method and nonlinear optimization is that the pose interference that can be solved is very limited. Practical verification has shown that when the pose rotation error is greater than 0.3 degrees, it is difficult for the algorithm to converge and obtain a seamless 360-degree BEV image.
[0049] When solving the photometric error optimization problem using non - linear optimization, it is easy to fall into local optimal values, thus unable to recover the optimal camera pose. To solve this technical problem, the present invention abandons the traditional non - linear optimization method and innovatively proposes a coarse - to - fine offset random search strategy based on the concurrent programming mode, realizing the optimization and recovery of the best pose even when there are large disturbances in the pose. The algorithm has both real - time performance and high precision, and its accuracy and robustness are tested in a large number of simulation environments and real scenarios. The rotation perturbation accuracy can reach 0.01 degrees, and it can solve pose rotation perturbations within 10 degrees.
[0050] To solve the above - mentioned technical problems, the embodiments of the present application provide a method for self - calibrating the extrinsic parameters of an omnidirectional fisheye camera, including the following steps: First, establish the mapping relationship between the camera source image and the BEV image; BEV is the bird's - eye view perspective. Then, based on the mapping relationship, project multiple camera images onto the same BEV space to obtain BEV images respectively corresponding to the camera images one - to - one; and extract the region of interest based on the BEV image corresponding to the camera image. Next, based on the BEV image corresponding to the camera image, extract the texture points in the BEV perspective, and back - project the texture points in the BEV perspective to convert the BEV perspective to the camera perspective to obtain the texture points in the camera perspective. Finally, use the adaptive threshold binaryzation method to process the BEV image and the camera image respectively, and perform gray - scale extraction to optimize the photometric difference between the texture points in the BEV perspective and the texture points in the camera perspective. The embodiments of the present application provide a method and system for self - calibrating the extrinsic parameters of an omnidirectional fisheye camera, which can calibrate the pose relationship of multiple cameras in the omnidirectional system and generate a seamless 360 - degree BEV image in a very short time for multi - scenario applications of autonomous / assisted driving.
[0051] The following will elaborate on each embodiment of the present application with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in each embodiment of the present application, many technical details are proposed to help readers better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented.
[0052] See Figure 1 , the embodiments of the present application provide a method for self - calibrating the extrinsic parameters of an omnidirectional fisheye camera, including the following steps:
[0053] Step S1: Establish the mapping relationship between the camera source image and the BEV image; BEV is the bird's - eye view perspective.
[0054] Step S2: Based on the mapping relationship, project multiple camera images onto the same BEV space to obtain BEV images respectively corresponding to the camera images one - to - one; and extract the region of interest based on the BEV image corresponding to the camera image.
[0055] Step S3: Based on the BEV image corresponding to the camera image, extract the texture points from the BEV perspective, and back-project the texture points from the BEV perspective to convert the BEV perspective into the camera perspective to obtain the texture points from the camera perspective.
[0056] Step S4: Process the BEV image and the camera image respectively using the adaptive threshold binaryzation method, and perform gray-scale extraction to optimize the photometric difference between the texture points from the BEV perspective and the texture points from the camera perspective.
[0057] In the present invention, the present application proposes a brand-new multi-camera online calibration algorithm for a surround view system. First, based on IPM sparse projection, establish the mapping relationship between the camera source image and the BEV image; based on the mapping relationship, project multiple camera images into the same BEV space to obtain BEV images respectively corresponding to the camera images one by one; and based on the BEV image corresponding to the camera image, extract the region of interest (ROI). Next, extract the texture points from the BEV perspective in this ROI region, and back-project the texture points from the BEV perspective to convert the BEV perspective into the camera perspective to obtain the texture points from the camera perspective; process the BEV image and the camera image respectively using the adaptive threshold binaryzation method, and perform gray-scale extraction, and optimize the photometric difference between the texture points from the BEV perspective and the texture points from the camera perspective, and use the offset random search method to solve the problem of estimating the pose directly based on minimizing the photometric error, and still can calibrate the multi-camera pose of an effective surround view system even in the case of excessive pose perturbation; in addition, the present application introduces the problem of exposure difference of the multi-camera system with the adaptive threshold binaryzation result, and introduces the semantic segmentation method from the BEV perspective to solve the problem of high interference in the optimization problem based on the traditional IPM algorithm.
[0058] See Figure 2The embodiment of the present application also provides a surround fisheye camera extrinsic parameter self-calibration system, including: a sparse projection module 101, a BEV perspective texture extraction module 102 and a hierarchical optimization module 103 connected in sequence; the sparse projection module 101 is used to establish a mapping relationship between a camera source image and a BEV image; BEV is a bird's-eye view perspective; based on the mapping relationship, multiple camera images are projected into the same BEV space to obtain BEV images corresponding to the camera images one by one; and based on the BEV image corresponding to the camera image, an area of interest is extracted; the BEV perspective texture extraction module 102 is used to extract texture points of the BEV perspective according to the BEV image corresponding to the camera image, and back-project the texture points of the BEV perspective, convert the BEV perspective into a camera perspective perspective, and obtain texture points of the camera perspective; the hierarchical optimization module 103 is used to process the BEV image and the camera image respectively using an adaptive threshold binarization method, and perform grayscale extraction to optimize the photometric difference between the texture points of the BEV perspective and the texture points of the camera perspective.
[0059] The surround fisheye camera extrinsic self-calibration system provided in the embodiment of the present application is mainly divided into three modules: an IPM-based sparse projection module, a BEV perspective texture extraction module, and a hierarchical optimization module. Among them, the IPM-based sparse projection module includes two parts: the conversion of the camera perspective view to BEV, and the conversion of the BEV perspective to the camera perspective view.
[0060] (1) Conversion from camera perspective to BEV: This part of the application is based on the IPM algorithm concept. It is assumed that the road surface points are highly consistent with the BEV camera, and a mapping relationship is established between each camera source image and the BEV image, as shown in formula (1):
[0061] (1)
[0062] in, represents the texture points extracted from the BEV perspective, Represents the back projection of the texture point to the camera space C j The depth of Indicates BEV camera to camera C j The posture, , Represent the internal parameters of the camera and the BEV camera respectively. This application uses this mapping relationship to project multiple camera images into the same BEV space to obtain their respective BEV images and extract the region of interest (ROI) by post-processing the BEV images, such as Figure 3 As shown, texture points are subsequently extracted within the region.
[0063] (2)Conversion from BEV perspective to camera perspective: The principle of this part is the same as that of (1). The difference is that in this application, only the texture points extracted in the texture extraction module are back-projected (from BEV to camera perspective). Since the optimization item (photometric error) of this application comes from the photometric difference between the texture points in the BEV perspective and the texture points in the camera perspective, only projecting the texture points will greatly accelerate the algorithm convergence speed.
[0064] In some embodiments, based on the BEV image corresponding to the camera image, extracting texture points in the BEV perspective includes: defining the image gray-scale gradient; when the image gray-scale gradient meets the preset conditions, using the road surface semantic segmentation mask to extract the texture points in the BEV perspective; the image gray-scale gradient is as shown in formula (2):
[0065] (2)
[0066] Wherein, represents the gray-scale value at (x, y) of the gray-scale image; the preset conditions are: = 2; satisfies formula (3):
[0067] , (3)
[0068] Wherein, mask is the road surface semantic segmentation mask. If the pixel point with coordinates (x, y) simultaneously satisfies , and , then this pixel point is a texture point.
[0069] Specifically, in the part of the BEV perspective texture extraction module, this application first defines the image gray-scale gradient, as shown in formula (2). In the algorithm, this application sets = 2. When satisfies formula (3), this application sets this pixel as a texture point, where mask is the road surface semantic segmentation mask. The texture point effect is as shown in Figure 4 , and the visualization result of the road surface mask is as shown in Figure 5 .
[0070] In the part of the hierarchical optimization module, as mentioned above, the optimization item of this application comes from the photometric difference between the texture points in the BEV perspective and the corresponding texture points in the camera perspective. To solve the problem of multi-camera system differences, this application introduces adaptive threshold binarization to process the image and then perform gray-scale extraction.
[0071] In some embodiments, the adaptive threshold binarization method is used to process the BEV image and the camera image respectively, and gray-scale extraction is performed. The expressions are as shown in formula (4) and formula (5) respectively:
[0072] (4)
[0073] (5)
[0074] Among them, adb ( ) represents the adaptive threshold binary operation on the image, and i and j respectively represent the adjacent camera indices in the multi-camera system.
[0075] It should be noted that in the stage of extracting BEV texture in this application, the application does not perform the adaptive threshold binary operation on the BEV image, but only uses the pixel gray value after this operation (the calculation of the pixel gray value is shown in formula (4)) when constructing the error term in the optimization stage; this is because this operation will significantly strengthen the texture point gradient and reduce the texture difference, which is not conducive to the texture clustering effect, as Figure 6 shown. The effect of back-projecting the texture points back to the camera perspective view (after adaptive threshold binary) is as Figure 7 shown.
[0076] In some embodiments, a random search method is used to optimize the photometric difference between the texture points in the BEV view and the texture points in the camera view.
[0077] In some embodiments, the random search method includes: in each round of random search, by calculating the photometric loss and judging whether the sum of the photometric losses of adjacent cameras is less than the photometric loss calculated by the optimal pose of the current round; if so, update the optimal pose of the current round; if not, do nothing.
[0078] In some embodiments, in each round of random search, calculating the photometric loss includes: in the first round of random search, calculating the photometric loss through the initial optimization pose; the initial optimization pose is obtained by multiplying the pose correction by the initial pose of the current round; in each round of random search in the second and third rounds of optimization stages, calculating the photometric loss through the temporary optimization pose; the temporary optimization pose is obtained by multiplying the pose correction by the optimal pose of the current round.
[0079] In some embodiments, the expression of the initial optimization pose is as shown in formula (6):
[0080] (6)
[0081] Among them, represents the initial optimization pose, represents the pose correction, represents the initial pose of this stage; Ci and Cj represent adjacent cameras, k represents the kth round of random search in the first stage; calculating the photometric loss through the initial optimization pose, as shown in formula (7):
[0082] (7)
[0083] Wherein, is the texture pixel in the common view of the BEV images of two adjacent cameras.
[0084] In some exemplary embodiments, in the first round of random search, the expression of the photometric loss of the BEV image generated by the optimal pose is as shown in formula (8):
[0085] (8)
[0086] If the sum of the photometric losses of adjacent cameras is less than the BEV image generated by the optimal pose in the first round, update the optimal pose in the first round, as shown in formula (9):
[0087] (9)
[0088] In each round of random search in the second and third round optimization phases, if the sum of the photometric losses of adjacent cameras in the current round is less than the BEV image generated by the optimal pose then update the optimal pose ; wherein, is the pose generated by the k-th round of random search in the second and third round optimization phases.
[0089] The following introduces the random search method of the present application through a specific embodiment.
[0090] Taking adjacent cameras C i , C j as an example, optimize the pose (Euler angle + displacement) of C i with C j as the reference. Initially, obtain the initial optimized pose , as shown in formula (6).
[0091] In each round, the present application uses to calculate the photometric loss, as shown in formula (7).
[0092] If the sum of the photometric losses of adjacent cameras Ci and Cj is less than the BEV image generated by the optimal pose the algorithm updates , as shown in formula (8) and formula (9).
[0093] In each round of random search in the second and third round y optimization phases, the present application multiplies the pose correction ( ) by the optimal pose ( )(This is different from the first stage. Instead of multiplying it by the initial pose of this stage, the optimal pose is used in this application. As before, if the photometric loss sum of this round is less than the optimal pose For the generated BEV image, this application uses to update .
[0094] It is worth mentioning that this application introduces this random search method, which is similar to the gradient descent method in nonlinear optimization. However, this application can avoid getting stuck in local optima. That is, in the second or third stage of random search, as long as a parameter that can generate a BEV image with a smaller brightness loss than the temporary optimal pose is found, it is immediately updated, and at the same time, random search continues around the updated pose within the range set by the algorithm. Therefore, even if this application is at the optimal value, since the search range of this algorithm is not determined by the gradient and must be larger than the step size around the optimal value calculated by the gradient, it can deviate from the local optimum.
[0095] However, this application does not take this measure in the first random search stage because the correction in the first stage (such as ) is relatively large compared to the other two stages. Taking this measure may risk the algorithm even obtaining a pose far from the true value. Therefore, it should be noted that in the first stage, this application randomly searches for a rough pose within a fixed range, and in the second or third stage of random search, the set search range is randomly searched around the updated pose This optimization mode is called the offsetable random search strategy. In practical engineering, this application borrows the concurrent programming mode to set threads with different search ranges to recover the pose at a faster speed.
[0096] To sum up, on the one hand, this application introduces adaptive threshold binarization to solve the exposure difference problem of the multi-camera system; on the other hand, this application introduces BEV semantic segmentation to solve the influence of the traditional IPM algorithm's unsolvable highly interfering problem on the direct method of minimizing photometric error; in addition, a coarse-to-fine offsetable random search strategy is proposed to solve the optimization problem of the direct method of minimizing photometric error, making up for the drawback that the traditional nonlinear optimization cannot solve the pose recovery when the pose perturbation is too large.
[0097] Compared with the prior art, the advantages of the present invention are as follows: The method provided by this application can still optimize and recover the pose of the multi-camera system when the pose perturbation is too large. Moreover, this application introduces adaptive threshold binarization to solve the exposure difference problem of the multi-camera system, and introduces a BEV semantic segmentation module to filter out texture points with high interference, enabling the pose optimization to reach a higher accuracy.
[0098] Through simulations in multiple scenarios (simulation scenarios and real scenarios), experiments have proven that the algorithm of this application has high precision, robustness, and real-time performance. Compared with the prior art, the method provided by this application can recover a more accurate pose under greater perturbations.
[0099] Refer to Figure 8 Another embodiment of this application provides an electronic device, including: at least one processor 110; and a memory 111 communicatively connected to the at least one processor; wherein, the memory 111 stores instructions executable by the at least one processor 110, and when the instructions are executed by the at least one processor 110, the at least one processor 110 is enabled to execute any of the above method embodiments.
[0100] Among them, the memory 111 and the processor 110 are connected by a bus. The bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors 110 and the memory 111 together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver may be one component or multiple components, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor 110 is transmitted over the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor 110.
[0101] The processor 110 is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory 111 can be used to store data used by the processor 110 when executing operations.
[0102] Another embodiment of this application relates to a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the above method embodiments are implemented.
[0103] That is, those skilled in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the above methods of each embodiment of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0104] With the above technical solutions, the embodiments of the present application provide a method and system for self-calibrating the extrinsic parameters of an omnidirectional fisheye camera. The method includes the following steps: First, establish a mapping relationship between the camera source image and the BEV image; the BEV is the bird's-eye view perspective; then, based on the mapping relationship, project multiple camera images into the same BEV space to obtain BEV images respectively corresponding to the camera images one by one; and extract the region of interest based on the BEV image corresponding to the camera image; Next, extract the texture points from the BEV perspective based on the BEV image corresponding to the camera image, and back-project the texture points from the BEV perspective to convert the BEV perspective into the camera perspective to obtain the texture points from the camera perspective; Finally, use the adaptive threshold binary method to process the BEV image and the camera image respectively, and perform gray-scale extraction to optimize the photometric difference between the texture points from the BEV perspective and the texture points from the camera perspective.
[0105] The present application provides a method and system for self-calibrating the extrinsic parameters of an omnidirectional fisheye camera. By proposing a brand-new online calibration algorithm for multiple cameras in an omnidirectional system, it can calibrate the pose relationship of multiple cameras in the omnidirectional system and generate a seamless 360-degree BEV image in a very short time for multi-scenario applications in autonomous / assisted driving. Compared with traditional methods, the present invention innovatively introduces an offset random search strategy to solve the problem of estimating the pose directly based on minimizing the photometric error, and can still calibrate the effective multi-camera poses of the omnidirectional system even when the pose perturbation is too large. In addition, the problem of exposure difference in the multi-camera system is introduced by the adaptive threshold binary result, and the semantic segmentation method from the BEV perspective is introduced to solve the problem of high interference in the optimization of the traditional IPM algorithm.
[0106] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application. In actual applications, various changes can be made in form and details without departing from the spirit and scope of the present application. Any person skilled in the art can make their own changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be subject to the scope defined by the claims.
Claims
1. A method for self-calibrating the external parameters of an omnidirectional fisheye camera, characterized in that, Including the following steps: Establish the mapping relationship between the camera source image and the BEV image; the BEV is the bird's-eye view perspective; Based on the mapping relationship, project multiple camera images into the same BEV space to obtain BEV images respectively corresponding to the camera images one by one; and extract the region of interest based on the BEV image corresponding to the camera image; Based on the BEV image corresponding to the camera image, extract the texture points from the BEV perspective, and back-project the texture points from the BEV perspective to convert the BEV perspective to the camera perspective to obtain the texture points from the camera perspective; Process the BEV image and the camera image respectively using the adaptive threshold binaryzation method, and perform gray-scale extraction to optimize the photometric difference between the texture points from the BEV perspective and the texture points from the camera perspective; Use the random search method to optimize the photometric difference between the texture points from the BEV perspective and the texture points from the camera perspective; The random search method includes: in each round of random search, calculate the photometric loss, and judge whether the sum of the photometric losses of adjacent cameras is less than the photometric loss calculated by the optimal pose in the current round; if so, update the optimal pose in the current round; if not, do not process; In each round of random search, calculating the photometric loss includes: In the first round of random search, calculate the photometric loss through the initial optimized pose; the initial optimized pose is obtained by multiplying the pose correction by the initial pose in the current round; in each round of random search in the second and third round optimization stages, calculate the photometric loss through the temporary optimized pose; the temporary optimized pose is obtained by multiplying the pose correction by the optimal pose in the current round.
2. The method for self-calibrating the external parameters of an omnidirectional fisheye camera according to claim 1, characterized in that, The mapping relationship is shown in formula (1): (1) Among them, represents the texture points extracted from the BEV perspective, represents the back-projection of the texture points into the camera space C j depth, represents the pose from the BEV camera to the camera C j pose, and represent the internal parameters of the camera and the BEV camera respectively.
3. The method for self-calibrating the external parameters of an omnidirectional fisheye camera according to claim 1, characterized in that, Based on the BEV image corresponding to the camera image, extracting the texture points from the BEV perspective includes: Define the image gray-scale gradient; When the image gray-scale gradient meets the preset conditions, use the road surface semantic segmentation mask to extract the texture points from the BEV perspective; The image gray-scale gradient is shown in formula (2): (2) Among them, represents the gray value at the gray image (x, y); The preset condition is as follows: = 2; Satisfy formula (3): , (3) where mask is the road surface semantic segmentation mask; If the pixel point with coordinates (x, y) meets the preset conditions, the pixel point with coordinates (x, y) is a texture point.
4. The method for self-calibrating the external parameters of an omnidirectional fisheye camera according to claim 1, characterized in that, Process the BEV image and the camera image respectively using the adaptive threshold binaryzation method, and perform gray-scale extraction. The expressions are shown in formula (4) and formula (5) respectively: (4) (5) Among them, adb ( ) represents that the image undergoes an adaptive threshold binaryzation operation, and i and j respectively represent the adjacent camera indices in the multi-camera system.
5. The method for self-calibrating the external parameters of an omnidirectional fisheye camera according to claim 1, characterized in that, The expression of the initial optimized pose is shown in formula (6): (6) Among them, represents the initial optimized pose, represents pose correction, represents the initial pose of this stage; Ci and Cj represent adjacent cameras, and k represents the k-th round of random search in the first stage; the photometric loss is calculated through the initial optimized pose, as shown in formula (7): (7) Among them, represents texture pixels in the common view of the BEV images of two adjacent cameras.
6. The method for self-calibrating the external parameters of an omnidirectional fisheye camera according to claim 5, characterized in that,In the first round of random search, the expression of the photometric loss of the BEV image generated by the optimal pose is shown in formula (8): (8) If the sum of the photometric losses of adjacent cameras is less than the BEV image generated by the optimal pose in the first round, update the optimal pose in the first round, as shown in formula (9): (9) In each round of random search in the second-round and third-round optimization phases, if the sum of photometric losses of adjacent cameras in the current round is less than the optimal pose of the generated BEV image, then update the optimal pose ; where is the pose generated in the k-th round of random search in the second-round and third-round optimization phases.
7. An external parameter self-calibration system for a surround-view fisheye camera, characterized in that, Including: A sequentially connected sparse projection module, a BEV perspective texture extraction module, and a hierarchical optimization module; where, The sparse projection module is used to establish the mapping relationship between the camera source image and the BEV image; the BEV is the bird's-eye view perspective; based on the mapping relationship, project multiple camera images into the same BEV space to obtain BEV images corresponding to the camera images one by one; and extract the region of interest based on the BEV image corresponding to the camera image. The BEV perspective texture extraction module is used to extract the texture points of the BEV perspective according to the BEV image corresponding to the camera image, and back-project the texture points of the BEV perspective to convert the BEV perspective into the camera perspective to obtain the texture points of the camera perspective. The hierarchical optimization module is used to process the BEV image and the camera image respectively by using the adaptive threshold binary method, and perform gray-scale extraction to optimize the photometric difference between the texture points of the BEV perspective and the texture points of the camera perspective. The random search method is used to optimize the photometric difference between the texture points of the BEV perspective and the texture points of the camera perspective. The random search method includes: in each round of random search, calculate the photometric loss, and judge whether the sum of the photometric losses of adjacent cameras is less than the photometric loss calculated by the optimal pose of the current round; if so, update the optimal pose of the current round; if not, do nothing. In each round of random search, calculating the photometric loss includes: In the first round of random search, calculate the photometric loss through the initial optimization pose; the initial optimization pose is obtained by multiplying the pose correction by the initial pose of the current round; in each round of random search in the second and third round optimization stages, calculate the photometric loss through the temporary optimization pose; the temporary optimization pose is obtained by multiplying the pose correction by the optimal pose of the current round.
Citation Information
Patent Citations
Relative pose estimation method based on BEV perception, neural network and training method thereof
CN116452654A
Automatic calibration system based on visual guidance
WO2022120567A1