Non-overlapping view field binocular camera external parameter calibration method and system
By obtaining three-dimensional coordinates with a total station and optimizing the initial values using the LM algorithm, combined with the Transformer network and bundle adjustment method, the accuracy and generalization problems of binocular camera external parameter calibration in complex scenes are solved, achieving high-precision external parameter calibration.
Patent Information
- Application Number
- CN202510665453.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-10-10
AI Technical Summary
Existing binocular camera external parameter calibration methods have problems such as high deployment cost, insufficient robustness, limited accuracy, inability to update in real time, and weak generalization ability in complex scenes. In particular, they are unable to meet the high-precision calibration requirements in dynamic scenes.
A total station is used to obtain the three-dimensional coordinates of the control points. The LM algorithm is used to optimize the initial values of the rotation and translation matrices. A Transformer network is designed for cross-modal feature association. A weighted multi-objective loss function including projection error, feature similarity and geometric constraints is constructed, and the extrinsic parameter matrix is optimized through the bundle adjustment method.
It achieves high-precision external parameter calibration under non-overlapping fields of view, overcomes the traditional method's reliance on field of view overlap, solves the scale drift problem in deep learning, supports high-precision calibration of complex scenes, and significantly improves calibration accuracy and scene generalization capabilities.
Smart Images

Figure CN120765756A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a method and system for calibrating extrinsic parameters of a binocular camera with a non-overlapping field of view. Background Art
[0002] In modern industry, optical measurement technology based on binocular cameras is widely used in the positioning and measurement of targets with unknown depth of field. The calibration accuracy of the binocular camera's external parameters will directly affect whether the entire optical measurement task can be successfully completed.
[0003] Traditional extrinsic parameter calibration methods rely on artificial calibration objects such as checkerboards and calibration plates. They manually arrange control points and collect multiple sets of data for parameter solution. However, they have significant limitations in complex scenarios (such as large-scale outdoor environments or dynamic scenes): First, their deployment cost is high, and the position of the calibration objects needs to be repeatedly adjusted to meet the common visibility requirements, which consumes a lot of manpower and time costs. Second, this method is not robust enough. The calibration objects are easily affected by factors such as changes in ambient lighting and occlusion, resulting in feature point extraction failures, which directly affects the reliability of parameter solution. In addition, the calibration accuracy is limited by the manufacturing errors of the calibration objects themselves, and it is impossible to achieve real-time updates of extrinsic parameters, making it difficult to meet the high-precision calibration requirements in dynamic scenes. These problems seriously restrict the application effectiveness of traditional methods in complex engineering scenarios. Although deep learning-based calibration methods (such as end-to-end Transformers and CNNs) can directly regress extrinsic parameters from images or point clouds, they still have significant drawbacks. First, they are highly sensitive to feature noise. Image blur, low-texture areas, or point cloud sparsity can easily lead to feature matching failures, significantly increasing the cumulative error (for example, projection error can increase by more than three times in low-light scenes). Second, the purely data-driven approach lacks geometric constraints, resulting in a mismatch in the scale decoupling of translation vectors and rotation matrices (for example, the physical units of the predicted translation amount and the rotation angle are inconsistent), causing scale drift. Furthermore, such methods have weak generalization capabilities, and model training relies heavily on specific scene data (such as fixed-focal-length cameras or specific LiDAR models). Parameter recalibration is required when migrating across devices (such as changing lenses) or scenes (such as indoors to outdoors). These issues make it difficult for existing deep learning methods to meet the comprehensive requirements of robustness, accuracy consistency, and generalization capabilities in practical engineering.
[0004] Cross-modal calibration methods provide new solutions to the problems of the above methods through geometric constraints and semantic fusion. However, cross-modal calibration methods (such as LiDAR-camera and radar-camera) have significant deficiencies in the geometric and semantic fusion mechanisms: First, although methods based on simple feature-level fusion (such as feature splicing or weighted fusion) can extract multimodal data, they do not establish geometric association constraints between modalities (for example, the geometric consistency between the normal vector of the unaligned point cloud and the edge features of the image), which makes the extrinsic parameter optimization easily fall into local optimality and has poor adaptability to changes in sensor pose; second, the strategy of post-processing optimization separation (first roughly estimating the extrinsic parameters through deep learning, and then optimizing through traditional algorithms such as ICP and PnP) will introduce error accumulation due to the decoupled calculation of the two-step method (for example, the initial pose deviation causes the iteration to converge to the error extreme value), and end-to-end optimization cannot be achieved in dynamic scenes. These limitations make it difficult for existing methods to deeply integrate high-precision geometric measurements (such as millimeter-level spatial positioning of total stations) and semantic features (such as feature matching guided by semantic segmentation masks), resulting in cross-modal calibration accuracy and scene generalization capabilities being unable to meet the accuracy requirements in complex environments such as high-precision autonomous driving or robotic systems. Summary of the Invention
[0005] Purpose of the invention: The first purpose of the present invention is to provide a method for calibrating the extrinsic parameters of binocular cameras with non-overlapping field of view, which has both accuracy and scene generalization capability. The second purpose is to provide a system for calibrating the extrinsic parameters of binocular cameras with non-overlapping field of view.
[0006] Technical solution: A method for calibrating extrinsic parameters of a binocular camera with a non-overlapping field of view, comprising the following steps:
[0007] S1. Set several control points in the non-overlapping field of view, obtain the three-dimensional coordinates of the control points using a total station, and capture images containing the control points using a binocular camera;
[0008] S2. Use the three-dimensional coordinates of the control points to construct a reprojection error function, and use the LM algorithm to solve the initial values of the binocular camera rotation matrix and the initial values of the translation matrix;
[0009] S3. Design a Transformer network. The image branch of the Transformer encoder uses the ResNet-50 algorithm to extract image features captured by the binocular camera. The point cloud branch of the Transformer encoder uses PointNet++ to extract the 3D coordinate features of the control points. A cross-attention mechanism is used to associate cross-modal features between the image branch and the point cloud branch. A weighted multi-objective loss function is constructed that includes projection error, feature similarity, and geometric constraints. The output of the Transformer encoder is input into the Transformer decoder. After training, matching information between image features and 3D coordinate features is obtained.
[0010] S4. Establish a bundle adjustment model that integrates feature similarity weights, input the initial value of the rotation matrix, the initial value of the translation matrix, and the matching information of the image features and the three-dimensional coordinate features, build a reprojection error model, and output the optimized binocular camera rotation matrix and translation matrix after training.
[0011] Specifically, the reprojection error function is:
[0012]
[0013] Where: E is the projection error function, R is the rotation matrix, T is the translation matrix, π is the camera projection model, P i are the three-dimensional coordinates of the control points, and N is the total number of control points.
[0014] Specifically, in step S2, the iterative convergence conditions of the LM algorithm include: the error decrease is less than 0.01% for three consecutive iterations; the Euclidean distance of the rotation matrix increment is less than 10 -6 The modulus of the translation vector is less than 0.1 mm; the maximum number of iterations is less than or equal to 60.
[0015] Specifically, in step S3, the weighted multi-objective loss function formula is:
[0016] L=λ1L proj +λ2L feat +λ3L geo
[0017] L proj =||π(R·P i +T)-P i ||
[0018] L feat =1-CosSim(F img ,F pts )
[0019] L geo =||det(R)-1||+||R T ·RI||
[0020] Where: L is the weighted multi-objective loss function, λ1, λ2, λ3 are weight parameters, L proj is the projection error term, L feat is the feature similarity term, L geo is the geometric constraint term, π is the camera projection model, R is the rotation matrix, T is the translation matrix, P i is the three-dimensional coordinate of the control point, CosSim is the cosine similarity function, F img is the image feature, F ptsis the coordinate feature, det is the function for calculating the matrix determinant, and I is the identity matrix.
[0021] Specifically, in step S3, the Transformer encoder adopts a multi-layer encoder stacking structure, and each layer of the encoder is configured with a multi-head self-attention mechanism. The Transformer decoder adopts a multi-layer decoder stacking structure, and each layer of the decoder is configured with a multi-head self-attention mechanism.
[0022] Specifically, in step S3, in the cross-attention mechanism, the three-dimensional coordinates of the control point are encoded into a high-dimensional vector through a multi-layer perceptron as a key-value pair input, and the spatial distance constraint matrix is introduced in the attention weight calculation. The formula is:
[0023]
[0024] Where: Attention is the attention weight parameter, Q is the query vector, K is the key vector, V is the value vector, d k is the dimension of the key vector, and M is the Euclidean distance matrix between the three-dimensional coordinates of the control points.
[0025] Specifically, in step S4, the bundle adjustment model is:
[0026]
[0027] Where: ε is the sum of the reprojection errors, π is the camera projection model, R is the rotation matrix, T is the translation matrix, ΔR is the rotation matrix increment, ΔT is the translation matrix increment, P i are the three-dimensional coordinates of the control point, γ is the regularization coefficient, and controls the weight of the rotation matrix increment, ||.|| F represents the Frobenius norm, which calculates the overall change in the rotation matrix increment, ||.|| 2 Represents the square norm of the observation error, that is, the reprojection error.
[0028] Specifically, in step S4, the LM algorithm is used to optimize and solve the bundle adjustment model.
[0029] Specifically, the output of the bundle adjustment model satisfies that the root mean square of the reprojection error is less than or equal to 0.5 mm, and the spatial scale drift rate is less than or equal to 0.01%.
[0030] The present invention also provides a non-overlapping field of view binocular camera extrinsic parameter calibration system, comprising:
[0031] Data acquisition module: used to set several control points in the non-overlapping field of view, obtain the three-dimensional coordinates of the control points using the total station, and collect images containing the control points through the binocular camera;
[0032] Initial value calculation module: used to construct the reprojection error function using the three-dimensional coordinates of the control points, and use the LM algorithm to solve the initial values of the binocular camera rotation matrix and the initial values of the translation matrix;
[0033] Network training module: This module is used to design the Transformer encoder. The image branch of the Transformer encoder uses the ResNet-50 algorithm to extract image features captured by the binocular camera. The point cloud branch of the Transformer encoder uses PointNet++ to extract the 3D coordinate features of the control points. A cross-attention mechanism is used to associate cross-modal features between the image and point cloud branches. A weighted multi-objective loss function is constructed that includes projection error, feature similarity, and geometric constraints. The output of the Transformer encoder is input into the Transformer decoder to obtain matching information between image features and 3D coordinate features.
[0034] BA optimization module: used to establish a bundle adjustment model that integrates feature similarity weights. It inputs the initial values of the rotation matrix, the initial values of the translation matrix, and the matching information of image features and three-dimensional coordinate features. After training, it outputs the optimized binocular camera rotation matrix and translation matrix.
[0035] Beneficial effects: Compared with the existing technology, the significant effects of the present invention are: the present invention adopts a total station to obtain high-precision three-dimensional coordinates of the control points, and uses the LM algorithm to optimize the projection error of the control points, constraining the initial value error of the extrinsic parameters to ±0.5mm, and then constructs a Transformer network including a point cloud branch and an image branch, realizes cross-modal feature association through the cross-attention mechanism, and designs a multi-objective loss function that integrates projection error, feature similarity and geometric constraints to jointly optimize the extrinsic parameter matrix and network parameters of the binocular camera. The present invention overcomes the dependence of traditional calibration methods on field of view overlap, and solves the scale drift problem that exists when using deep learning for extrinsic parameter calibration, supports calibration of complex scenes with low field of view overlap, and has a significant improvement in accuracy compared with existing calibration methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a flow chart of the method of the present invention.
[0037] Figure 2 Schematic diagram of control points and camera positions of the present invention.
[0038] Figure 3 This is the Transformer network flow chart of the present invention.
[0039] Figure 4 It is a flow chart of the optimization solution of the bundle adjustment model of the present invention.
[0040] Figure 5It is a schematic diagram comparing the deflection of the camera before and after correction of the present invention. DETAILED DESCRIPTION
[0041] The present invention will be further described below with reference to the accompanying drawings.
[0042] Example 1
[0043] See also Figure 1 As shown, this embodiment provides a method for calibrating extrinsic parameters of a binocular camera with a non-overlapping field of view, comprising the following steps:
[0044] S1. Set N control points (N≥6) in the non-overlapping field of view, and use the total station to obtain the three-dimensional coordinates of the control points {P i |i=1,2,...N}, and collect images containing control points {I j |j=1,2,...M}. PTP (Precision Time Protocol) network protocol is used to synchronize the clocks of the two cameras. The master control unit distributes a 10MHz global clock signal (synchronization error <100ns) through the FPGA.
[0045] Please refer to Figure 2 As shown in the figure, it is a schematic diagram of the device in a specific scene. There is no overlapping field of view between the first camera 1 and the second camera 2. There are 6 control points set in the scene, evenly distributed on both sides of the binocular camera. The first control point 3, the second control point 4 and the third control point 5 are on the side of the first camera 1, and the fourth control point 6, the fifth control point 7 and the sixth control point 8 are on the side of the second camera 2.
[0046] S2. Use the three-dimensional coordinates of the control points to construct the reprojection error function, and use the LM algorithm to solve the initial values of the binocular camera rotation matrix and the initial values of the translation matrix.
[0047] The reprojection error function is:
[0048]
[0049] Where: E is the projection error function, R is the rotation matrix, T is the translation matrix, π is the camera projection model, P i are the three-dimensional coordinates of the control points, and N is the total number of control points.
[0050] The present invention uses the Levenberg-Marquardt algorithm (LM algorithm) to optimize the initial value R0 of the rotation matrix and the initial value T0 of the translation matrix. The algorithm iteration termination condition needs to comprehensively consider the speed and accuracy of convergence. In this embodiment, the comprehensive judgment is based on the following three criteria:
[0051] (1) Error convergence threshold: monitor the rate of change of the objective function. When the error decreases by less than 0.01% in three consecutive iterations, convergence is determined to avoid over-fitting of the model due to excessive optimization.
[0052] (2) Parameter update threshold: The Euclidean distance of the rotation matrix increment ΔR is less than 10 -6 The iteration is terminated when the modulus of the translation vector is less than 0.1 mm to ensure that the geometric or physical meaning of the parameter adjustment meets the actual requirements.
[0053] (3) Iteration threshold: In dynamic scenarios or under noise interference, the maximum number of iterations is limited to 60, and the iteration is forcibly terminated to prevent the algorithm from falling into invalid oscillation of the local optimal solution.
[0054] The above three criteria complement each other and can ensure the computational efficiency of the algorithm while maintaining the stability and reliability of the results in complex optimization problems.
[0055] In this embodiment, an adaptive weight factor is introduced into the LM algorithm:
[0056] w i =exp(-||P i -π(R·P i +T)|| / σ)
[0057] Where σ is an adaptive adjustment factor that controls the rate at which the weight decays: a larger σ means a more gradual decay of the weight as the error increases, resulting in a higher tolerance for larger errors. The adaptive weight factor constrains the error range to ±0.5 mm.
[0058] S3. Design a Transformer encoder. The image branch of the Transformer encoder uses the ResNet-50 algorithm to extract the image features F collected by the binocular camera. img ∈R^{H×W×C}, the point cloud branch of the Transformer encoder uses PointNet++ to extract the three-dimensional coordinate features F of the control points pts ∈R^{N×D}, use the cross-attention mechanism to associate cross-modal features between the image branch and the point cloud branch: Attention(Q=F img ,K=V=F pts ), and construct a weighted multi-objective loss function including projection error, feature similarity and geometric constraints, input the output of the Transformer encoder into the Transformer decoder to obtain the matching information of image features and three-dimensional coordinate features.
[0059] The weighted multi-objective loss function formula is:
[0060] L=λ1Lproj +λ2L feat +λ3L geo
[0061] L proj =||π(R·P i +T)-P i ||
[0062] L feat =1-CosSim(F img ,F pts )
[0063] L geo =||det(R)-1||+||R T ·RI||
[0064] Where: L is the weighted multi-objective loss function, λ1, λ2, λ3 are weight parameters, L proj is the projection error term, L feat is the feature similarity term, L geo is the geometric constraint term, π is the camera projection model, R is the rotation matrix, T is the translation matrix, P i is the three-dimensional coordinate of the control point, CosSim is the cosine similarity function, F img is the image feature, F pts is the coordinate feature, det is the function for calculating the matrix determinant, and I is the identity matrix.
[0065] In the cross-attention mechanism, the three-dimensional coordinates of the control points are encoded into high-dimensional vectors through a multi-layer perceptron (MLP) as key-value pair inputs. The spatial distance constraint matrix is introduced in the attention weight calculation, and the formula is:
[0066]
[0067] Where: Attention is the attention weight parameter, Q is the query vector, K is the key vector, V is the value vector, d k is the dimension of the key vector, and M is the Euclidean distance matrix between the three-dimensional coordinates of the control points.
[0068] In the present invention, the Transformer network parameter configuration designed for the calibration task needs to take into account both computational efficiency and feature expression capabilities, and the following optimization design is performed based on the traditional Transformer network.
[0069] See also Figure 3As shown in the figure, in terms of network structure, both the Transformer encoder and the Transformer decoder adopt a multi-layer stacked structure, and a multi-head attention mechanism is configured in each layer, connected to the fully connected feedforward layer. Residual connections and layer normalization are set between the multi-head attention mechanism and the fully connected feedforward layer. Taking the Transformer encoder as an example, in this embodiment, the Transformer encoder adopts a 4-layer encoder stacking structure, each layer of the encoder is configured with an 8-head attention mechanism, and the hidden layer dimension is fixed to 512. Through multi-level attention interaction, the fusion association of multimodal features (images, point clouds, IMU data, etc.) is fully captured, while avoiding the computational redundancy caused by overly deep networks.
[0070] In addition, the Transformer network designed in this invention also introduces a position encoding mechanism. To enhance the perception of spatiotemporal consistency, it expands the three-dimensional spatial coordinate (x, y, z) encoding and timestamp encoding on the basis of traditional sequence position encoding, enabling the network to explicitly model the dynamic association of sensor data in physical space and time series dimensions.
[0071] Finally, the training strategy is optimized with an initial learning rate of 10 -4 The cosine annealing scheduling strategy balances convergence speed and stability, combined with a gradient update frequency of batch size = 32, and a weighted geometric loss function with a rotation error weight of 0.7 and a translation error weight of 0.3 is designed to prioritize the optimization accuracy of the rotation parameters, thereby adapting to the asymmetric sensitivity characteristics of the rigid transformation parameters in the calibration task.
[0072] S4. Establish a bundle adjustment (BA optimization) model that integrates feature similarity weights, input the initial value of the rotation matrix, the initial value of the translation matrix, and the matching information of the image features and three-dimensional coordinate features, build a reprojection error model, and output the optimized binocular camera rotation matrix and translation matrix after training.
[0073] The bundle adjustment model is:
[0074]
[0075] Where: ε is the sum of the reprojection errors, π is the camera projection model, R is the rotation matrix, T is the translation matrix, ΔR is the rotation matrix increment, ΔT is the translation matrix increment, P i are the three-dimensional coordinates of the control point, γ is the regularization coefficient, and controls the weight of the rotation matrix increment, ||.|| F represents the Frobenius norm, which calculates the overall change in the rotation matrix increment, ||.|| 2 Represents the square norm of the observation error, that is, the reprojection error.
[0076] By fusing the feature similarity weights output by the Transformer network, the external parameters are weighted optimized.
[0077] See also Figure 4 As shown, in this embodiment, the bundle adjustment is optimized based on the LM algorithm:
[0078] First, the intrinsic parameters of the binocular camera, the initial values of the camera extrinsic parameters obtained in step S2, and the matching information of the image features and three-dimensional coordinate features obtained in step S3 are used as input data. Then, a reprojection error model is constructed, and the optimization parameters of the LM algorithm are set, including the learning rate, the number of iterations, and the convergence condition. Then, iterative optimization is performed until the algorithm converges and the output of the bundle adjustment model satisfies the requirements that the root mean square of the reprojection error is less than or equal to 0.5 mm and the spatial scale drift rate is less than or equal to 0.01%. The optimized camera extrinsic parameters are output, and after post-processing, outliers are eliminated and the data is smoothed to complete the calculation, and the final binocular camera rotation matrix and translation matrix are output.
[0079] To verify the optimization effect of the present invention, the above method is compared with the existing optimization method on the public dataset KITTI. The experimental results are shown in Table 1 below.
[0080] Table 1
[0081]
[0082] As shown in Table 1, compared to the traditional ICP+PnP method and the LCCNet model based entirely on machine learning, the calibration method provided by this invention achieves significant improvements in both translation and rotation errors. Specifically, the reprojection error of binocular vision extrinsic calibration with non-overlapping fields of view is significantly reduced. Compared to traditional methods, the time consumption is significantly reduced, and compared to methods based entirely on machine learning, the increased time consumption is acceptable while significantly improving accuracy.
[0083] Please refer to Figure 5 As shown in the figure, the above method is used to calibrate the extrinsic parameters of the binocular camera. At the same time, a stable single camera is used as a reference to obtain the deflection curves of a series of image sequences. It can be clearly seen that the deflection curves of the calibrated binocular camera are almost completely consistent with those of the stable single camera. At the same time, there is a significant improvement compared to the uncalibrated binocular camera, which proves that the calibration method of the present invention can be effectively applied in real-world scenarios.
[0084] Example 2
[0085] This embodiment provides an extrinsic calibration system for binocular cameras with non-overlapping fields of view, including the following modules:
[0086] Data acquisition module: used to set several control points in the non-overlapping field of view, obtain the three-dimensional coordinates of the control points using the total station, and collect images containing the control points through the binocular camera;
[0087] Initial value calculation module: used to construct the reprojection error function using the three-dimensional coordinates of the control points, and use the LM algorithm to solve the initial values of the binocular camera rotation matrix and the initial values of the translation matrix;
[0088] Network training module: This module is used to design the Transformer encoder. The image branch of the Transformer encoder uses the ResNet-50 algorithm to extract image features captured by the binocular camera. The point cloud branch of the Transformer encoder uses PointNet++ to extract the 3D coordinate features of the control points. A cross-attention mechanism is used to associate cross-modal features between the image and point cloud branches. A weighted multi-objective loss function is constructed that includes projection error, feature similarity, and geometric constraints. The output of the Transformer encoder is input into the Transformer decoder to obtain matching information between image features and 3D coordinate features.
[0089] BA optimization module: used to establish a bundle adjustment model that integrates feature similarity weights. It inputs the initial values of the rotation matrix, the initial values of the translation matrix, and the matching information of image features and three-dimensional coordinate features. After training, it outputs the optimized binocular camera rotation matrix and translation matrix.
Claims
1. A method for calibrating extrinsic parameters of a binocular camera with non-overlapping field of view, characterized in that: The following steps are involved: S1. Set several control points in the non-overlapping field of view, obtain the three-dimensional coordinates of the control points using a total station, and capture images containing the control points using a binocular camera; S2. Use the three-dimensional coordinates of the control points to construct a reprojection error function, and use the LM algorithm to solve the initial values of the binocular camera rotation matrix and the initial values of the translation matrix; S3. Design a Transformer network. The image branch of the Transformer encoder uses the ResNet-50 algorithm to extract image features captured by the binocular camera. The point cloud branch of the Transformer encoder uses PointNet++ to extract the 3D coordinate features of the control points. A cross-attention mechanism is used to associate cross-modal features between the image branch and the point cloud branch. A weighted multi-objective loss function is constructed that includes projection error, feature similarity, and geometric constraints. The output of the Transformer encoder is input into the Transformer decoder. After training, matching information between image features and 3D coordinate features is obtained. S4. Establish a bundle adjustment model that integrates feature similarity weights, input the initial value of the rotation matrix, the initial value of the translation matrix, and the matching information of the image features and the three-dimensional coordinate features, build a reprojection error model, and output the optimized binocular camera rotation matrix and translation matrix after training.
2. The method for calibrating extrinsic parameters of a binocular camera with a non-overlapping field of view according to claim 1, wherein: The reprojection error function is: Where: E is the projection error function, R is the rotation matrix, T is the translation matrix, π is the camera projection model, P i are the three-dimensional coordinates of the control points, and N is the total number of control points.
3. The method for calibrating extrinsic parameters of a binocular camera with a non-overlapping field of view according to claim 1, wherein: In step S2, the iterative convergence conditions of the LM algorithm include: the error decrease of three consecutive iterations is less than 0.01%; the Euclidean distance of the rotation matrix increment is less than 10 -6 The modulus of the translation vector is less than 0.1 mm; the maximum number of iterations is less than or equal to 60.
4. The method for calibrating extrinsic parameters of a binocular camera with a non-overlapping field of view according to claim 1, wherein: In step S3, the weighted multi-objective loss function formula is: L=λ1L proj +λ2L feat +λ3L geo L proj =||π(R·P i +T)-P i || L feat =1-CosSim(F img ,F pts ) L geo =||det(R)-1||+||R T ·RI|| Where: L is the weighted multi-objective loss function, λ1, λ2, λ3 are weight parameters, L proj is the projection error term, L feat is the feature similarity term, L geo is the geometric constraint term, π is the camera projection model, R is the rotation matrix, T is the translation matrix, P i is the three-dimensional coordinate of the control point, CosSim is the cosine similarity function, F img is the image feature, F pts is the coordinate feature, det is the function for calculating the matrix determinant, and I is the identity matrix.
5. The method for calibrating extrinsic parameters of a binocular camera with a non-overlapping field of view according to claim 1, wherein: In step S3, the Transformer encoder adopts a multi-layer encoder stacking structure, and each layer of encoder is configured with a multi-head self-attention mechanism; the Transformer decoder adopts a multi-layer decoder stacking structure, and each layer of decoder is configured with a multi-head self-attention mechanism.
6. The method for calibrating extrinsic parameters of a binocular camera with a non-overlapping field of view according to claim 1, wherein: In step S3, in the cross attention mechanism, the three-dimensional coordinates of the control point are encoded into a high-dimensional vector through a multi-layer perceptron as a key-value pair input, and the spatial distance constraint matrix is introduced in the attention weight calculation. The formula is: Where: Attention is the attention weight parameter, Q is the query vector, K is the key vector, V is the value vector, d k is the dimension of the key vector, and M is the Euclidean distance matrix between the three-dimensional coordinates of the control points.
7. The method for calibrating extrinsic parameters of a binocular camera with a non-overlapping field of view according to claim 1, wherein: In step S4, the bundle adjustment model is: Where: ε is the sum of the reprojection errors, π is the camera projection model, R is the rotation matrix, T is the translation matrix, ΔR is the rotation matrix increment, ΔT is the translation matrix increment, P i are the three-dimensional coordinates of the control points, γ is the regularization coefficient, ‖.‖ F represents the Frobenius norm, ‖.‖ 2 represents the squared norm of the observation error.
8. The method for calibrating extrinsic parameters of a binocular camera with a non-overlapping field of view according to claim 1, wherein: In step S4, the LM algorithm is used to optimize and solve the bundle adjustment model.
9. The method for calibrating extrinsic parameters of a binocular camera with a non-overlapping field of view according to claim 1, wherein: The output of the bundle adjustment model satisfies that the root mean square of the reprojection error is less than or equal to 0.5 mm, and the spatial scale drift rate is less than or equal to 0.01%.
10. A non-overlapping field of view binocular camera extrinsic calibration system, characterized in that: include: Data acquisition module: used to set several control points in the non-overlapping field of view, obtain the three-dimensional coordinates of the control points using the total station, and collect images containing the control points through the binocular camera; Initial value calculation module: used to construct the reprojection error function using the three-dimensional coordinates of the control points, and use the LM algorithm to solve the initial values of the binocular camera rotation matrix and the initial values of the translation matrix; Network training module: This module is used to design the Transformer encoder. The image branch of the Transformer encoder uses the ResNet-50 algorithm to extract image features captured by the binocular camera. The point cloud branch of the Transformer encoder uses PointNet++ to extract the 3D coordinate features of the control points. A cross-attention mechanism is used to associate cross-modal features between the image and point cloud branches. A weighted multi-objective loss function is constructed that includes projection error, feature similarity, and geometric constraints. The output of the Transformer encoder is input into the Transformer decoder to obtain matching information between image features and 3D coordinate features. BA optimization module: used to establish a bundle adjustment model that integrates feature similarity weights. It inputs the initial values of the rotation matrix, the initial values of the translation matrix, and the matching information of image features and three-dimensional coordinate features. After training, it outputs the optimized binocular camera rotation matrix and translation matrix.
Citation Information
Cited By
Three-dimensional projection deformation measurement and error compensation method based on single camera
CN121564704A
A single-camera-based three-dimensional projection deformation measurement and error compensation method
CN121564704B
Total station self-learning measurement adjustment method and system based on rolling prediction
CN121655575A