Multi-small-ball auxiliary multi-camera calibration method and system based on deep learning

By using a deep learning-based multi-ball-assisted multi-camera calibration method, the problems of incomplete field of view and error accumulation in traditional multi-camera calibration are solved, achieving global calibration and high-precision multi-camera pose estimation, thus improving calibration efficiency and accuracy.

CN120833385APending Publication Date: 2025-10-24EAST CHINA JIAOTONG UNIVERSITY

Patent Information

Application Number
CN202511319055.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Traditional multi-camera calibration methods suffer from incomplete field of view of the calibration object, high rate of unqualified calibration images, slow processing speed of feature points, inability to perform global calibration, and are prone to error accumulation and getting trapped in local optima.

Method used

A deep learning-based multi-ball-assisted multi-camera calibration method is adopted. By acquiring multiple sets of image datasets, the image coordinates of speckle marker points are extracted using a target detection algorithm. The projection matrix is ​​solved by combining multi-view geometric relationships and Euclidean structure restoration algorithm. Nonlinear optimization is performed using bundle adjustment algorithm. Finally, the pose calibration of the multi-camera is achieved by matching with a checkerboard calibration board.

Benefits of technology

It achieves global calibration of multiple cameras, improves calibration accuracy and efficiency, avoids occlusion and environmental limitations, can completely preserve spatial coordinate information, reduces noise and unnecessary details, and improves calibration accuracy and speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833385A_ABST
    Figure CN120833385A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-small-ball auxiliary multi-camera calibration method and system based on deep learning, and the method comprises the steps: obtaining a plurality of groups of image data sets, and extracting the image coordinates of all speckle mark points on a calibration part of each group of images from the image data sets through a target detection algorithm; based on the image coordinates and the geometric information of the calibration component, solving a projection matrix of multiple cameras by using a multi-view geometric relationship and an Euclidean structure recovery algorithm; taking the projection matrix as an initial value, and performing nonlinear optimization on the projection matrix by using a bundle adjustment algorithm to obtain an optimized calibration result; and performing world coordinate system matching on the optimized calibration result and a fixed checkerboard calibration plate to realize multi-camera pose calibration. The method disclosed by the invention is simple and practical, has good robustness, and can quickly and accurately identify the calibration object.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of computer vision, and in particular to a multi-ball assisted multi-camera calibration method based on deep learning and a system thereof. BACKGROUND

[0002] Camera calibration remains a fundamental requirement in image metrology and industrial machine vision systems. The accuracy of calibration directly affects the precision of downstream applications, including 3D reconstruction, defect detection, and dimension measurement, ultimately determining the effectiveness of the operation results. The field of view of a single camera system is limited and often cannot meet the needs of these applications. In cases where the use of mechanical auxiliary equipment such as a turntable is impractical, multi-camera configurations become the preferred solution. Multi-camera systems require full-system calibration to spatially align all cameras - a key factor in achieving accurate 3D object reconstruction. This process requires a precisely designed calibration target. Predefined reference points on the target enable mathematical mapping between 3D spatial coordinates and their corresponding 2D image coordinates.

[0003] Traditional 3D calibration targets, while effective for rigid camera configurations, require pre-measured feature coordinates and customized geometric constraints. These prerequisites limit their ability to adapt to different system architectures, especially in large-scale applications, where manufacturing costs are prohibitively high. Moreover, due to occlusions inherent in the calibration components, it is not possible to simultaneously calibrate all cameras in a multi-camera system. Instead, only two adjacent cameras can be calibrated sequentially, with a final coordinate transformation to unify them under a single coordinate system. This method is time-consuming, and the extrinsic parameters often have cumulative errors. Yang et al. proposed a calibration framework based on a circular projection model, bypassing the traditional central projection assumption. By directly relating geometric features on the calibration target to elliptical projections in the image, this method iteratively optimizes mapping parameters through nonlinear least squares fitting. While effective in close-range static environments, in long-range scenarios, the reduced feature resolution and sub-pixel scale of the projected circles make accurate detection complex. However, due to the limited precision of the calibration object and errors in the calculation process, the accuracy of calibration is limited. Therefore, there is a need for a multi-camera calibration method that can quickly and accurately identify calibration objects and overcome the problem of local optimal solution in pose estimation, thereby completing the matching of the world coordinate system and obtaining the spatial coordinate information of the landmark points. SUMMARY

[0004] To solve the problems of traditional multi-camera calibration, such as incomplete field of view of the calibration object, a large number of unqualified calibration pictures, slow processing of a large number of feature points, inability to globally calibrate, easy error accumulation, and easy to fall into local optimal solution, the present disclosure proposes a multi-ball assisted multi-camera calibration method based on deep learning to solve the above problems.

[0005] According to an aspect of the present disclosure, a deep learning-based multi-ball assisted multi-camera calibration method is provided, comprising: S10, a plurality of image data sets are obtained, wherein the plurality of image data sets are obtained by simultaneously photographing, by a plurality of cameras, a calibration component composed of a plurality of balls with speckle markers on surfaces, the calibration component making irregular rigid body movements in space; S20, a target detection algorithm is used to extract image coordinates of each speckle marker on the calibration component from the image data sets; S30, based on the image coordinates and geometric information of the calibration component, a multi-view geometric relationship and a Euclidean structure recovery algorithm are used to solve projection matrices of the plurality of cameras; S40, the projection matrices are used as initial values, and a bundle adjustment algorithm is used to non-linearly optimize the projection matrices to obtain an optimized calibration result; S50, the optimized calibration result is matched with a fixed checkerboard calibration board in a world coordinate system to realize pose calibration of the plurality of cameras.

[0006] Preferably, a target detection algorithm is used to extract image coordinates of each speckle marker on the calibration component from the image data sets, and the target detection algorithm comprises: detecting a bounding box of all speckle markers in the image data sets; calculating pixel coordinates of a center point of each bounding box and taking the pixel coordinates as image coordinates of the markers.

[0007] Preferably, the multi-view geometric relationship and the Euclidean structure recovery algorithm are used to solve the projection matrices of the plurality of cameras, comprising: based on three-dimensional space constraints between the speckle markers on the calibration component, initial extrinsic parameter matrices of all cameras are solved by a linear equation set; based on the initial extrinsic parameter matrices, projection matrices in a Euclidean space are obtained by non-linear optimization.

[0008] Preferably, the bundle adjustment algorithm is used to non-linearly optimize the projection matrices, comprising: the bundle adjustment algorithm takes the intrinsic parameter matrices, the extrinsic parameter matrices of the cameras, and three-dimensional coordinates of the markers as optimization variables, and takes a sum of squared Euclidean distances between image observation coordinates of the markers and estimated re-projection coordinates as a cost function for non-linear optimization.

[0009] Preferably, the sum of squared Euclidean distances between the image observation coordinates of the markers and the estimated re-projection coordinates is taken as the cost function, and the cost function is expressed as: , wherein, is a transformed three-dimensional point, is a normalization factor, is a three-dimensional point in the original space, is a point in the target space, For the parameter The transformation matrix of the control, n is the number of three-dimensional point pairs.

[0010] Preferably, the bundle adjustment algorithm is used to perform nonlinear optimization on the projection matrix, which can be expressed as: , in, represents the first image captured by the jth camera in the i-th observation landmark points, π represents the function of projecting a 3D point onto a 2D image plane through the camera model and distortion model, is the intrinsic parameter matrix of the camera, represents the distortion parameter, is the rotation matrix of the camera relative to the calibration object, is the corresponding translation vector, represents the number of landmark points that can be observed in the image taken by the j-th camera in the i-th observation.

[0011] Preferably, matching the optimized calibration result with the fixed checkerboard calibration plate in the world coordinate system includes: Use multiple cameras to shoot the checkerboard calibration plate and detect the pixel coordinates of its corner points; Based on the known physical dimensions of the checkerboard calibration plate, the 3D physical coordinates of the corner points are calculated, and the extrinsic parameter matrix of each camera to the checkerboard calibration plate coordinate system is solved by combining the pixel coordinates. The center of the chessboard is selected as the origin of the world coordinate system, and the extrinsic parameter matrices of all cameras are optimized by the least squares method to complete the global pose alignment of multiple cameras.

[0012] According to one aspect of the present disclosure, a multi-ball assisted multi-camera calibration system based on deep learning is provided, comprising: a module for acquiring multiple sets of image data sets, wherein the multiple sets of image data sets are obtained by simultaneously photographing a calibration component consisting of multiple small balls with speckle markers on their surfaces performing irregular rigid body motion in space using multiple cameras; an image coordinate extraction module, which extracts the image coordinates of each speckle marker point on the calibration component of each set of images from the image data set using a target detection algorithm; A multi-camera projection matrix solving module solves the multi-camera projection matrix based on the image coordinates and the geometric information of the calibration components using the multi-view geometric relationship and the Euclidean structure recovery algorithm; A nonlinear optimization module uses the projection matrix as an initial value and utilizes a bundle adjustment algorithm to perform nonlinear optimization on the projection matrix to obtain an optimized calibration result; The pose estimation module matches the optimized calibration results with the fixed checkerboard calibration plate in the world coordinate system to achieve multi-camera pose calibration.

[0013] According to one aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to: execute the above-mentioned deep learning-based multi-ball-assisted multi-camera calibration method.

[0014] According to one aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the multi-ball-assisted multi-camera calibration method based on deep learning is implemented.

[0015] Compared with the prior art, the beneficial effects of the present disclosure are: 1) The present disclosure uses a deep learning method to overcome the limitations of traditional image recognition and reduce the time for image recognition and feature point coordinate acquisition.

[0016] 2) This paper uses a spherical calibration object that does not cause occlusion in the multi-camera working area, avoiding the limitations of traditional calibration methods such as being cumbersome and requiring high environmental conditions. It can achieve global calibration of multiple cameras and can fully preserve spatial coordinate information by matching the world coordinate system with the checkerboard calibration plate, thereby improving calibration accuracy.

[0017] 3) This disclosure can improve the efficiency and accuracy of multi-camera calibration, combining deep learning with DIC speckle analysis on the surface reflectivity of a sphere. This approach reduces noise and unnecessary details while preserving important image structure and texture information.

[0018] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure.

[0019] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.

[0021] Figure 1 The flowchart of the multi-ball assisted multi-camera calibration method based on deep learning is shown; Figure 2A multi-camera calibration principle diagram in the embodiment of the present disclosure is shown. Figure 3 A tree structure diagram of camera pose estimation in the embodiment of the present disclosure is shown. Figure 4 A multi-ball assisted multi-camera calibration system based on deep learning in the embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0022] Various exemplary embodiments, features and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference signs in the drawings represent functionally identical or similar elements. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless specifically indicated.

[0023] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any implementation described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other implementations.

[0024] The term "and / or", merely describes an associated relationship, which means that there can be three relationships, for example, A and / or B, which can mean that A exists alone, A and B exist together, and B exists alone. In addition, the term "at least one" herein means any one of the plurality or any combination of at least two of the plurality, for example, including at least one of A, B, and C, which means including any one or more elements selected from the set consisting of A, B, and C.

[0025] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the specific embodiments below. Those skilled in the art should understand that the present disclosure can also be implemented without certain specific details. In some examples, methods, means, elements and circuits well known to those skilled in the art are not described in detail, in order to highlight the main idea of the present disclosure.

[0026] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0027] Embodiment 1 Based on the above idea, the present application proposes a multi-ball assisted multi-camera calibration method based on deep learning. Figure 1 A flowchart of the multi-ball assisted multi-camera calibration method based on deep learning is shown. The method comprises: S10, obtain a plurality of image data sets, wherein the plurality of image data sets are obtained by simultaneously photographing, by a plurality of cameras, a calibration component composed of a plurality of balls having speckle markers on surfaces, the calibration component performing irregular rigid body motion in space; S20, extract image coordinates of each speckle marker on the calibration component in each group of images from the image data sets by using a target detection algorithm; S30, based on the image coordinates and geometric information of the calibration component, solve projection matrices of the plurality of cameras by using a multi-view geometry relationship and a Euclidean structure recovery algorithm; S40, use the projection matrices as initial values, and perform non-linear optimization on the projection matrices by using a bundle adjustment algorithm to obtain an optimized calibration result; S50, match the optimized calibration result with a fixed checkerboard calibration board in a world coordinate system to realize pose calibration of the plurality of cameras.

[0028] The embodiment of the disclosure provides a multi-ball assisted multi-camera calibration method based on deep learning, and specifically includes the following steps: S10, obtain a plurality of image data sets, wherein the plurality of image data sets are obtained by simultaneously photographing, by a plurality of cameras, a calibration component composed of a plurality of balls having speckle markers on surfaces, the calibration component performing irregular rigid body motion in space.

[0029] The calibration component in the embodiment of the disclosure is composed of three collinear balls, and performs rigid body motion in an effective area of the plurality of cameras. Image data sets of three balls having black speckle markers on surfaces are collected by simultaneously photographing the three balls by the plurality of cameras, the calibration component composed of the balls performs irregular rigid body motion in space, and the motion is tracked by using a DIC technology. For each motion of the calibration component, each camera in the plurality of cameras captures an image, thereby forming a group of images corresponding to a current pose of the calibration component. A plurality of image data sets in different poses are obtained by performing multiple motions of the calibration component. A multi-camera calibration principle diagram is shown in Figure 2 The plurality of cameras (such as C1, C2, C3 and C4) are composed of a plurality of cameras, arranged around a workspace to be calibrated, and ensure that all cameras can cover the motion range of the calibration component, the calibration component performs irregular rigid body motion in space, and ensures that the plurality of cameras can collect a plurality of image data sets of the calibration component from different angles.

[0030] The multi-camera calibration requires simultaneously photographing feature pictures of the calibration component in different placement positions in the field of view thereof, which are used as a training set during deep learning. Spatial position information obtained from the same feature picture and pixel coordinate information in the feature picture are combined to obtain extrinsic information of the plurality of cameras. Intrinsic information of the camera can be obtained by using a traditional Zhang calibration method.

[0031] S20, extracting image coordinates of each speckle marker point on the calibration component in each group of images from the image data set by using a target detection algorithm.

[0032] extracting image coordinates of each speckle marker point on the calibration component in each group of images from the image data set by using a target detection algorithm, the target detection algorithm comprising: detecting a bounding box of all speckle marker points in the image data set; calculating pixel coordinates of a center point of each bounding box and taking the pixel coordinates as image coordinates of the marker point.

[0033] In this embodiment, each group of images corresponds to a specific pose of the calibration component. For a certain marker point on the calibration component in a certain pose, the image coordinates of the marker point are extracted from each image in the group of images corresponding to the pose, and the obtained image coordinates are taken as a group of image point correspondences. Since the calibration component has multiple marker points and the calibration component undergoes multiple motions, the image coordinates of all the marker points and the correspondence relationships between multiple groups of image points can be obtained from multiple groups of image data sets. The multiple groups of image data sets of the calibration component captured by the multiple cameras are input into a computer. For a certain feature point of the calibration component in a certain specific pose, the image coordinates of the feature point are extracted from the corresponding image data set using YOLOV8. These extracted image coordinates constitute the correspondence relationship between image points. Since the calibration component has multiple feature points and undergoes multiple motions, multiple groups of image data sets can provide image coordinates of all the feature points and their correspondence relationships between different image data sets.

[0034] Specifically, the same frame of image between different cameras. Subsequently, an approximate nearest neighbor (ANN) algorithm is used to match the features between the feature points in different image data sets. Each group of image data sets corresponds to a pose of the calibration component. For a certain marker point on the calibration component in a certain pose, the image coordinates of the marker point are extracted from each image in the group of images corresponding to the pose using YOLOV8, and the obtained image coordinates are taken as a group of image point correspondences.

[0035] For a feature point in a feature image, the nearest neighbor feature point and the second nearest neighbor feature point .

[0036] , wherein, is the feature vector of the feature point.

[0037] By matching each feature point with the corresponding feature point identified by its nearest neighbor description, the confidence is set to 0.8 in this embodiment, and the algorithm eliminates mismatched point pairs.

[0038] S30, based on the image coordinates and geometric information of the calibration component, solving the projection matrix of the multi-camera by using multi-view geometry relationship and Euclidean structure recovery algorithm.

[0039] Solving the projection matrix of the multi-camera by using multi-view geometry relationship and Euclidean structure recovery algorithm, comprising: obtaining initial extrinsic matrix of all cameras by solving linear equations according to three-dimensional space constraints between each speckle mark point on the calibration component; and obtaining the projection matrix in Euclidean space by nonlinear optimization based on the initial extrinsic matrix.

[0040] In the embodiment, the image coordinates, the corresponding relationship and the geometric information of the calibration component are recovered by using Euclidean geometry information, and then the projection matrix of the multi-camera is solved.

[0041] S40, taking the projection matrix as an initial value, performing nonlinear optimization on the projection matrix by using a bundle adjustment algorithm to obtain an optimized calibration result.

[0042] The nonlinear optimization on the projection matrix by using the bundle adjustment algorithm comprises: the bundle adjustment algorithm takes the intrinsic matrix, the extrinsic matrix of the camera and the three-dimensional coordinates of the mark point as optimization variables, and performs nonlinear optimization on the cost function of the Euclidean distance square sum of the image observation coordinates of the mark point and the estimated re-projection coordinates.

[0043] In the embodiment, the projection matrix result of the multi-camera solved is taken as an initial value, and the initial value is nonlinearly optimized by using a bundle adjustment (BA) algorithm to obtain a more accurate calibration result. The multi-camera is calibrated at one time. The reason for obtaining the relative camera pose is to connect them to describe the pose of each camera in the setting. If the camera is regarded as a node in the graph, all poses can be calculated relative to one of the cameras given the spanning tree as shown in Figure 3 , and these poses can be used as the initial solution of the bundle adjustment. Since the relative poses have different scales, each relative pose must be normalized before splicing.

[0044] Taking the minimum spanning tree connecting each camera as an example, according to the pose of the camera C1 and the relative poses C1→C2 and C2→C3, the pose of C3 relative to C1→C3 can be found by using the following formula. The tree structure diagram of the camera pose estimation in the embodiment of the disclosure is shown in Figure 3 .

[0045] , , , , where R1, R2 and R3 are rotation matrices of different cameras, t2, t3 are translation matrices of different cameras, is a rotation matrix between camera 1 and camera 2, is a rotation matrix between camera 1 and camera 2, is a translation matrix between camera 1 and camera 2, is a translation matrix between camera 2 and camera 3, s represents the relative scale between two relative poses, which can be calculated as the scale difference between the common triangulated points.

[0046] In this embodiment, the three-dimensional point coordinates in space are , and the projected pixel coordinates are .

[0047] , wherein, is a non-zero scale factor, represents the Lie group symbolic representation of the external parameter [R, T], and K represents the camera intrinsic parameter. Considering the uncertainty of the camera pose and the existence of image noise, the sum of errors is solved by the least square method, and a Lie algebra element is found, so that the square sum of the Euclidean distance between all observed points and the three-dimensional points transformed by the camera intrinsic matrix K and the Lie algebra element is minimized.

[0048] Further, the square sum of the Euclidean distance between the image observation coordinates of the landmark points and the estimated re-projection coordinates is taken as the cost function, and the cost function is represented as: , wherein, is the transformed three-dimensional point, is a normalization factor, is a three-dimensional point in the original space, is a point in the target space, is a transformation matrix controlled by the parameters , and n is the number of three-dimensional point pairs. The error term represents the difference between the predicted position of the three-dimensional point after the camera imaging and the observed position.

[0049] wherein the bundle adjustment algorithm is used to nonlinearly optimize the projection matrix to obtain a more accurate calibration result, and the bundle adjustment algorithm is represented as: , wherein, represents the jth camera in the ith observation, a landmark point, p denotes a function that projects a three-dimensional point onto a two-dimensional image plane through a camera model and a distortion model, is an intrinsic matrix of the camera, denotes distortion parameters, is a rotation matrix of the camera relative to the calibration object, is a corresponding translation vector, is the number of observable landmark points in the image taken by the jth camera in the ith observation.

[0050] S50, match the optimized calibration result with the fixed checkerboard calibration plate in the world coordinate system to realize the pose calibration of the multi-camera.

[0051] Matching the optimized calibration result with the fixed checkerboard calibration plate in the world coordinate system comprises: taking the checkerboard calibration plate by the multi-camera, detecting the pixel coordinates of the corner points thereof; calculating the three-dimensional physical coordinates of the corner points based on the known physical size of the checkerboard calibration plate, and solving the extrinsic matrix of each camera to the checkerboard calibration plate coordinate system in combination with the pixel coordinates; selecting the center of the checkerboard as the origin of the world coordinate system, and optimizing the extrinsic matrix of all cameras by the least square method to complete the global pose alignment of the multi-camera. In this way, not only the complete pose estimation of the multi-camera can be directly obtained, but also the coordinate information of the landmark points in space can be obtained.

[0052] In the embodiment, the obtained calibration data is matched with the world coordinate system to complete the pose estimation of the multi-camera. As the initial solution of the bundle adjustment, all the obtained camera poses are based on the pose of the first camera. In order to convert these poses into the global reference system with correct scale, the rigid transformation (including rotation, scaling and translation) between the triangulation points and their real positions in the world is estimated. Once this transformation is determined, these poses can be projected into the new reference system. In this way, the specific coordinates of the multi-camera and the landmark points in the space coordinate system can be obtained.

[0053] In order to verify the accuracy of the method proposed in the embodiment, it is necessary to define a camera accuracy evaluation standard. The world coordinate system coordinates of the target recognition points obtained in the calibration process of the main camera are used. The calibration parameters obtained by the camera calibration are substituted into the camera imaging model to obtain new pixel coordinates . Then these new coordinates are compared with the pixel coordinates of the target recognition points used for calibration , and the average error between them is calculated:

[0054] .

[0055] The embodiment of the present disclosure proposes a multi-ball assisted multi-camera calibration method based on deep learning. The method can calibrate any number of cameras with a common field of view. The spherical calibration object does not cause occlusion in the multi-camera working area, avoids the cumbersome single camera calibration of traditional calibration methods, and overcomes the problems of low feature image recognition efficiency and easy local optimal solution of camera pose estimation. A multi-camera calibration method with simplicity, practicality, high accuracy and good robustness is designed, and the efficiency and accuracy of multi-camera calibration are improved.

[0056] Embodiment 2 As another aspect of the embodiment of the present disclosure, a multi-ball assisted multi-camera calibration system 100 based on deep learning is also provided, as shown in Figure 4 The system comprises: A plurality of image data set acquisition modules 1 acquire a plurality of image data sets, wherein the plurality of image data sets are obtained by irregular rigid body motion of a calibration component composed of a plurality of balls with speckle markers on the surface in space and photographed by a plurality of cameras at the same time; An image coordinate extraction module 2 extracts the image coordinates of each speckle marker on the calibration component from each group of image data using a target detection algorithm; A multi-camera projection matrix solving module 3 solves the projection matrix of the multi-camera based on the image coordinates and the geometric information of the calibration component using multi-view geometry relationship and Euclidean structure recovery algorithm; A nonlinear optimization module 4 uses bundle adjustment algorithm to perform nonlinear optimization on the projection matrix with the projection matrix as the initial value to obtain the optimized calibration result; A pose estimation module 5 matches the optimized calibration result with a fixed checkerboard calibration board in the world coordinate system to realize the pose calibration of the multi-camera.

[0057] Without contradiction, the above-mentioned modules in the system of the embodiment of the present disclosure can realize any of the embodiments in the above-mentioned method.

[0058] Based on the description of the above-mentioned embodiments, the embodiment of the present disclosure can achieve the following technical effects: 1) The method of the present disclosure using deep learning can overcome the limitations of traditional image recognition and reduce the image recognition and feature point coordinate collection time.

[0059] 2) The present disclosure uses a spherical calibration object that does not cause occlusion in the multi-camera working area, avoids the limitations of traditional calibration methods such as complexity and high environmental requirements, and can realize global calibration of multi-camera. The spatial coordinate information is completely preserved by matching with the checkerboard calibration board in the world coordinate system, and the calibration accuracy is improved.

[0060] 3) The present disclosure can improve the efficiency and accuracy of multi-camera calibration, combining the accuracy level of deep learning and DIC speckle on the reflectivity of the small ball surface. While preserving important structural and texture information of the image, reduce noise and unnecessary details.

[0061] The embodiments of the present disclosure also provide an electronic device, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the above-mentioned deep learning-based multi-small ball assisted multi-camera calibration method. The electronic device can be provided as a terminal, a server or other forms of devices.

[0062] The embodiments of the present disclosure also provide a computer-readable storage medium having stored thereon computer program instructions, which, when executed by a processor, implement the above-mentioned deep learning-based multi-small ball assisted multi-camera calibration method. The computer-readable storage medium can be a non-volatile computer-readable storage medium.

[0063] Those skilled in the art can understand that, in the above-mentioned deep learning-based multi-small ball assisted multi-camera calibration method and system of the specific embodiments, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process, and the specific execution order of each step should be determined by its function and possible internal logic.

[0064] The flowcharts and block diagrams in the drawings show the possible implementation architecture, function and operation of the systems, methods and computer program products according to the embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams can represent a module, a program segment or a part of instructions, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks can also occur in different order from that noted in the drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and sometimes they can be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and the combination of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0065] Having described above several embodiments of the disclosure, any modifications and variations that fall within the scope of the described embodiments are also intended to be within the scope of the disclosure. As will be apparent to those skilled in the art, some modifications and variations to the embodiments described above can be practiced while staying within the scope and spirit of the described embodiments. The foregoing description of the described embodiments has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the described embodiments to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. It is intended that the disclosed embodiments be limited only by the claims.

Claims

1. A multi-sphere assisted multi-camera calibration method based on deep learning, characterized in that, The method comprises the following steps: S10, obtaining a plurality of image data sets, wherein the plurality of image data sets are obtained by simultaneously photographing, by a plurality of cameras, a calibration component composed of a plurality of small balls with speckle markers on the surfaces of the small balls irregularly rigidly moving in space; S20, extracting, by using a target detection algorithm, image coordinates of each speckle marker on the calibration component in each group of images from the image data sets; S30, solving projection matrices of the plurality of cameras by using a multi-view geometric relationship and a Euclidean structure recovery algorithm based on the image coordinates and geometric information of the calibration component; S40, performing nonlinear optimization on the projection matrices by using a bundle adjustment algorithm with the projection matrices as initial values to obtain optimized calibration results; S50, matching the optimized calibration results with a fixed checkerboard calibration board in a world coordinate system to realize pose calibration of the plurality of cameras.

2. The method of claim 1, wherein, The target detection algorithm comprises: detecting a bounding box of each speckle marker in the image data sets; calculating pixel coordinates of the center points of the bounding boxes and taking the pixel coordinates as the image coordinates of the markers.

3. The method of claim 1, wherein, Solving the projection matrices of the plurality of cameras by using the multi-view geometric relationship and the Euclidean structure recovery algorithm comprises: obtaining initial extrinsic matrices of all the cameras by solving linear equations based on three-dimensional space constraints among the speckle markers on the calibration component; obtaining the projection matrices in a Euclidean space by performing nonlinear optimization based on the initial extrinsic matrices.

4. The method of claim 1, wherein, Performing nonlinear optimization on the projection matrices by using the bundle adjustment algorithm comprises: The bundle adjustment algorithm takes the intrinsic matrices and extrinsic matrices of the cameras and three-dimensional coordinates of the markers as optimization variables, and performs nonlinear optimization with the sum of squared Euclidean distances between the image observation coordinates of the markers and the estimated re-projection coordinates as a cost function.

5. The method of claim 4, wherein, The cost function is represented as: , in, is the transformed 3D point, is the normalization factor, is a three-dimensional point in the original space, is a point in the target space, For the parameter The transformation matrix of the control, n is the number of three-dimensional point pairs.

6. The method of claim 1, wherein, The nonlinear optimization on the projection matrices by using the bundle adjustment algorithm is represented as: , wherein, represents the j-th marker point photographed by the j-th camera in the i-th observation, π denotes a function that projects a three-dimensional point onto a two-dimensional image plane through a camera model and a distortion model, is an intrinsic matrix of the camera, denotes a distortion parameter, is a rotation matrix of the camera with respect to the calibration object, is a corresponding translation vector, is the number of observable marker points in the image photographed by the j-th camera in the i-th observation.

7. The method of claim 1, wherein, Matching the optimized calibration results with the fixed checkerboard calibration board in the world coordinate system comprises: photographing the checkerboard calibration board by the plurality of cameras to detect pixel coordinates of the corners of the checkerboard calibration board; calculating three-dimensional physical coordinates of the corners based on the known physical size of the checkerboard calibration board, and combining the pixel coordinates to solve extrinsic matrices of the cameras to the coordinate system of the checkerboard calibration board; selecting the center of the checkerboard as the origin of the world coordinate system, and optimizing the extrinsic matrices of all the cameras by the least square method to complete global pose alignment of the plurality of cameras.

8. A multi-sphere assisted multi-camera calibration system based on deep learning, characterized in that, The method comprises: a plurality of image data set obtaining module, which obtains a plurality of image data sets, wherein the plurality of image data sets are obtained by simultaneously photographing, by a plurality of cameras, a calibration component composed of a plurality of small balls with speckle markers on the surfaces of the small balls irregularly rigidly moving in space; an image coordinate extracting module, which extracts image coordinates of each speckle marker on the calibration component in each group of images from the image data sets by using a target detection algorithm; a plurality of camera projection matrix solving module, which solves projection matrices of the plurality of cameras by using a multi-view geometric relationship and a Euclidean structure recovery algorithm based on the image coordinates and geometric information of the calibration component; a nonlinear optimization module, which performs nonlinear optimization on the projection matrices by using a bundle adjustment algorithm with the projection matrices as initial values to obtain optimized calibration results; a nonlinear optimization module, which performs nonlinear optimization on the projection matrices by using a bundle adjustment algorithm with the projection matrices as initial values to obtain optimized calibration results; The pose estimation module matches the optimized calibration result with the fixed checkerboard calibration board in a world coordinate system, and realizes pose calibration of the multiple cameras.

9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to realize the deep learning-based multi-ball assisted multi-camera calibration method in any one of claims 1 to 7.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to realize the deep learning-based multi-ball assisted multi-camera calibration method in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and apparatus for standardization of multiple camera system

    CN101226638A

  • Method for improving camera calibration accuracy by using multi-plane calibration board

    CN111429532A

  • Calibration method and system of multi-camera system

    CN114399554A

  • Multi-camera calibration method based on spherical calibration block

    CN114581526A

  • Light source calibration method and device based on combined target and related medium

    CN117557656A

Cited By

  • Method and system for back projection of 2D defect to 3D digital model

    CN121414571A

  • Thermopile array and depth camera combined calibration method and system

    CN121725076A