Large-view-field multi-view camera calibration method, system, equipment and medium

By acquiring multi-view target board images with a monocular camera and decoding the target code value, combined with a 3D point coordinate optimization algorithm, the problem of insufficient calibration accuracy of large field-of-view multi-view cameras is solved, realizing high-precision multi-view camera calibration and large-size object measurement.

CN121304802APending Publication Date: 2026-01-09SHANGHAI BAOSTEEL METALLURGICAL CONSTRUCTION CORP
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511400070.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively improve the calibration accuracy of large field-of-view multi-camera systems, especially in scenarios requiring high-precision full-field measurements, such as industrial inspection and motion capture. Insufficient calibration accuracy among multiple cameras can negatively impact measurement results.

Method used

Multi-view target images are acquired using a monocular camera, target feature points are extracted and decoded into target code values, and multi-view camera calibration is performed using a 3D point coordinate optimization algorithm. Nonlinear optimization strategies are used to reduce noise and errors, and accurate coordinate transformation between multi-view cameras is achieved.

Benefits of technology

It improves the accuracy of multi-view camera calibration, reduces noise and errors in 3D reconstruction, and enhances the measurement accuracy of large objects in industrial production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121304802A_ABST
    Figure CN121304802A_ABST
Patent Text Reader

Abstract

The invention provides a large-view-field multi-view camera calibration method, system and device and a medium. The method comprises the following steps: acquiring a multi-view-field target plate image in a large-view-field plane by using a monocular camera; the target plate comprises a plurality of coding targets; obtaining target feature points of the plurality of coding targets based on the target plate image so as to correspondingly obtain target code values based on the target feature points; and obtaining a three-dimensional point coordinate of each target feature point based on each target code value, and carrying out multi-view camera calibration based on the three-dimensional point coordinates. According to the invention, the large-view-field multi-view camera can be calibrated, and the calibration precision is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of image processing technology, and in particular relates to a method, system, device and medium for calibrating a large field of view multi-view camera. Background Technology

[0002] Wide field-of-view multi-camera calibration is a key technology for high-precision 3D measurement of vast spaces through the collaborative work of multiple cameras. Its core lies in utilizing multiple cameras distributed in space (such as in a ring or array arrangement) to simultaneously capture target images covering a large area of ​​the scene, and then using optimization algorithms to uniformly calculate the intrinsic and extrinsic parameters of all cameras. This method effectively solves problems such as the difficulty of single-camera or small field-of-view systems in covering the entire scene and significant edge distortion, and is particularly suitable for scenarios requiring high-precision full-field measurement, such as industrial inspection and motion capture. Therefore, improving the accuracy of wide field-of-view multi-camera calibration is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0003] The purpose of this application is to provide a method, system, device, and medium for calibrating a wide field of view multi-view camera, addressing the technical problem of how to improve the accuracy of wide field of view multi-view camera calibration.

[0004] In a first aspect, this application provides a method for calibrating a large field-of-view multi-view camera, the method comprising:

[0005] A monocular camera is used to acquire multi-view images of a target board within a large field of view; the target board includes several coded targets.

[0006] Based on the target plate image, target feature points of several coded targets are obtained, and target code values ​​are obtained based on each target feature point.

[0007] The three-dimensional point coordinates of each target feature point are obtained based on the target code values, and multi-view camera calibration is performed based on the three-dimensional point coordinates.

[0008] In one implementation of the first aspect, obtaining target feature points of several coded targets based on the target plate image includes:

[0009] The target plate image is preprocessed to obtain multiple processed images;

[0010] The multiple processed images are detected to obtain target feature points of several coded targets.

[0011] In one implementation of the first aspect, obtaining the target code value based on the corresponding target feature points includes:

[0012] Convert each of the target feature points into a binary sequence;

[0013] The binary sequence is subjected to a cyclic shift operation, and the binary string with the smallest lexicographical order among all cyclic shift results is selected as the canonical encoding;

[0014] The standard encoding is converted into a decimal value to serve as the corresponding target code value.

[0015] In one implementation of the first aspect, obtaining the three-dimensional point coordinates of each target feature point based on each target code value includes:

[0016] Based on the target code values, obtain the target feature matching points of each coded target under multiple perspectives;

[0017] The single-view extrinsic parameters of the monocular camera are obtained based on the target feature matching points of any single viewpoint.

[0018] Based on the single set of extrinsic parameters, obtain the three-dimensional point coordinates corresponding to the target feature matching point of the single set of perspectives;

[0019] Based on the single set of extrinsic parameters and the corresponding three-dimensional point coordinates, the extrinsic parameters of the other perspectives are obtained, and the corresponding three-dimensional point coordinates are obtained by combining the target feature matching points of the corresponding perspectives, so as to obtain the three-dimensional point coordinates of each target feature point under multiple perspectives.

[0020] In one implementation of the first aspect, obtaining the three-dimensional point coordinates corresponding to the target feature matching point based on the single set of extrinsic parameters includes:

[0021] Based on the known intrinsic parameters of the monocular camera, the pixel coordinates of the target feature matching points are converted into corresponding normalized coordinates;

[0022] The projection matrix under a single viewpoint is obtained based on the normalized coordinates and the single viewpoint extrinsic parameters.

[0023] Based on the projection matrix under a single perspective, a system of equations is established, and the system of equations is solved to obtain the corresponding three-dimensional point coordinates.

[0024] In one implementation of the first aspect, after obtaining the three-dimensional point coordinates of each target feature point, the method further includes:

[0025] The projection points of the three-dimensional points are obtained based on the coordinates of the three-dimensional points, the intrinsic parameters of the monocular camera, and the extrinsic parameters of the multi-viewpoint.

[0026] The reprojection error of the target feature points is optimized based on the image points of the three-dimensional points in multiple viewpoints and the projection points in the corresponding viewpoints; wherein, in the optimization process, the visibility weight determines whether the three-dimensional points are visible in a viewpoint.

[0027] In one implementation of the first aspect, multi-camera calibration based on the three-dimensional point coordinates includes:

[0028] Within the large field of view plane, one camera is designated as the main camera, and the remaining cameras are designated as sub-cameras;

[0029] Based on the three-dimensional point coordinates, obtain the region coordinate transformation matrix between each region;

[0030] Based on the 3D point coordinates, camera intrinsic parameters, target feature points, image coordinates of the 3D points under the regional camera, and the camera's projection function, obtain the regional camera coordinate transformation matrix between each camera and the corresponding region;

[0031] Based on the region coordinate transformation matrix and the region camera coordinate transformation matrix, the local coordinate system of each sub-camera is transformed to the global coordinate system of the main camera for multi-camera calibration.

[0032] Secondly, this application provides a wide field-of-view multi-view camera calibration system, the system comprising:

[0033] The first acquisition module is used to acquire multi-view images of a target board within a large field of view using a monocular camera; the target board includes several coded targets;

[0034] The second acquisition module acquires target feature points of several coded targets based on the target plate image, and acquires target code values ​​based on each target feature point.

[0035] The calibration module is used to obtain the three-dimensional point coordinates of each target feature point based on each target code value, so as to perform multi-view camera calibration based on the three-dimensional point coordinates.

[0036] Thirdly, this application provides an electronic device, the electronic device comprising: a processor and a memory; the memory for storing a computer program; the processor for executing the computer program stored in the memory, so as to cause the electronic device to perform the above-described method.

[0037] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by an electronic device, implements the above-described method.

[0038] As described above, the wide field-of-view multi-camera calibration method, system, device, and medium of this application have the following beneficial effects: This application obtains accurate three-dimensional point coordinates of target feature points through monocular camera calibration, and further performs multi-camera calibration based on the three-dimensional point coordinates. In the process of obtaining the three-dimensional point coordinates and multi-cameras, this application utilizes a nonlinear optimization strategy to significantly reduce noise and errors in three-dimensional reconstruction, as well as reduce positioning errors between multi-cameras, effectively improving the accuracy of multi-camera calibration. Attached Figure Description

[0039] Figure 1 This is a flowchart illustrating the wide field-of-view multi-view camera calibration method according to an embodiment of this application.

[0040] Figure 2 The diagram shown is a schematic representation of the coded target in an embodiment of this application.

[0041] Figure 3 The diagram shows a flowchart of obtaining target code values ​​based on the corresponding target feature points in this application embodiment.

[0042] Figure 4 The diagram shown is a partial representation of the encoded targets and their corresponding binary sequences and decimal decoded code values ​​in the embodiments of this application.

[0043] Figure 5 The diagram shows a flowchart of obtaining the three-dimensional point coordinates of each target feature point based on the target code value in this embodiment of the application.

[0044] Figure 6 The diagram shown is a schematic of non-overlapping large field-of-view multi-view camera calibration in an embodiment of this application.

[0045] Figure 7 The diagram shown is a structural schematic of a large field-of-view multi-view camera calibration system in an embodiment of this application.

[0046] Figure 8 The diagram shown is a structural schematic of the electronic device of this application in one embodiment. Detailed Implementation

[0047] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, unless otherwise specified, the following embodiments and features in the embodiments can be combined with each other.

[0048] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this application. Therefore, the drawings only show the components related to this application and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0049] Furthermore, the use of terms such as "first" and "second" in this application is for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of the various embodiments can be combined with each other, but only on the basis of being achievable by those skilled in the art. If the combination of technical solutions is contradictory or impossible to implement, such a combination of technical solutions should be considered non-existent and not within the scope of protection claimed in this application.

[0050] Currently, large-scale steel plates are frequently used in industrial production. However, unevenness in the steel plates can cause problems during processing or welding, leading to dimensional deviations or shape defects in the finished products, thus affecting their performance and reliability. Therefore, accurate detection of steel plate flatness is a crucial step in industrial production. Since large-sized steel plates cannot be captured in a single camera field of view, multiple cameras are needed to achieve a wide field of view and simultaneously measure the flatness information of steel plates of various sizes. However, calibration is required between these cameras to unify the coordinates of the three-dimensional information, and the calibration accuracy directly affects the yield rate of industrial products. Therefore, improving the calibration accuracy of multi-view cameras is a technical problem that urgently needs to be solved by those skilled in the art.

[0051] To address at least the aforementioned problems, the following embodiments of this application provide a method for calibrating a large field-of-view multi-view camera. This method can calibrate non-overlapping large field-of-view multi-view cameras, thereby enabling the measurement of large-sized targets that cannot appear in the same camera's field of view. Applying this application can effectively improve the accuracy of multi-view camera calibration and enhance the quality of industrial production.

[0052] The technical solutions in the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0053] Figure 1 This is a flowchart illustrating the wide field-of-view multi-view camera calibration method according to an embodiment of this application. Figure 1 As shown, the method includes steps S1-S3.

[0054] S1. Use a monocular camera to acquire multi-view images of a target board within a large field of view; the target board includes several coded targets.

[0055] In some embodiments, a target plate is provided in a large field-of-view plane, and the target plate includes several different coded targets.

[0056] In some embodiments, the coded target can be a concentric circle coded target with 12-bit binary encoding. This target consists of a centrally located concentric circle as the central positioning point, and an outer concentric coded ring. The specific target structure is as follows: Figure 2 As shown, the concentric coded rings are divided into 12 equal parts, with each division point becoming a coded bit. Starting from any coded bit, the coded strip is coded sequentially in a specific order (clockwise or counterclockwise), with black coded as "0" and white coded as "1," thus forming the corresponding target codes. The target plate contains several targets with different coded numerical information. In this way, during the shooting process, the target's position and orientation in space can be determined by identifying the target's coded information, thereby improving the accuracy of calibration.

[0057] In some embodiments, a monocular camera with precise intrinsic parameter calibration is used to capture images of the target board within a large field of view from different perspectives, thereby acquiring multi-view target board images. Different targets will be captured in the target board images from different perspectives, and some targets may be partially missing in some perspectives. Therefore, it is necessary to use a monocular camera to densely capture target board images from different perspectives to ensure that all complete targets on the target board are captured.

[0058] S2. Based on the target plate image, obtain target feature points of several coded targets, and obtain target code values ​​based on each target feature point.

[0059] In some embodiments, when a monocular camera acquires multiple target board images from different perspectives (i.e., multi-view target board images), target feature points of several coded targets can be extracted based on the target board images. Target feature points are key locations or structures in an image or scene that are salient and distinguishable, used for precise matching, localization, or measurement. In the concentric circle coded target of the above embodiments, the central positioning point and the black or white coded points on the coded positions are all target feature points. The black coded points are not actually explicitly displayed.

[0060] In some embodiments, obtaining target feature points of several coded targets based on the target board image includes: preprocessing the target board image to obtain multiple processed images; and detecting the multiple processed images to obtain target feature points of several coded targets. The image preprocessing of the target board image includes grayscale conversion of the color image, image denoising, grayscale stretching, and binarization. After obtaining the binarized image, a line scanning algorithm is used to extract targets from the image, and the extracted targets, i.e., target feature points, are marked on the original color target board image.

[0061] Furthermore, after acquiring the target feature points, target code values ​​are obtained based on the corresponding target feature points. Each target actually has a unique target code value, thereby enabling rapid identification and localization by utilizing the target code value to correspond to the target. In some embodiments, after a monocular camera captures an image containing a target, the target feature points on the target are decoded using the principle of cyclic shift least binary encoding to extract the unique code value of the target. Figure 3 This is a schematic diagram illustrating the process of obtaining target code values ​​based on the corresponding target feature points in an embodiment of this application. For example... Figure 3 As shown, obtaining the target code value based on each target feature point includes steps S21 to S23.

[0062] S21. Convert each of the target feature points into a binary sequence.

[0063] S22. Perform a cyclic shift operation on the binary sequence to select the binary string with the smallest lexicographical order from all cyclic shift results as the standard encoding.

[0064] S23. Convert the standard code into a decimal value to serve as the corresponding target code value.

[0065] Because the target image is captured from different viewpoints, the same target may rotate under different fields of view. Therefore, to avoid rotation interference and correctly decode the target, the above embodiment uses an encoding method that finds the unique smallest binary value through cyclic shift operations. The core principle is to compare all cyclic shift forms of the binary code and select the smallest value as the standard representation of the code. For example, taking a four-bit binary sequence 1100, all cyclic shift results are 1100, 1001, 0011, and 0110, where 0011 is the smallest value. Therefore, the final representation of this code is 0011. This encoding method is often used to solve rotational symmetry problems (such as ring-coded targets), ensuring that the system can correctly identify its unique code regardless of the target's rotation. Figure 4 The images shown are some of the encoded targets and their corresponding binary sequences and decimal decoded code values ​​in the embodiments of this application.

[0066] S3. Obtain the three-dimensional point coordinates of each target feature point based on the target code value, and perform multi-view camera calibration based on the three-dimensional point coordinates.

[0067] Because the positions of each target on the target board are fixed, the same targets and target feature points appearing in target board images from different viewpoints actually correspond to the real coordinates in a world coordinate system. Therefore, a code value matching strategy is used to associate feature point pairs from different viewpoints, and then the corresponding 3D point coordinates are obtained using the matched feature point pairs. Figure 5 This is a schematic diagram illustrating the process of obtaining the three-dimensional point coordinates of each target feature point based on the target code values ​​in an embodiment of this application. For example... Figure 5 As shown, obtaining the three-dimensional point coordinates of each target feature point based on each target code value includes steps S31 to S34.

[0068] S31. Based on the target code values, obtain the target feature matching points of each coded target under multiple perspectives.

[0069] Once the target code value is extracted and the code value of the same target is consistent under different viewpoints, a local coordinate system is established by combining the geometric features of the target. The code value is used as an index to associate the corresponding targets under different viewpoints. Finally, the algorithm can be optimized to eliminate mismatches and achieve high-precision multi-view target feature point matching to obtain the target feature matching point.

[0070] S32. Obtain the single-view extrinsic parameters of the monocular camera based on the target feature matching points of any single viewpoint.

[0071] In some embodiments, two viewpoints are arbitrarily selected from the multi-view images as a single set of viewpoints, and the extrinsic parameters of the two viewpoints are obtained based on the target feature matching points extracted from the images of the single set of viewpoints. For example, after obtaining the target feature matching points under the two viewpoints through feature matching, the essential matrix between the two different viewpoint images is calculated based on epipolar geometry, and the relative rotation matrix and relative translation matrix of the monocular camera with respect to the target under the two viewpoints are decomposed to obtain the extrinsic parameters of the single set of viewpoints.

[0072] S33. Based on the single set of external parameters, obtain the three-dimensional point coordinates corresponding to the target feature matching point of the single set of external parameters.

[0073] In some embodiments, obtaining the 3D point coordinates corresponding to the target feature matching point of the single viewpoint based on the single viewpoint extrinsic parameters includes: converting the pixel coordinates of the target feature matching point into corresponding normalized coordinates based on the known intrinsic parameters of the monocular camera; obtaining the projection matrix under the single viewpoint based on the normalized coordinates and the single viewpoint extrinsic parameters; solving a system of equations based on the projection matrix under the single viewpoint, and solving the system of equations to obtain the corresponding 3D point coordinates.

[0074] S34. Based on the single set of external parameters and the corresponding three-dimensional point coordinates, obtain the external parameters of the other external parameters, and combine them with the target feature matching points of the corresponding external parameters to obtain the corresponding three-dimensional point coordinates, so as to obtain the three-dimensional point coordinates of each target feature point under multiple perspectives.

[0075] In some embodiments, after obtaining the extrinsic parameters and 3D point coordinates of a single view, new views are gradually added based on the extrinsic parameters and 3D points of these two views, and the extrinsic parameters of the corresponding view are calculated. Then, the corresponding 3D point coordinates are obtained by combining the target feature matching points of the corresponding view. New views are continuously added until the 3D point coordinates of each target feature point under all views are obtained.

[0076] Furthermore, after acquiring the multi-view extrinsic parameters and all three-dimensional point coordinates, this embodiment of the application performs joint optimization on the acquired extrinsic parameters and three-dimensional point coordinates to further improve accuracy. In some embodiments, the projection points of the three-dimensional points are acquired based on the three-dimensional position of the target feature points, the intrinsic parameters of the monocular camera, and the multi-view extrinsic parameters; the reprojection error of the target feature points is optimized based on the image points of the three-dimensional points in the multi-view and the projection points in the corresponding view; wherein, during the optimization process, the visibility weight determines whether the three-dimensional point is visible in a view. The specific process of jointly optimizing the acquired extrinsic parameters and three-dimensional point coordinates can be shown in Equation (1).

[0077]

[0078] Among them, C j Let X be the camera extrinsic parameter for the j-th viewpoint. i Let x be the coordinates of the i-th 3D point. For each camera viewpoint, use the observed image point x... ij (3D point X) i (Image point at the j-th viewpoint) and according to the camera extrinsic C j The camera intrinsic parameter K' and the projection points obtained from the calculation of 3D point coordinates The Levenberg-Marquardt algorithm is used to minimize the sum of squared reprojection errors of all target feature points to obtain more accurate extrinsic parameters and 3D point coordinates. Where w ij Visibility weights are used to indicate the visibility weights of a 3D point X. i Whether it is visible in the j-th viewpoint. This allows for the accurate acquisition of the 3D point cloud X of the calibration object and the coordinates X of each 3D point. i The correspondence between feature points in each viewpoint.

[0079] After multi-view calibration is completed using a monocular camera to obtain 3D point coordinates and multi-view extrinsic parameters, multi-view camera calibration can be performed for a large field of view. The multi-view camera calibration based on the 3D point coordinates includes: determining one camera as the main camera within the large field of view plane, and designating the remaining cameras as sub-cameras; obtaining the region coordinate transformation matrix between each region based on the 3D point coordinates; obtaining the region camera coordinate transformation matrix between each camera and its corresponding region based on the 3D point coordinates, camera intrinsic parameters, target feature points, image coordinates of the 3D points under the region cameras, and the camera's projection function; and transforming the local coordinate system of each sub-camera to the global coordinate system of the main camera based on the region coordinate transformation matrix and the region camera coordinate transformation matrix to perform multi-view camera calibration.

[0080] In some embodiments, a multi-camera system is constructed using four cameras. To minimize costs, the four cameras are installed with non-overlapping fields of view. Camera 1 is designated as the master camera, and the remaining cameras are designated as sub-cameras, with each camera's field of view representing a corresponding region. To achieve coordinate system unification, the local coordinate systems of each sub-camera need to be transformed to the global coordinate system of the master camera. The core of this transformation process lies in solving the transformation matrix of each sub-camera relative to the master camera, i.e., multi-camera calibration. The specific transformation is as follows... Figure 6 As shown.

[0081] The multi-camera localization process described above will be specifically explained using the conversion between the main camera 1 and the sub-camera 4 in the region as an example. If... As the coordinates of a point in region 4 in the region 4 coordinate system, Assuming the coordinates of a point in region 4 within the coordinate system of region 1, then:

[0082]

[0083] in, It is the rotation matrix between the region 1 coordinate system and the camera 1 coordinate system. It is the translation matrix between the region 1 coordinate system and the camera 1 coordinate system. It is the rotation matrix between the region's 4-coordinate system and the camera's 4-coordinate system. It is the translation matrix between the region 4 coordinate system and the camera 4 coordinate system. It is the rotation matrix between the camera 4 coordinate system and the camera 1 coordinate system. It is the translation matrix between the camera 4 coordinate system and the camera 1 coordinate system. It is the rotation matrix between the coordinate system of region 1 and the coordinate system of region 4. This is the translation matrix between the coordinate systems of region 1 and region 4. Then, by combining equations (1) and (2), the transformation matrix between camera 1 and camera 4 can be obtained as shown in equation (3):

[0084]

[0085] in, and The coordinates between the three-dimensional points obtained in steps S1 to S3 can be directly obtained. The only unknowns are... That is, the transformation relationship between each region and the corresponding camera. Taking the transformation relationship between camera 1 and region 1 as an example, the optimized transformation relationship can be obtained by nonlinearly optimizing the minimum reprojection error, as shown in equation (4).

[0086]

[0087] Where i represents the i-th target point, N represents the number of target points, and π(·) represents the function that projects the 3D point onto the 2D image plane through the camera model and distortion model. i1 This represents the coordinates of a 3D point on a 2D image plane. K1 is the intrinsic parameter of camera 1, k1 is the distortion coefficient of camera 1, and X... i1 These are the coordinates of a three-dimensional point.

[0088] The 3D point coordinates acquired by camera 1 are reprojected onto the image plane where camera 1 is located based on the camera model and distortion model. The coordinates of the 2D image plane are obtained, which are the reprojected coordinates. The reprojection error is obtained by subtracting the reprojected coordinates from the 3D point coordinates. The optimized transformation relationship between region 1 and camera 1 is obtained by minimizing the reprojection error. The transformation relationship between region 4 and camera 4 can also be obtained through the same steps. Equation (3) can then be solved to obtain the transformation relationship between camera 1 and camera 4.

[0089] The above process transforms the local coordinate system of the sub-camera into the global coordinate system of the main camera, thus completing the multi-camera calibration.

[0090] The scope of protection of the wide field-of-view multi-view camera calibration method described in this application is not limited to the execution order of the steps listed in this embodiment. Any solution implemented by adding, subtracting, or replacing steps in the prior art based on the principles of this application is included within the scope of protection of this application.

[0091] This application also provides a wide field-of-view multi-view camera calibration system. The wide field-of-view multi-view camera calibration system can implement the wide field-of-view multi-view camera calibration method described in this application. However, the implementation device of the wide field-of-view multi-view camera calibration system described in this application includes, but is not limited to, the structure of the wide field-of-view multi-view camera calibration system listed in this embodiment. All structural modifications and substitutions of the prior art made based on the principles of this application are included within the protection scope of this application.

[0092] like Figure 7As shown, in one embodiment, the wide field-of-view multi-view camera calibration system of this application includes a first acquisition module 41, a second acquisition module 42, and a calibration module 43.

[0093] The first acquisition module 41 is used to acquire multi-view target board images within a large field of view using a monocular camera; the target board includes several coded targets.

[0094] The second acquisition module 42 acquires target feature points of several coded targets based on the target plate image, and acquires target code values ​​based on each target feature point.

[0095] The calibration module 43 is used to obtain the three-dimensional point coordinates of each target feature point based on each target code value, so as to perform multi-view camera calibration based on the three-dimensional point coordinates.

[0096] The structure and principle of the first acquisition module 41, the second acquisition module 42, and the calibration module 43 correspond one-to-one with the steps in the above-mentioned large field-of-view multi-view camera calibration method, so they will not be described again here.

[0097] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, or methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules / units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or units may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of apparatuses or modules or units may be electrical, mechanical, or other forms.

[0098] The modules / units described as separate components may or may not be physically separate. The components shown as modules / units may or may not be physical modules; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules / units can be selected to achieve the objectives of the embodiments of this application, depending on actual needs. For example, the functional modules / units in the various embodiments of this application may be integrated into one processing module, or each module / unit may exist physically separately, or two or more modules / units may be integrated into one module / unit.

[0099] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0100] This application also provides a computer-readable storage medium. Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing a processor. The program can be stored in a computer-readable storage medium, which is a non-transitory medium, such as random access memory, read-only memory, flash memory, hard disk, solid-state drive, magnetic tape, floppy disk, optical disk, and any combination thereof. The storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital video disc (DVD)), or a semiconductor medium (e.g., solid-state drive (SSD)).

[0101] This application also provides an electronic device. The electronic device includes a processor and a memory.

[0102] The memory is used to store computer programs.

[0103] The memory includes various media capable of storing program code, such as ROM, RAM, magnetic disk, USB flash drive, memory card, or optical disk.

[0104] The processor is connected to the memory and is used to execute the computer program stored in the memory so that the electronic device performs the above-described wide field-of-view multi-view camera calibration method.

[0105] Preferably, the processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0106] like Figure 8 As shown, the electronic device of this application is embodied in the form of a general-purpose computing device. The components of the electronic device may include, but are not limited to: one or more processors or processing units 51, memory 52, and bus 53 connecting different system components (including memory 52 and processing unit 51).

[0107] Bus 53 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. Examples of these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.

[0108] Electronic devices typically include a variety of computer-readable media. These media can be any available media that can be accessed by the electronic device, including volatile and non-volatile media, and removable and non-removable media.

[0109] Memory 52 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 521 and / or cache memory 522. The electronic device may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 523 may be used to read and write non-removable, non-volatile magnetic media (…). Figure 8 Not shown; usually referred to as a "hard drive"). Although Figure 8Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 53 via one or more data media interfaces. Memory 52 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of this application.

[0110] A program / utility 524 having a set (at least one) of program modules 5241 may be stored, for example, in memory 52. ​​Such program modules 5241 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 5241 typically perform the functions and / or methods described in the embodiments of this application.

[0111] The electronic device can also communicate with one or more external devices (e.g., keyboard, pointing device, display, etc.), one or more devices that enable a user to interact with the electronic device, and / or any device that enables the electronic device to communicate with one or more other computing devices (e.g., network interface card, modem, etc.). This communication can be performed through input / output (I / O) interface 54. Furthermore, the electronic device can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) through network adapter 55. Figure 7 As shown, network adapter 55 communicates with other modules of the electronic device via bus 53. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0112] This application embodiment may also provide a computer program product comprising one or more computer instructions. When the computer instructions are loaded and executed on a computing device, all or part of the processes or functions described in this application embodiment are generated. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0113] When the computer program product is executed by a computer, the computer performs the method described in the foregoing method embodiments. The computer program product can be a software installation package; when the foregoing method is required, the computer program product can be downloaded and executed on the computer.

[0114] This application provides a method, system, device, and medium for calibrating a large field-of-view multi-view camera, which can accurately calibrate a non-overlapping large field-of-view multi-view camera.

[0115] The descriptions of the processes or structures corresponding to the above figures each have their own emphasis. For parts of a process or structure that are not described in detail, please refer to the relevant descriptions of other processes or structures.

[0116] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.

Claims

1. A method for calibrating a large field-of-view multi-view camera, characterized in that, The method includes: A monocular camera is used to acquire multi-view images of a target board within a large field of view; the target board includes several coded targets. Based on the target plate image, target feature points of several coded targets are obtained, and target code values ​​are obtained based on each target feature point. The three-dimensional point coordinates of each target feature point are obtained based on the target code values, and multi-view camera calibration is performed based on the three-dimensional point coordinates.

2. The method for calibrating a large field-of-view multi-view camera according to claim 1, characterized in that, Obtaining target feature points of several coded targets based on the target plate image includes: The target plate image is preprocessed to obtain multiple processed images; The multiple processed images are detected to obtain target feature points of several coded targets.

3. The method for calibrating a large field-of-view multi-view camera according to claim 1, characterized in that, The target code value is obtained based on the corresponding target feature points, including: Convert each of the target feature points into a binary sequence; The binary sequence is subjected to a cyclic shift operation, and the binary string with the smallest lexicographical order among all cyclic shift results is selected as the canonical encoding; The standard encoding is converted into a decimal value to serve as the corresponding target code value.

4. The method for calibrating a large field-of-view multi-view camera according to claim 1, characterized in that, Obtaining the three-dimensional point coordinates of each target feature point based on the target code values ​​includes: Based on the target code values, obtain the target feature matching points of each coded target under multiple perspectives; The single-view extrinsic parameters of the monocular camera are obtained based on the target feature matching points of any single viewpoint. Based on the single set of extrinsic parameters, obtain the three-dimensional point coordinates corresponding to the target feature matching point of the single set of perspectives; Based on the single set of extrinsic parameters and the corresponding three-dimensional point coordinates, the extrinsic parameters of the other perspectives are obtained, and the corresponding three-dimensional point coordinates are obtained by combining the target feature matching points of the corresponding perspectives, so as to obtain the three-dimensional point coordinates of each target feature point under multiple perspectives.

5. The method for calibrating a large field-of-view multi-view camera according to claim 4, characterized in that, Obtaining the 3D point coordinates corresponding to the target feature matching point based on the single set of extrinsic parameters includes: Based on the known intrinsic parameters of the monocular camera, the pixel coordinates of the target feature matching points are converted into corresponding normalized coordinates; The projection matrix under a single viewpoint is obtained based on the normalized coordinates and the single viewpoint extrinsic parameters. Based on the projection matrix under a single perspective, a system of equations is established, and the system of equations is solved to obtain the corresponding three-dimensional point coordinates.

6. The method for calibrating a large field-of-view multi-view camera according to claim 1, characterized in that, After obtaining the three-dimensional point coordinates of each target feature point, the method further includes: The projection points of the three-dimensional points are obtained based on the coordinates of the three-dimensional points, the intrinsic parameters of the monocular camera, and the extrinsic parameters of the multi-viewpoint. The reprojection error of the target feature points is optimized based on the image points of the three-dimensional points in multiple viewpoints and the projection points in the corresponding viewpoints; wherein, in the optimization process, the visibility weight determines whether the three-dimensional points are visible in a viewpoint.

7. The method for calibrating a large field-of-view multi-view camera according to claim 1, characterized in that, Multi-view camera calibration based on the aforementioned three-dimensional point coordinates includes: Within the large field of view plane, one camera is designated as the main camera, and the remaining cameras are designated as sub-cameras; Based on the three-dimensional point coordinates, obtain the region coordinate transformation matrix between each region; Based on the 3D point coordinates, camera intrinsic parameters, target feature points, image coordinates of the 3D points under the regional camera, and the camera's projection function, obtain the regional camera coordinate transformation matrix between each camera and the corresponding region; Based on the region coordinate transformation matrix and the region camera coordinate transformation matrix, the local coordinate system of each sub-camera is transformed to the global coordinate system of the main camera for multi-camera calibration.

8. A calibration system for a large field of view multi-view camera, characterized in that, The system includes: The first acquisition module is used to acquire multi-view images of a target board within a large field of view using a monocular camera; the target board includes several coded targets; The second acquisition module acquires target feature points of several coded targets based on the target plate image, and acquires target code values ​​based on each target feature point. The calibration module is used to obtain the three-dimensional point coordinates of each target feature point based on each target code value, so as to perform multi-view camera calibration based on the three-dimensional point coordinates.

9. An electronic device, characterized in that, The electronic device includes: a processor and a memory; The memory is used to store computer programs; The processor is used to execute the computer program stored in the memory to cause the electronic device to perform the large field-of-view multi-view camera calibration method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by an electronic device, the program implements the large field-of-view multi-view camera calibration method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • A global calibration method for active phase target large field of view multi-camera measurement system

    CN122473282A