Multi-modal multi-channel complementary visual data calibration method and apparatus
By aligning the complementary cameras and auxiliary cameras in time and space, and using the original calibration data of the auxiliary cameras for data calibration, the defects of multi-modal multi-path complementary visual data calibration are solved, and more efficient data annotation and wider scene coverage are achieved, improving the quality and applicability of the data set.
Patent Information
- Application Number
- PCT/CN2024/116943
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-20
- Filing Date
- 2024-09-04
- Publication Date
- 2025-06-26
AI Technical Summary
There is a lack of suitable calibration methods in the prior art to calibrate multimodal multi-path complementary visual data, resulting in insufficient data set labeling information, limited sample number, narrow scene coverage, complex collection methods and limited sensor types.
By obtaining the original calibration data of the complementary camera and the auxiliary camera based on the pre-selected calibration object and the pre-constructed calibration data acquisition device, adjusting the camera time alignment and forming a binocular camera calibration pair for spatial alignment, and using the original calibration data of the auxiliary camera for data calibration to obtain the target calibration data of the complementary camera.
It realizes effective calibration of multimodal multi-path complementary visual data, improves the annotation information and scene coverage of the data set, enhances the diversity and applicability of the data, and supports higher-precision computer vision and deep learning tasks.
Smart Images

Figure CN2024116943_26062025_PF_FP_ABST
Abstract
Description
Multimodal multi-channel complementary visual data calibration method and device
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to Chinese patent application No. 2023117678302, filed on December 20, 2023, entitled “Multimodal Multi-channel Complementary Visual Data Calibration Method and Device,” which is incorporated herein by reference in its entirety. Technical Field
[0003] The present application relates to the field of camera calibration technology, and in particular to a multi-modal multi-channel complementary visual data calibration method and device. Background Art
[0004] A multimodal dataset is one that contains multiple different types of data, such as images, text, and audio. A multi-path dataset is a specific form of a multimodal dataset, where different data modalities are acquired through separate data paths, completing parallel and asynchronous data transmission. Alignment and synchronization between data paths relies on hardware timestamps and relative spatial positions.
[0005] What these datasets share in common is multimodal data, diverse collection platforms and locations, and a variety of sensor types, offering ample application potential. These datasets can be widely used for tasks such as computer vision, deep learning, moving object detection, and object tracking, enabling applications in diverse fields such as autonomous driving, robotic navigation, and environmental monitoring. These datasets provide rich information to solve complex problems. For example, combining images and text can be used for image annotation and automatic description, and combining audio and text can be used for speech recognition and machine translation. By studying and analyzing multimodal datasets, we can better understand the relationships between different modalities and develop more innovative and high-performance multimodal intelligent systems. The use of multimodal datasets not only brings higher accuracy and effectiveness, but also better meets user needs and provides more personalized services.
[0006] However, most of these datasets lack annotation information and suffer from limitations such as limited sample size, narrow scene coverage, complex collection methods, and a limited number of sensor types. The emergence of complementary vision sensors (CVS) offers a new approach to addressing this problem. However, since complementary cameras are a new type of neuromorphic camera with output characteristics from different modalities, the labels required for data collection from complementary cameras often come from multimodal, multi-viewing sensor systems. Currently, there is no suitable calibration method for the collected multimodal, multi-channel complementary vision data.
[0007] In summary, how to calibrate multimodal and multi-channel complementary visual data is an important issue that needs to be urgently solved in the industry.
[0008] Summary of the Invention
[0009] The present application provides a multimodal multi-channel complementary visual data calibration method and device to solve the defect in the prior art that there is no suitable calibration for multimodal multi-channel complementary visual data, and to realize the calibration of multimodal multi-channel complementary visual data.
[0010] This application provides a multimodal multi-channel complementary visual data calibration method, comprising:
[0011] Acquiring raw calibration data collected by the complementary camera and the auxiliary camera based on a pre-selected calibration object and a pre-built calibration data acquisition device; the pre-built calibration data acquisition device includes a complementary camera and at least one auxiliary camera; the raw calibration data of the complementary camera is multi-modal multi-channel complementary visual data;
[0012] adjusting, based on the original calibration data, time-aligning the complementary camera and the auxiliary camera;
[0013] The complementary camera and any auxiliary camera form a binocular camera calibration pair; for each binocular camera calibration pair, the complementary camera and the auxiliary camera are spatially aligned according to the original calibration data to obtain a mapping relationship between the complementary camera and the auxiliary camera;
[0014] According to the mapping relationship, the original calibration data of the auxiliary camera is used to calibrate the original calibration data of the complementary camera to obtain target calibration data.
[0015] According to a multimodal multi-channel complementary visual data calibration method provided by the present application, the complementary camera and the auxiliary camera are arranged in the same plane with their optical axes aligned and parallel and their sensor planes parallel; the complementary camera is located in the center of the plane, and the auxiliary camera is arranged around the complementary camera; the complementary camera and the auxiliary camera are connected to the same hardware trigger signal.
[0016] According to a multimodal multi-channel complementary visual data calibration method provided by the present application, the complementary camera and the auxiliary camera are spatially aligned according to the original calibration data to obtain a mapping relationship between the complementary camera and the auxiliary camera, which also includes:
[0017] Monocular camera calibration is performed on the complementary camera and the auxiliary camera respectively according to the original calibration data to obtain intrinsic and extrinsic parameter matrices of the complementary camera and the auxiliary camera.
[0018] According to a multimodal multi-path complementary visual data calibration method provided by the present application, the auxiliary camera in the binocular camera calibration pair includes a high dynamic range camera or a high-speed camera, and the slow path of the complementary camera and the auxiliary camera are triggered in the same path; the complementary camera and the auxiliary camera are spatially aligned according to the original calibration data to obtain a mapping relationship between the complementary camera and the auxiliary camera, specifically including:
[0019] Solving the relative translation matrix of the complementary camera and the auxiliary camera according to the intrinsic and extrinsic parameter matrix;
[0020] Substitute the relative translation matrix into the pre-constructed binocular calibration equation to obtain the mapping relationship of the binocular camera calibration pair.
[0021] According to a multimodal multi-path complementary visual data calibration method provided by the present application, the auxiliary camera in the binocular camera calibration pair includes an event camera, and the fast path of the complementary camera and the event camera are triggered in the same path; the complementary camera and the auxiliary camera are spatially aligned according to the original calibration data to obtain a mapping relationship between the complementary camera and the auxiliary camera, specifically comprising:
[0022] Performing a projective transformation on the auxiliary camera according to the intrinsic and extrinsic parameter matrices of the complementary camera and the auxiliary camera;
[0023] Acquire second original calibration data acquired by the complementary camera and the auxiliary camera after projection transformation according to the moving calibration object;
[0024] Adjusting the exposure time of the complementary camera fast path so that the complementary camera and the auxiliary camera have the same movement amplitude;
[0025] Performing grayscale reconstruction based on the second original calibration data of the complementary camera and the auxiliary camera to obtain a grayscale reconstruction result;
[0026] Binarizing the grayscale reconstruction result to obtain a grayscale binarization result;
[0027] Comparing the grayscale binarization results of the complementary camera and the auxiliary camera, terminating the calibration if the overlap reaches a threshold, and recording the mapping relationship of the binocular camera calibration pair.
[0028] According to a multimodal multi-channel complementary visual data calibration method provided by the present application, grayscale reconstruction is performed based on the second original calibration data of the complementary camera to obtain a grayscale reconstruction result, specifically including:
[0029] Based on the Poisson equation, the relationship between the divergence and gradient of the spatially differential field is iterated according to the second original calibration data to obtain a grayscale reconstruction result of the complementary camera.
[0030] According to a multimodal multi-channel complementary visual data calibration method provided by the present application, adjusting the complementary camera and the auxiliary camera to be time-aligned based on the original calibration data specifically includes:
[0031] A preset time delay is added to the auxiliary camera based on a hardware trigger signal, and the image data included in the original calibration data ensures that the starting points of the actual exposure times of the auxiliary camera and the complementary camera are completely consistent.
[0032] According to a multimodal multi-channel complementary visual data calibration method provided by the present application, the complementary camera and the auxiliary camera are calibrated as monocular cameras respectively according to the original calibration data to obtain the intrinsic and extrinsic parameter matrices of the complementary camera and the auxiliary camera, specifically comprising:
[0033] Using the Zhang Zhengyou calibration method, for the complementary camera and each of the auxiliary cameras, a preset number of calibration objects at preset distances are calibrated according to the original calibration data, and the intrinsic and extrinsic parameter matrices of the complementary camera and the auxiliary camera are calculated using the maximum likelihood method.
[0034] According to a multimodal multi-channel complementary visual data calibration method provided by the present application, the data calibration includes label addition and / or copy backup; according to the mapping relationship, the original calibration data of the auxiliary camera is used to calibrate the original calibration data of the complementary camera to obtain target calibration data, and then further includes:
[0035] The calibration data acquisition device acquires complementary camera data and auxiliary camera data according to the calibration, and the complementary camera data, the auxiliary camera data and the mapping relationship determined by calibration constitute a data set.
[0036] The present application also provides a device, comprising:
[0037] A data acquisition unit, configured to acquire raw calibration data acquired by the complementary camera and the auxiliary camera based on a preselected calibration object and a pre-built calibration data acquisition device; the pre-built calibration data acquisition device includes a complementary camera and at least one auxiliary camera; the raw calibration data of the complementary camera is multimodal multi-channel complementary visual data;
[0038] a time alignment unit, configured to adjust the complementary camera and the auxiliary camera to be time aligned according to the original calibration data;
[0039] a spatial alignment unit, configured to form a binocular camera calibration pair with the complementary camera and any one of the auxiliary cameras; and for each binocular camera calibration pair, spatially aligning the complementary camera and the auxiliary camera according to the original calibration data to obtain a mapping relationship between the complementary camera and the auxiliary camera;
[0040] The target calibration unit is configured to perform data calibration on the original calibration data of the complementary camera using the original calibration data of the auxiliary camera according to the mapping relationship to obtain target calibration data.
[0041] The present application also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the multimodal multi-channel complementary visual data calibration method as described above is implemented.
[0042] The present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the multi-modal multi-channel complementary visual data calibration method as described in any one of the above is implemented.
[0043] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the multi-modal multi-channel complementary visual data calibration methods described above.
[0044] The multimodal, multi-channel complementary visual data calibration method and apparatus provided by the present application acquires raw calibration data collected by a complementary camera and an auxiliary camera based on a pre-selected calibration object and a pre-built calibration data acquisition device; the pre-built calibration data acquisition device includes a complementary camera and at least one auxiliary camera; the raw calibration data of the complementary camera is multimodal, multi-channel complementary visual data; based on the raw calibration data, the complementary camera and the auxiliary camera are adjusted to be time-aligned; the complementary camera and any one of the auxiliary cameras are combined into a binocular camera calibration pair; for each binocular camera calibration pair, the complementary camera and the auxiliary camera are spatially aligned based on the raw calibration data to obtain a mapping relationship between the complementary camera and the auxiliary camera; based on the mapping relationship, the raw calibration data of the auxiliary camera is used to perform data calibration on the raw calibration data of the complementary camera to obtain target calibration data. The present application proposes a calibration method for unconventional multimodal, multi-channel visual data based on a pre-designed calibration data acquisition device, in which the complementary camera and each auxiliary camera are simultaneously spatiotemporally calibrated. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions in the present application or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0046] FIG1 is a flow chart of a multimodal multi-channel complementary visual data calibration method according to an embodiment of the present application;
[0047] FIG2 is a schematic diagram of the basic arrangement of calibration data acquisition equipment according to an embodiment of the multimodal multi-channel complementary visual data calibration method provided in an embodiment of the present application;
[0048] FIG3 is an example of a camera space calibration model and a connection method of an embodiment of a multimodal multi-channel complementary visual data calibration method provided in an embodiment of the present application;
[0049] FIG4 is a second flow chart of the multimodal multi-channel complementary visual data calibration method provided in an embodiment of the present application;
[0050] FIG5 is a schematic diagram of a grayscale reconstruction process according to an embodiment of a multimodal multi-channel complementary visual data calibration method provided in an embodiment of the present application;
[0051] FIG6 is a schematic structural diagram of a multimodal multi-channel complementary visual data calibration device provided in an embodiment of the present application;
[0052] FIG7 is a schematic structural diagram of an electronic device provided in an embodiment of the present application.
[0053] Reference numerals: 610: data acquisition unit; 620: time alignment unit; 630: space alignment unit; 640: target calibration unit. DETAILED DESCRIPTION
[0054] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions in this application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.
[0055] The collection system for multimodal datasets typically consists of a spatially compact sensor array. Since the core application of multimodal datasets is multimodal tasks, which require good consistency, synchronization, and alignment in space and time, the integration and synchronization of the collection system are extremely important.
[0056] Existing collection methods include, but are not limited to, rooftop sensor arrays, handheld sensor arrays, head-mounted sensor arrays, and drone payloads. Integrating with IoT technology has also enabled the collection of multimodal data through online channels. These datasets are generally multi-channel, multi-modal, multi-task, large in size, and challenging to process.
[0057] As shown in Table 1, Table 1 lists the existing common multimodal vision datasets.
[0058] Table 1 Existing common multimodal vision datasets
[0059] The aforementioned datasets all use multimodal data collected using different types of sensors, including visual and depth information. Furthermore, these datasets were collected using a variety of platforms, including head-mounted devices, handheld devices, vehicle-mounted devices, and drones. The data was collected in a variety of locations, including indoors, outdoors, in cities, and in forests. These datasets also use a variety of sensor types, including Prophesee Gen4, DAVIS346, LiDAR, and Vision sensors. Furthermore, most of these datasets lack annotation information, with the exception of the KITTI dataset, which covers 37 categories, and the DSEC and M3ED datasets, which each cover 11 categories.
[0060] Existing neuromorphic sensors have significant shortcomings in terms of information integrity. Most current multimodal cameras introduce time-varying information, but this is incomplete information required for motion perception. The complementary vision sensor (CVS) is a new type of neuromorphic vision sensor based on the theory of complementary perception. The complementary theory refers to the characteristics of human vision and requires that the acquired data belong to multiple different pathways, and the data structure of each pathway should be complementary in different properties. The characteristic of CVS is that it imitates the human retina and outputs multiple data modalities with a single vision chip. These different modal data have strong complementary properties, including in terms of sampling accuracy, sampling speed, dynamic range, sensitivity, color range, spatial resolution, etc. These pathways are different from each other and can complement each other to ensure information integrity, so they are also called complementary cameras. Complementary cameras based on complementary vision sensors can solve the problem of incomplete information in current multimodal cameras.
[0061] Based on this, this application proposes a multimodal multi-channel complementary visual data calibration method based on complementary cameras.
[0062] The following describes the multimodal multi-channel complementary visual data calibration method of the present application in conjunction with Figures 1 to 5, which is one of the flow charts of the multimodal multi-channel complementary visual data calibration method provided in an embodiment of the present application. As shown in Figure 1, the method includes:
[0063] Step 110: Based on a preselected calibration object and a pre-built calibration data acquisition device, original calibration data collected by the complementary camera and the auxiliary camera are obtained; the pre-built calibration data acquisition device includes a complementary camera and at least one auxiliary camera; the original calibration data of the complementary camera is multimodal multi-channel complementary visual data.
[0064] The calibration data acquisition device is pre-built and designed primarily for recording calibration test data for complementary cameras. Specifically, the calibration data acquisition device is a data acquisition device centered around the complementary camera, comprising the complementary camera and at least one auxiliary camera. In one embodiment, the auxiliary camera includes an HDR (High-Dynamic Range) camera, a high-speed camera, and a DVS (Dynamic Vision Sensor) camera, as shown in Figure 2.
[0065] It can be understood that a complementary camera is a camera that can simultaneously achieve high dynamic range, high sampling accuracy, and high sampling speed. Complementary cameras construct a hybrid pixel arrangement in the image sensor and design a hybrid data readout circuit. Therefore, the same CMOS chip can output RGB, spatial difference, and temporal difference data information, and encode this information in different modalities. This information is transmitted using different data channels. In extreme scenarios, ordinary high dynamic range cameras, high-speed cameras, and event cameras cannot provide data labels for them alone. In order to obtain calibrated multi-modal, multi-channel complementary visual data, it is necessary to combine multiple cameras and use hardware synchronization, and perform data labeling through multi-sensor fusion methods. Therefore, for the complementary camera in the calibration data acquisition device, it is first necessary to synchronize it with multiple image sensors and calibrate the multi-eye system.
[0066] Furthermore, since time-response pattern data is involved, the calibration object must be something that responds to time changes. It should be understood that this application does not limit the calibration object; in practice, the calibration object can be selected based on actual requirements. In some embodiments, a high-speed turntable is used as the calibration object.
[0067] Step 120: Adjust the complementary camera and the auxiliary camera to be time-aligned according to the original calibration data.
[0068] First, the complementary camera and the auxiliary camera are time-aligned. It's understood that time alignment means that the actual exposure times of each camera are completely consistent. In the design of the calibration data acquisition device, the cameras are triggered using hardware synchronization, but exposure is not completely synchronized.
[0069] Furthermore, in some embodiments, adjusting the time alignment of the complementary camera and the auxiliary camera based on the original calibration data specifically includes:
[0070] A preset time delay is added to the auxiliary camera based on a hardware trigger signal, and the image data included in the original calibration data ensures that the starting points of the actual exposure times of the auxiliary camera and the complementary camera are completely consistent.
[0071] Specifically, taking into account the differences in exposure time and response speed of each camera, in some embodiments, a delay of the µs level is added to each camera based on the hardware trigger signal. The delays for the complementary camera, high dynamic range camera, high-speed camera, and event camera are 0, t1, t2, and t3, respectively. That is, after the complementary camera starts exposing, the other cameras wait for t µs before starting exposure, and the image data is used to ensure that the starting points of their actual exposure times are completely consistent.
[0072] Step 130: The complementary camera and any auxiliary camera form a binocular camera calibration pair; for each binocular camera calibration pair, the complementary camera and the auxiliary camera are spatially aligned according to the original calibration data to obtain a mapping relationship between the complementary camera and the auxiliary camera.
[0073] In some embodiments, spatially aligning the complementary camera and the auxiliary camera according to the original calibration data to obtain a mapping relationship between the complementary camera and the auxiliary camera further includes:
[0074] Monocular camera calibration is performed on the complementary camera and the auxiliary camera respectively according to the original calibration data to obtain intrinsic and extrinsic parameter matrices of the complementary camera and the auxiliary camera.
[0075] In existing camera imaging theory, the camera contains four coordinate systems: the world coordinate system, the camera coordinate system, the image coordinate system, and the pixel coordinate system. It is understood that before performing spatial alignment, each camera needs to be calibrated for a monocular camera to obtain the camera intrinsic and extrinsic parameter matrices of each camera. Furthermore, the intrinsic and extrinsic parameter matrix H' is the product of the intrinsic parameter matrix A and the extrinsic parameter matrix [R, T]. The intrinsic parameter matrix depends on the internal parameters of the camera and is independent of the calibration method. The extrinsic parameter matrix is obtained by the aforementioned monocular camera calibration method.
[0076] After the monocular camera calibration is completed, the complementary camera needs to be spatially aligned with each auxiliary camera. In this embodiment, a binocular camera calibration pair is generated between the complementary camera and any auxiliary camera, and the binocular camera calibration pair is used to spatially align the complementary camera and the auxiliary camera.
[0077] It should be understood that in this process, it is necessary to ensure that the three-axis rotation angles in the camera coordinate system are fixed and aligned with high precision. In actual operation, the housing design of each camera in the data acquisition device can be calibrated to achieve this, so that the left and right cameras of the binocular camera calibration pair do not rotate relative to each other (the rotation is known to be 0).
[0078] In a specific embodiment, the auxiliary camera includes a high dynamic range camera, an event camera, and a high-speed camera. In this case, the complementary camera needs to be spatially aligned with the high dynamic range camera, the event camera, and the high-speed camera.
[0079] Step 140: According to the mapping relationship, the original calibration data of the auxiliary camera is used to calibrate the original calibration data of the complementary camera to obtain target calibration data.
[0080] Specifically, in some embodiments, according to the calibrated calibration data acquisition device, training and verification data are constructed based on complementary cameras, and the data generated by the auxiliary camera is used as a calibration copy and label to obtain target calibration data for use by the deep learning algorithm.
[0081] Based on the above embodiment, in this method, the complementary camera and the auxiliary camera are arranged in the same plane with their optical axes aligned and parallel and their sensor planes parallel; the complementary camera is located in the center of the plane, and the auxiliary camera is arranged around the complementary camera; the complementary camera and the auxiliary camera are connected to the same hardware trigger signal.
[0082] Specifically, in a specific embodiment, a schematic diagram of the calibration data acquisition device is shown in Figure 2, which includes a complementary camera (located in the middle), an event camera, a high dynamic range camera, and a high-speed camera. After each camera is packaged in a module, the optical axes are aligned and parallel, and the sensor planes are adjusted to be parallel. Due to the size limitation of the packaging shell, the camera's phase plane will have a fixed deviation in the xyz direction, which is one of the core problems that need to be solved in calibration. Since the complementary camera has an RGB mode, the classic binocular camera calibration is used for spatial alignment, and the calibration of high-speed cameras and high dynamic range cameras can be completed in static mode. In some embodiments, the complementary camera includes a fast path and a slow path, and the signals of the fast path and the slow path run at a fixed multiple and there is a synchronization point.
[0083] Based on the above embodiment, in this method, the auxiliary camera in the binocular camera calibration pair includes a high dynamic range camera or a high-speed camera, and the slow path of the complementary camera and the auxiliary camera are triggered in the same path; the spatial alignment of the complementary camera and the auxiliary camera based on the original calibration data to obtain the mapping relationship between the complementary camera and the auxiliary camera specifically includes:
[0084] Solving the relative translation matrix of the complementary camera and the auxiliary camera according to the intrinsic and extrinsic parameter matrix;
[0085] Substitute the relative translation matrix into the pre-constructed binocular calibration equation to obtain the mapping relationship of the binocular camera calibration pair.
[0086] Specifically, as shown in Figure 3(a), for an HDR-CVS binocular camera calibration pair, HDR is used to calibrate the static imaging component and spatial gradient component of the complementary camera. In other words, simply solving the relative translation matrix of the complementary camera and the auxiliary camera and substituting it into the pre-built binocular calibration equation yields the mapping relationship for the camera pair. In some embodiments, the calibration process is implemented using relevant functions in the open-source software OpenCV.
[0087] First, define the following variables:
[0088] K1: intrinsic parameter matrix of camera 1;
[0089] K2: intrinsic parameter matrix of camera 2;
[0090] R1: rotation matrix of camera 1;
[0091] R2: rotation matrix of camera 2;
[0092] T1: translation vector of camera 1;
[0093] T2: translation vector of camera 2;
[0094] The specific binocular camera calibration process is as follows:
[0095] First, by collecting a series of image pairs of checkerboard patterns, the checkerboard corners in the image pairs are detected using the calibration board corner detection algorithm. Here, the findChessboardCorners function is used.
[0096] Next, set the world coordinate system. Choose a fixed local coordinate system and align it with the corners of the checkerboard. Typically, the top left corner of the checkerboard is chosen as the origin, and the dimensions of the checkerboard are used as the units of the world coordinate system (e.g., meters or millimeters).
[0097] Next, the correspondence between the checkerboard corner points and the image coordinate system is used to calculate the intrinsic parameter matrices K1 and K2 of the two cameras by calling the corresponding camera calibration function (such as the calibrateCamera function in OpenCV).
[0098] Next, we use the checkerboard corners in the image pairs to align the binocular cameras and calculate the two camera rotation matrices R1 and R2, as well as the translation vectors T1 and T2, using the stereoCalibrate function in OpenCV.
[0099] The above parameters include the geometric relationship between the two cameras and the registration method. Finally, the calibration results need to be evaluated. Based on the correction parameters and the calibration image pairs obtained during the calibration process, the accuracy of the calibration can be assessed by calculating metrics such as reprojection error. Reprojection error refers to the difference between the corner points in the world coordinate system after reprojection into the image coordinate system and the actual detected corner points.
[0100] For high-speed camera-CVS binocular camera calibration, the high-speed camera is used to calibrate the static imaging component and spatial gradient component of the complementary camera. In other words, simply solving the relative translation matrix of the complementary camera and the auxiliary camera and substituting it into the pre-built binocular calibration equation yields the mapping relationship for the camera pair. In some embodiments, the calibration process is implemented using relevant functions in the open-source software OpenCV.
[0101] Based on the above embodiment, in this method, the auxiliary camera in the binocular camera calibration pair includes an event camera, and the fast path of the complementary camera and the event camera are triggered in the same way; the spatial alignment of the complementary camera and the auxiliary camera based on the original calibration data to obtain a mapping relationship between the complementary camera and the auxiliary camera specifically includes:
[0102] Performing a projective transformation on the auxiliary camera according to the intrinsic and extrinsic parameter matrices of the complementary camera and the auxiliary camera;
[0103] Acquire second original calibration data acquired by the complementary camera and the auxiliary camera after projection transformation according to the moving calibration object;
[0104] Adjusting the exposure time of the complementary camera fast path so that the complementary camera and the auxiliary camera have the same movement amplitude;
[0105] Performing grayscale reconstruction based on the second original calibration data of the complementary camera and the auxiliary camera to obtain a grayscale reconstruction result;
[0106] Binarizing the grayscale reconstruction result to obtain a grayscale binarization result;
[0107] Comparing the grayscale binarization results of the complementary camera and the auxiliary camera, terminating the calibration if the overlap reaches a threshold, and recording the mapping relationship of the binocular camera calibration pair.
[0108] Because the data collected by complementary cameras, including SD (Spatial Difference) and TD (Temporal Difference) data, has fast data speed, low codeword accuracy, and a large dynamic range, traditional high-speed cameras and high dynamic range cameras are not suitable for use as labels for this modality. The output results of event cameras are also difficult to align with the RGB information obtained in CVS.
[0109] In this embodiment, the event camera and the CVS camera are spatially aligned and time-synchronized. Specifically, since the DVS and CVS-TD modes only respond to time changes, a moving calibration object is used to obtain the second original calibration data. Since the moving checkerboard can bring sufficient response, in some embodiments, the calibration object is selected as a high-speed turntable with chess and card grids. In other embodiments, a fixed motion device (a fixed disc is used here) is added to the stationary calibration object with chess and card grids, and a servo motor is used to drive it to obtain a moving calibration object.
[0110] Grayscale reconstruction is then performed based on the second raw calibration data of each complementary camera and the auxiliary camera to obtain a grayscale reconstruction result. Specifically, due to the large error in the temporal response mode, the motion amplitudes of the complementary and auxiliary cameras need to be adjusted to correspond before grayscale reconstruction. During operation, grayscale reconstruction is performed on the DVS using event accumulation within a preset time. In some embodiments, the preset time is 100 μs. It is understood that this will retain 100 μs of motion blur. Simultaneously, the exposure time of the TD and SD data in the complementary cameras is adjusted to the preset time to ensure consistent motion amplitudes.
[0111] After the motion amplitude adjustment is completed, the complementary camera and the event camera are respectively used to construct grayscale images, and then the registration is performed. This calibration process has both temporal and spatial errors, so a set of coarse and fine calibration processes are required, as shown in Figure 4. The overall process is as follows:
[0112] S201: After the complementary camera and the event camera are calibrated for the monocular camera respectively, the distortion coefficient and the intrinsic and extrinsic parameter rotation matrix R and translation vector T are obtained;
[0113] S202: Obtain a rough relationship R', T' between CVS and DVS according to the symmetry transformation R, T;
[0114] S203: Turn on the turntable and perform projection transformation on the DVS according to the internal and external parameter matrix H';
[0115] S204: DVS and CVS perform grayscale reconstruction;
[0116] S205: Adjust the black accumulation time and trigger time difference of DVS according to the step size, keep the same frame rate as CVS, and find the setting with the minimum matching error;
[0117] S206: Recalculate the corresponding relationships R and T of the two cameras according to the time synchronization result under the setting;
[0118] S207: Record a video, perform grayscale reconstruction and binarization, and determine whether the overlap reaches a threshold;
[0119] S208: If the overlap reaches a threshold, the calibration is terminated and the CVS-DVS relationship is recorded; if the overlap does not reach the threshold, the process jumps to S204.
[0120] It should be noted that the threshold value can be selected according to actual needs, and this application does not impose any restrictions on this. In one embodiment, in order to minimize the influence of non-ideal factors of lighting and camera, the grayscale for calculating the overlap is first adaptively binarized to obtain a binarized checkerboard. The L1 error is then calculated for the two. Since the L1 error of the binarized image can be directly calculated using XOR, the result obtained is the ratio of different pixels. Taking into account noise and adaptive errors, the threshold value is selected here as 20% of the total number of pixels.
[0121] It should be noted that for CVS, since there is only fixed motion in the scene and the illumination remains unchanged, the response of TD and the response of SD encode the same information, and the result obtained by using SD for grayscale reconstruction is the same as the theoretical result of TD.
[0122] Furthermore, for HDR-CVS and DVS-CVS binocular camera calibration pairs, multiple hardware triggers are required for time synchronization. The hardware signals are provided by the same PC, as shown in Figure 3(b). The slow path in CVS and HDR maintain the same trigger path, while the fast paths in DVS and CVS maintain the same trigger path. Both signals operate at a fixed multiple to ensure synchronization. Due to the presence of time-mode cameras, the camera exposure time must be calibrated before obtaining the signal to ensure that the exposure time is consistent, allows for reasonable imaging, and is lower than the synchronization signal frame rate.
[0123] Based on the above embodiment, in this method, grayscale reconstruction is performed according to the second original calibration data of the complementary camera to obtain a grayscale reconstruction result, which specifically includes:
[0124] Based on the Poisson equation, the relationship between the divergence and gradient of the spatially differential field is iterated according to the second original calibration data to obtain a grayscale reconstruction result of the complementary camera.
[0125] Specifically, the grayscale reconstruction method of the CVS-differential data path is the prerequisite for spatiotemporal calibration of CVS and other high-speed cameras. Two different reconstruction algorithms are provided here, one using SD data and the other using TD data. Specifically, the method based on SD data is based on the Poisson equation, which describes the relationship between the divergence and gradient of the field. That is, it is equivalent to having a gradient field in the x and y directions provided by SD, and allowing a blank image to grow a grayscale image that conforms to the gradient field. This method is based on Poisson editing, that is, it is iterated based on the following formula:
[0126] Here f refers to the grayscale image, SD is the spatial difference, Ω is the image boundary, and f * is the reference value of the boundary value, and its solution will satisfy the Poisson equation: Δf = div SD. According to this method, SD can directly generate a continuous grayscale image. Based on this basic principle, this embodiment proposes a spatial difference-grayscale reconstruction algorithm for a scaled image pyramid, the process of which is shown in Figure 5:
[0127] S301: Initialize the zero matrix O;
[0128] S302: SD obtains the boundary result by integration and sets the default value to O;
[0129] S303: SD is downsampled K times and interpolated to its original size to obtain SD_k;
[0130] S304: Copy O to old_O;
[0131] S305: Calculate the divergence of SD_k; calculate the mean of O;
[0132] S306: Add the divergence of SD_k and the mean of O, and assign the value to O;
[0133] S307: Determine whether the change between O and old_O is lower than the threshold; if so, jump to step S308; if not, jump to step S304;
[0134] S308: Set K=K-1; if K>0, jump to step S303; if K=0, jump to step S309;
[0135] S309: Outputting a grayscale image.
[0136] Furthermore, the event camera uses the general E2VID method for grayscale reconstruction, an open-source neural network-based reconstruction method. Specifically, a recurrent convolutional neural network is trained to accumulate the continuous event stream into grayscale images.
[0137] Based on the above embodiment, in this method, performing monocular camera calibration on the complementary camera and the auxiliary camera respectively according to the original calibration data to obtain the intrinsic and extrinsic parameter matrices of the complementary camera and the auxiliary camera specifically includes:
[0138] Using the Zhang Zhengyou calibration method, for the complementary camera and each of the auxiliary cameras, a preset number of calibration objects at preset distances are calibrated according to the original calibration data, and the intrinsic and extrinsic parameter matrices of the complementary camera and the auxiliary camera are calculated using the maximum likelihood method.
[0139] Specifically, the Zhang Zhengyou calibration method is used to calibrate the complementary camera and the auxiliary camera for monocular calibration. During the calibration process, raw calibration data is obtained based on a calibration object, and the camera intrinsic and extrinsic parameter matrices are solved using the maximum likelihood method. This allows the imaging position of a target point in a given world coordinate to be determined. In some embodiments, n sufficiently distant chess and card squares are selected as calibration objects for calibration, where n is a preset number. Furthermore, in some embodiments, relevant functions in the open source software OpenCV are often used to implement the calibration process.
[0140] Based on the above embodiment, in this method, the data calibration includes label addition and / or copy backup; the original calibration data of the complementary camera is calibrated using the original calibration data of the auxiliary camera according to the mapping relationship to obtain target calibration data, and then the method further includes:
[0141] The calibration data acquisition device acquires complementary camera data and auxiliary camera data according to the calibration, and the complementary camera data, the auxiliary camera data and the mapping relationship determined by calibration constitute a data set.
[0142] The multimodal, multi-channel complementary visual data acquired by complementary cameras belongs to multiple different channels, and the data structure of each channel should be complementary in different properties. In some embodiments, the complementary properties include temporal resolution (fast and slow complementarity), spatial resolution (high and low complementarity), color (complementarity of spectral sensitivity ranges such as color, grayscale, infrared, and ultraviolet), sensitivity (i.e., high and low complementarity of response coefficients to light intensity), response modality (integrated intensity or differential change), and data accuracy (high and low complementarity).
[0143] The data structure of the data set of the present application requires that the data has any number of channels, wherein the properties of at least one primitive between any two channels are complementary. The data source can be independently implemented by complementary cameras. In some embodiments, the data source can also be generated by simulation of a separate APS (Active Pixel Sensor), DVS (Event Camera), DAVIS (Dynamic and Active-Pixel Vision Sensor) or spatial gradient camera, or can be generated by a combination of different types of cameras. Specifically, in some embodiments, two APS cameras or APS+DVS cameras with complementary primitive properties are used to acquire data. It should be understood that in the step of acquiring data for the data set, if complementary cameras are not used, any number of cameras (APS, DVS or complementary cameras) can be selected to acquire multi-channel data, and the physical implementation of the channel is any number of cameras. It can be expanded to any number of channels, requiring that the properties of at least one primitive between any two channels are complementary. As shown in Table 2, Table 2 provides a schematic diagram of the composition of various data source forms.
[0144] Table 2 Schematic diagram of the composition of various data sources
[0145] For the data independently acquired by the complementary cameras, due to the lack of labeled data, the data set provided in this embodiment will use CVS to construct training and verification data, and the data generated by other auxiliary cameras will be used as calibration copies and labels for use by the deep learning algorithm. The data acquisition requirements proposed in this application include the use of complementary visual sensors in the above-mentioned CMOS chip and the use of complementary data modalities outside the chip. Compared with existing multimodal data sets, higher requirements are placed on the mutual complementation and completion of different data. Due to the use of complementary visual sensors, this application can be more compact.
[0146] This dataset is provided for deep learning training. The dataset includes CVS and auxiliary camera data, as well as internal and external parameters determined by calibration. Furthermore, in some embodiments, the dataset additionally provides lidar and IMU data as auxiliary for camera pose estimation tasks and depth estimation tasks. In this way, the capabilities of the CVS camera can be maximized, and the sparsity and unavailability of corner case data can be solved from the source of the data. Densified corner case data can be introduced into the simulation data, and the complementary dataset itself can directly include corner cases that cannot be included in general datasets (such as high-speed HDR, flash, etc.). It is crucial to the safety and high performance of open-world robotics tasks.
[0147] Furthermore, each pathway in the complementary dataset is carefully designed, and a single pathway can be considered a complete data source, allowing for flexible design of processing algorithms tailored to each pathway. For example, the APS data structure is a compact matrix, suitable for ANNs such as CNNs and Transformers; while the DVS, a biomimetic temporal pulse event generated by changes in scene light intensity, is suitable for SNNs; and spatial gradients are sparse and can be represented as matrices, allowing for the selection of practical algorithms based on the task.
[0148] In the specific use process, the unlabeled data can be used for CVS applications such as image denoising, image dehazing, high-speed HDR image reconstruction, and super-resolution. For semantically annotated data, this application is designed to utilize the data completeness of the complementary camera itself, combined with the UNet-optical flow network, to train the reconstruction algorithm. This reconstruction method acts on complementary cameras to obtain image data with high frame rate and high dynamic range, and serves as a multimodal benchmark. Other data such as radar and IMU can also be used for alignment. Furthermore, for the obtained data set, large-scale pre-trained DETR2 and SAM can be used for data pre-labeling, and manual screening can be used to obtain detection and segmentation labels.
[0149] The multimodal multi-channel complementary visual data calibration method provided by the present application obtains raw calibration data collected by a complementary camera and an auxiliary camera based on a pre-selected calibration object and a pre-built calibration data acquisition device; the pre-built calibration data acquisition device includes a complementary camera and at least one auxiliary camera; the raw calibration data of the complementary camera is multimodal multi-channel complementary visual data; based on the raw calibration data, the complementary camera and the auxiliary camera are adjusted to be time-aligned; the complementary camera and any one of the auxiliary cameras are combined to form a binocular camera calibration pair; for each binocular camera calibration pair, the complementary camera and the auxiliary camera are spatially aligned based on the raw calibration data to obtain a mapping relationship between the complementary camera and the auxiliary camera; based on the mapping relationship, the raw calibration data of the auxiliary camera is used to perform data calibration on the raw calibration data of the complementary camera to obtain target calibration data. Based on the pre-designed calibration data acquisition device, the complementary camera and each auxiliary camera are simultaneously calibrated in time and space, and a calibration method for unconventional multimodal multi-channel visual data is proposed.
[0150] The following describes the multimodal multi-channel complementary visual data calibration device provided by the present application. The multimodal multi-channel complementary visual data calibration device described below can be referenced in correspondence with the multimodal multi-channel complementary visual data calibration method described above. Figure 6 is a schematic structural diagram of the multimodal multi-channel complementary visual data calibration device provided in an embodiment of the present application. As shown in Figure 6, the device includes:
[0151] A data acquisition unit 610 is configured to acquire raw calibration data collected by the complementary camera and the auxiliary camera based on a preselected calibration object and a pre-built calibration data acquisition device; the pre-built calibration data acquisition device includes a complementary camera and at least one auxiliary camera; the raw calibration data of the complementary camera is multimodal multi-channel complementary visual data;
[0152] a time alignment unit 620, configured to adjust the complementary camera and the auxiliary camera to be time aligned according to the original calibration data;
[0153] a spatial alignment unit 630 configured to form a binocular camera calibration pair with the complementary camera and any one of the auxiliary cameras; and for each binocular camera calibration pair, spatially aligning the complementary camera and the auxiliary camera based on the original calibration data to obtain a mapping relationship between the complementary camera and the auxiliary camera;
[0154] The target calibration unit 640 is configured to perform data calibration on the original calibration data of the complementary camera using the original calibration data of the auxiliary camera according to the mapping relationship to obtain target calibration data.
[0155] Based on the above embodiment, in the device, the complementary camera and the auxiliary camera are arranged on the same plane with their optical axes aligned and parallel and their sensor planes parallel; the complementary camera is located in the center of the plane, and the auxiliary camera is arranged around the complementary camera; the complementary camera and the auxiliary camera are connected to the same hardware trigger signal.
[0156] Based on the above embodiment, in the device, spatially aligning the complementary camera and the auxiliary camera according to the original calibration data to obtain a mapping relationship between the complementary camera and the auxiliary camera further includes:
[0157] Monocular camera calibration is performed on the complementary camera and the auxiliary camera respectively according to the original calibration data to obtain intrinsic and extrinsic parameter matrices of the complementary camera and the auxiliary camera.
[0158] Based on the above embodiment, in the device, the auxiliary camera for binocular camera calibration includes a high dynamic range camera or a high-speed camera, and the slow path of the complementary camera and the auxiliary camera are triggered in the same way; the spatial alignment unit 640 specifically includes:
[0159] Solving the relative translation matrix of the complementary camera and the auxiliary camera according to the intrinsic and extrinsic parameter matrix;
[0160] Substitute the relative translation matrix into the pre-constructed binocular calibration equation to obtain the mapping relationship of the binocular camera calibration pair.
[0161] Based on the above embodiment, in the device, the auxiliary camera for binocular camera calibration includes an event camera, and the fast path of the complementary camera and the event camera are triggered in the same way; the spatial alignment unit 630 specifically includes:
[0162] Performing a projective transformation on the auxiliary camera according to the intrinsic and extrinsic parameter matrices of the complementary camera and the auxiliary camera;
[0163] Acquire second original calibration data acquired by the complementary camera and the auxiliary camera after projection transformation according to the moving calibration object;
[0164] Adjusting the exposure time of the complementary camera fast path so that the complementary camera and the auxiliary camera have the same movement amplitude;
[0165] Performing grayscale reconstruction based on the second original calibration data of the complementary camera and the auxiliary camera to obtain a grayscale reconstruction result;
[0166] Binarizing the grayscale reconstruction result to obtain a grayscale binarization result;
[0167] Comparing the grayscale binarization results of the complementary camera and the auxiliary camera, terminating the calibration if the overlap reaches a threshold, and recording the mapping relationship of the binocular camera calibration pair.
[0168] Based on the above embodiment, in the device, grayscale reconstruction is performed according to the second original calibration data of the complementary camera to obtain a grayscale reconstruction result, which specifically includes:
[0169] Based on the Poisson equation, the relationship between the divergence and gradient of the spatially differential field is iterated according to the second original calibration data to obtain a grayscale reconstruction result of the complementary camera.
[0170] Based on the above embodiment, in the device, the time alignment unit 620 specifically includes:
[0171] A preset time delay is added to the auxiliary camera based on a hardware trigger signal, and the image data included in the original calibration data ensures that the starting points of the actual exposure times of the auxiliary camera and the complementary camera are completely consistent.
[0172] Based on the above embodiment, in the device, the monocular camera calibration of the complementary camera and the auxiliary camera is performed separately according to the original calibration data to obtain the intrinsic and extrinsic parameter matrices of the complementary camera and the auxiliary camera, specifically including:
[0173] Using the Zhang Zhengyou calibration method, for the complementary camera and each of the auxiliary cameras, a preset number of calibration objects at preset distances are calibrated according to the original calibration data, and the intrinsic and extrinsic parameter matrices of the complementary camera and the auxiliary camera are calculated using the maximum likelihood method.
[0174] Based on the above embodiment, in the device, the data calibration includes label addition and / or copy backup; the target calibration unit 650 further includes:
[0175] The calibration data acquisition device acquires complementary camera data and auxiliary camera data according to the calibration, and the complementary camera data, the auxiliary camera data and the mapping relationship determined by calibration constitute a data set.
[0176] The multimodal multi-channel complementary visual data calibration device provided by the present application obtains the original calibration data collected by the complementary camera and the auxiliary camera based on a pre-selected calibration object and a pre-built calibration data acquisition device; the pre-built calibration data acquisition device includes a complementary camera and at least one auxiliary camera; the original calibration data of the complementary camera is multimodal multi-channel complementary visual data; based on the original calibration data, the complementary camera and the auxiliary camera are adjusted to be time-aligned; the complementary camera and any one of the auxiliary cameras are combined into a binocular camera calibration pair; for each binocular camera calibration pair, the complementary camera and the auxiliary camera are spatially aligned based on the original calibration data to obtain a mapping relationship between the complementary camera and the auxiliary camera; based on the mapping relationship, the original calibration data of the auxiliary camera is used to perform data calibration on the original calibration data of the complementary camera to obtain target calibration data. Based on the pre-designed calibration data acquisition device, the complementary camera and each auxiliary camera are simultaneously calibrated in time and space, and a calibration method for unconventional multimodal multi-channel visual data is proposed.
[0177] Figure 7 illustrates a schematic diagram of the physical structure of an electronic device. As shown in Figure 7, the electronic device may include: a processor (processor) 710, a communication interface (Communications Interface) 720, a memory (memory) 730 and a communication bus 740, wherein the processor 710, the communication interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 can call logic instructions in the memory 730 to execute a multimodal multi-channel complementary visual data calibration method, which includes: obtaining raw calibration data collected by a complementary camera and an auxiliary camera based on a pre-selected calibration object and a pre-built calibration data acquisition device; the pre-built calibration data acquisition device includes a complementary camera and at least one auxiliary camera; the raw calibration data of the complementary camera is multimodal multi-channel complementary visual data; adjusting the complementary camera and the auxiliary camera to be time-aligned based on the raw calibration data; forming a binocular camera calibration pair with the complementary camera and any one of the auxiliary cameras; for each binocular camera calibration pair, spatially aligning the complementary camera and the auxiliary camera based on the raw calibration data to obtain a mapping relationship between the complementary camera and the auxiliary camera; and performing data calibration on the raw calibration data of the complementary camera using the raw calibration data of the auxiliary camera based on the mapping relationship to obtain target calibration data.
[0178] In addition, the logic instructions in the above-mentioned memory 730 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0179] On the other hand, the present application also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the multimodal multi-channel complementary visual data calibration method provided by the above methods, the method comprising: obtaining raw calibration data collected by a complementary camera and an auxiliary camera based on a pre-selected calibration object and a pre-built calibration data acquisition device; the pre-built calibration data acquisition device includes a complementary camera and at least one auxiliary camera; the raw calibration data of the complementary camera is multimodal multi-channel complementary visual data; based on the raw calibration data, adjusting the complementary camera and the auxiliary camera to achieve time alignment; forming a binocular camera calibration pair with the complementary camera and any one of the auxiliary cameras; for each binocular camera calibration pair, spatially aligning the complementary camera and the auxiliary camera based on the raw calibration data to obtain a mapping relationship between the complementary camera and the auxiliary camera; and based on the mapping relationship, performing data calibration on the raw calibration data of the complementary camera using the raw calibration data of the auxiliary camera to obtain target calibration data.
[0180] On the other hand, the present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the multimodal multi-channel complementary visual data calibration method provided by the above-mentioned methods, the method comprising: acquiring raw calibration data collected by a complementary camera and an auxiliary camera based on a pre-selected calibration object and a pre-built calibration data acquisition device; the pre-built calibration data acquisition device includes a complementary camera and at least one auxiliary camera; the raw calibration data of the complementary camera is multimodal multi-channel complementary visual data; adjusting the complementary camera and the auxiliary camera to be time-aligned based on the raw calibration data; forming a binocular camera calibration pair with the complementary camera and any one of the auxiliary cameras; for each binocular camera calibration pair, spatially aligning the complementary camera and the auxiliary camera based on the raw calibration data to obtain a mapping relationship between the complementary camera and the auxiliary camera; and calibrating the raw calibration data of the complementary camera using the raw calibration data of the auxiliary camera based on the mapping relationship to obtain target calibration data.
[0181] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0182] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0183] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A multi-modal multi-channel complementary visual data calibration method, comprising: Based on a pre-selected calibration object and a pre-built calibration data acquisition device, obtaining raw calibration data acquired by the complementary camera and the auxiliary camera; The pre-built calibration data acquisition device includes a complementary camera and at least one auxiliary camera; the original calibration data of the complementary camera is multi-modal multi-channel complementary visual data; According to the original calibration data, adjusting the complementary camera and the auxiliary camera to be time-aligned; The complementary camera and any auxiliary camera form a binocular camera calibration pair; for each binocular camera calibration pair, the complementary camera and the auxiliary camera are spatially aligned according to the original calibration data to obtain a mapping relationship between the complementary camera and the auxiliary camera; According to the mapping relationship, the original calibration data of the auxiliary camera is used to calibrate the original calibration data of the complementary camera to obtain target calibration data.
2. The multimodal multi-channel complementary visual data calibration method according to claim 1, wherein: The complementary camera and the auxiliary camera are arranged in the same plane with their optical axes aligned and parallel and their sensor planes parallel; the complementary camera is located in the center of the plane and the auxiliary camera is arranged around the complementary camera; the complementary camera and the auxiliary camera are connected to the same hardware trigger signal.
3. The multimodal multi-channel complementary visual data calibration method according to claim 2, wherein: The step of spatially aligning the complementary camera and the auxiliary camera according to the original calibration data to obtain a mapping relationship between the complementary camera and the auxiliary camera further includes: The complementary camera and the auxiliary camera are calibrated with a monocular camera respectively according to the original calibration data to obtain the intrinsic and extrinsic parameter matrices of the complementary camera and the auxiliary camera.
4. The multi-modal multi-channel complementary visual data calibration method according to claim 3, wherein: The auxiliary camera in the binocular camera calibration pair includes a high dynamic range camera or a high-speed camera, and the slow path of the complementary camera and the auxiliary camera are kept triggered in the same path; the complementary camera and the auxiliary camera are spatially aligned according to the original calibration data to obtain a mapping relationship between the complementary camera and the auxiliary camera, specifically including: Solving the relative translation matrix of the complementary camera and the auxiliary camera according to the intrinsic and extrinsic parameter matrix; Substitute the relative translation matrix into the pre-constructed binocular calibration equation to obtain the mapping relationship of the binocular camera calibration pair.
5. The multimodal multi-channel complementary visual data calibration method according to claim 3 or 4, wherein: The auxiliary camera in the binocular camera calibration pair includes an event camera, and the fast path of the complementary camera and the event camera are triggered in the same way; the complementary camera and the auxiliary camera are spatially aligned according to the original calibration data to obtain a mapping relationship between the complementary camera and the auxiliary camera, specifically including: Performing a projection transformation on the auxiliary camera according to the intrinsic and extrinsic parameter matrices of the complementary camera and the auxiliary camera; According to the moving calibration object, the complementary camera and the projection transformed Second original calibration data collected by the auxiliary camera; Adjusting the exposure time of the complementary camera fast path so that the complementary camera and the auxiliary camera have the same movement amplitude; Performing grayscale reconstruction according to the second original calibration data of the complementary camera and the auxiliary camera to obtain a grayscale reconstruction result; Binarizing the grayscale reconstruction result to obtain a grayscale binarization result; Compare the grayscale binarization results of the complementary camera and the auxiliary camera, and terminate the calibration if the overlap reaches a threshold value, and record the mapping relationship of the binocular camera calibration pair.
6. The multi-modal multi-channel complementary visual data calibration method according to claim 5, wherein: Performing grayscale reconstruction according to the second original calibration data of the complementary camera to obtain a grayscale reconstruction result specifically includes: Based on the Poisson equation, the relationship between the divergence and the gradient of the spatially differential field is iterated according to the second original calibration data to obtain a grayscale reconstruction result of the complementary camera.
7. The multimodal multi-channel complementary visual data calibration method according to claim 1 or 2, wherein: The adjusting, according to the original calibration data, to time align the complementary camera and the auxiliary camera specifically includes: A preset time delay is added to the auxiliary camera based on a hardware trigger signal, and the image data included in the original calibration data ensures that the starting points of the actual exposure time of the auxiliary camera and the complementary camera are completely consistent.
8. The multi-modal multi-channel complementary visual data calibration method according to claim 3, wherein: The performing monocular camera calibration on the complementary camera and the auxiliary camera respectively according to the original calibration data to obtain the internal and external parameter matrices of the complementary camera and the auxiliary camera specifically includes: Using Zhang Zhengyou calibration method, for the complementary camera and each of the auxiliary cameras, calibration objects at a preset number of preset distances are calibrated according to the original calibration data, and the internal and external parameter matrices of the complementary camera and the auxiliary camera are calculated using the maximum likelihood method.
9. The multi-modal multi-channel complementary visual data calibration method according to claim 1, wherein: The data calibration includes label addition and / or copy backup; the original calibration data of the complementary camera is calibrated with the original calibration data of the auxiliary camera according to the mapping relationship to obtain target calibration data, and then further includes: The calibration data acquisition device acquires complementary camera data and auxiliary camera data according to the calibration, and the complementary camera data, the auxiliary camera data and the mapping relationship determined by the calibration constitute a data set.
10. A multi-modal multi-channel complementary visual data calibration device, comprising: A data acquisition unit, configured to acquire original calibration data acquired by a complementary camera and an auxiliary camera based on a pre-selected calibration object and a pre-built calibration data acquisition device; the pre-built calibration data acquisition device comprises a complementary camera and at least one auxiliary camera; the original calibration data of the complementary camera is multi-modal multi-channel complementary visual data; A time alignment unit, configured to adjust the complementary camera and the auxiliary camera to be time aligned according to the original calibration data; A spatial alignment unit is used to form a binocular camera calibration pair with the complementary camera and any auxiliary camera; for each binocular camera calibration pair, according to the original calibration spatially aligning the complementary camera and the auxiliary camera using the data to obtain a mapping relationship between the complementary camera and the auxiliary camera; The target calibration unit is used to calibrate the original calibration data of the complementary camera using the original calibration data of the auxiliary camera according to the mapping relationship to obtain target calibration data.
Citation Information
Patent Citations
Fog penetrating target detection method and device based on multi-sensor fusion
CN114694011A
Calibration device and calibration method for telecentric camera
CN115526941A
Sensor fusion method based on binocular camera guidance
CN115937810A
Multi-modal multi-channel complementary visual data calibration method and device
CN117911525A
Method and system for time alignment calibration, event annotation and / or database generation
US20180308253A1
Cited By
Multi-modal multi-channel complementary visual data calibration method and device
CN117911525A