A camera calibration method and system based on deep learning

By using a deep learning-based camera calibration method, which divides the extrinsic parameter calibration group by the field of view and performs feature point matching to correct the camera extrinsic parameters, the problem of extrinsic parameter error accumulation during the use of the robot vision module is solved, and the accuracy and convenience of data processing are improved.

CN121582358BActive Publication Date: 2026-07-14HUNAN VOCATIONAL INST OF TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUNAN VOCATIONAL INST OF TECH
Filing Date
2026-01-27
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing robot vision modules suffer from extrinsic parameter errors due to camera rotation and translation during use, which affects the accuracy of data processing.

Method used

A deep learning-based camera calibration method is adopted. By obtaining the field of view of the peripheral camera, it is divided into extrinsic parameter calibration groups. Feature point matching is performed using video data from the reference camera and the camera to be calibrated to correct the extrinsic parameter data of the camera and eliminate errors.

Benefits of technology

It improves the convenience of camera calibration and the accuracy of data processing, adapts to various environments, and reduces dependence on specific laboratory environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121582358B_ABST
    Figure CN121582358B_ABST
Patent Text Reader

Abstract

The application discloses a camera calibration method and system based on deep learning, relates to the technical field of camera calibration, and solves the technical problem of low accuracy of related data processing caused by pose deviation of an existing robot vision sensor during use; comprising: acquiring a field of view range corresponding to a plurality of peripheral cameras; dividing the plurality of peripheral cameras into a plurality of extrinsic parameter calibration groups based on the field of view range; generating a reference three-dimensional map of the peripheral cameras based on original intrinsic parameter data, original extrinsic parameter data and first video data; acquiring second video data corresponding to a camera to be calibrated, performing feature point matching based on the first video data and the second video data to obtain a plurality of pairs of matched feature points, and correcting original extrinsic parameter data of the camera to be calibrated based on the matched feature points and the reference three-dimensional map to obtain current extrinsic parameter data; and the accuracy of subsequent related data processing is ensured by revising the extrinsic parameters of the camera.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of camera calibration technology, specifically a camera calibration method and system based on deep learning. Background Technology

[0002] Robots are typically equipped with multiple cameras, whose relative positions remain fixed, but whose individual angles can rotate; however, the overall position of the video acquisition module changes. Simultaneous Localization and Mapping (SLAM) refers to the technique by which a robot, in an unknown environment lacking prior information, estimates its own position and orientation using its own sensors and motion capabilities. With the deepening of theoretical research and the rapid development of sensor technology, multi-sensor fusion has become an important research direction in SLAM technology. When data from different sensors are used collaboratively in the same system, effective processing is essential. The acquisition of environmental position information depends on the coordinate system of each sensor; therefore, multi-sensor fusion first needs to address the problem of data fusion and unification. The reliability of the data directly determines the success or failure of the SLAM process; therefore, high accuracy of sensor data is crucial for achieving high-precision localization and mapping. The accuracy of pose estimation through the fusion of data from different sensors heavily depends on the extrinsic parameter calibration between the sensors. Extrinsic parameter calibration refers to the pose transformation relationship between different sensors.

[0003] Existing methods for calibrating robot vision modules often involve performing only one calibration of extrinsic and intrinsic parameters in a laboratory setting. During actual use, the camera of the vision module is generally not calibrated. However, because robot cameras undergo rotation and translation during use, each rotation or translation causes minor wear on the relevant mechanical parts. Over time, this leads to errors in the original camera extrinsic parameters, specifically errors in the relative poses between the cameras. This affects the accuracy of subsequent data processing that relies on sensor extrinsic parameters. Therefore, a camera calibration method and system based on deep learning is needed. Summary of the Invention

[0004] This application provides a camera calibration method and system based on deep learning, which solves the technical problem that pose deviations caused by existing robot vision sensors during use lead to low accuracy in related data processing.

[0005] To achieve the above objectives, this application adopts the following technical solution: Firstly, a deep learning-based camera calibration method is provided, including: Obtain the field of view corresponding to several peripheral cameras; the field of view is the entire field of view that the peripheral cameras can obtain during rotation and translation; it is worth noting that the relative positions of each peripheral camera in this embodiment are fixed, but the field of view angle of the peripheral cameras will change; based on the field of view, the several peripheral cameras are divided into several external parameter calibration groups; the external parameter calibration group includes a reference camera and several cameras to be calibrated. Acquire the original intrinsic and extrinsic parameter data of the reference camera, as well as the corresponding first video data; generate a reference 3D map of the peripheral camera based on the original intrinsic and extrinsic parameter data and the first video data; Acquire the second video data corresponding to the camera to be calibrated, perform feature point matching based on the first and second video data to obtain several pairs of matching feature points, and correct the original extrinsic parameter data of the camera to be calibrated based on the matching feature points and the reference 3D map to obtain the current extrinsic parameter data.

[0006] Based on the above technical solution, in the camera calibration method and system based on deep learning provided in this application, the following steps are taken: First, the field of view of several peripheral cameras is obtained. Then, the peripheral cameras are divided into several extrinsic parameter calibration groups based on the field of view. Each extrinsic parameter calibration group includes a reference camera and several cameras to be calibrated. The original intrinsic parameter data and original extrinsic parameter data of the reference camera, as well as the corresponding first video data, are obtained. A reference 3D map of the peripheral camera is generated based on the original intrinsic parameter data, the original extrinsic parameter data, and the first video data. Second video data of the corresponding camera to be calibrated is obtained. Feature point matching is performed based on the first video data and the second video data to obtain several pairs of matching feature points. The original extrinsic parameter data of the camera to be calibrated is corrected based on the matching feature points and the reference 3D map to obtain the current extrinsic parameter data. By revising the camera's extrinsic parameters, i.e., correcting the errors in the relative pose between cameras caused by camera rotation and movement wear, the accuracy of subsequent related data processing is ensured.

[0007] Meanwhile, the above-mentioned method for correcting extrinsic parameters can adapt to most environments, avoiding the limitation of calibration only in specific laboratory environments; the ability to correct the extrinsic parameters of each vision camera during robot movement greatly improves the convenience of camera calibration.

[0008] In conjunction with the first aspect above, in one possible implementation, the division of several peripheral cameras into several external parameter calibration groups based on the field of view includes: S1: Obtain the field of view of each peripheral camera, select any peripheral camera as the reference camera, and calculate the number of peripheral cameras whose field of view overlaps with the field of view of the reference camera. S2: Select the reference camera with the most peripheral cameras as the reference camera, and set the peripheral cameras whose field of view overlaps with the reference camera as the cameras to be calibrated corresponding to the reference camera; and divide the reference camera and the cameras to be calibrated into an external parameter calibration group; determine whether there are any unassigned peripheral cameras; if yes, proceed to S3; if no, proceed to S4. S3: Remove the reference camera, select any other camera to be calibrated or an unassigned peripheral camera as the proposed reference camera, and calculate the number of unassigned peripheral cameras whose field of view overlaps with the field of view of the proposed reference camera; proceed to S2. S4: Output each external parameter calibration group.

[0009] In conjunction with the first aspect above, one possible implementation also includes setting the number of the external parameter calibration group based on the external parameter calibration group; Extract the reference camera from each external parameter calibration group; When the reference camera is not the camera to be calibrated in other external parameter calibration groups, the number of the external parameter calibration group is set to 1; When the reference camera is a camera to be calibrated in another external parameter calibration group, the number of the external parameter calibration group is set to N+1; where N is the number of the external parameter calibration group corresponding to the camera to be calibrated. The corresponding number in the external parameter calibration group indicates the execution priority of the external parameter calibration group. External parameter calibration groups with the same number have the same execution priority, and the external parameter calibration group with the number 1 has the highest execution priority.

[0010] In conjunction with the first aspect above, in one possible implementation, generating the reference 3D map of the peripheral camera based on the original intrinsic parameter data, the original extrinsic parameter data, and the first video data includes: The first video data is frame extracted to obtain several usable frame images, each containing several feature points. The usable frame images are then integrated into the first inter-frame motion data according to the extraction order. Extract several usable frame images from the first inter-frame motion data; preprocess the usable frame images based on the original intrinsic parameter data to obtain a reference image; Extract several feature points from the reference image, and convert these feature points into three-dimensional point coordinates based on the original extrinsic data; sequentially obtain the three-dimensional point coordinates corresponding to each feature point in each reference image, and construct a reference three-dimensional map based on the three-dimensional point coordinates.

[0011] In conjunction with the first aspect above, in one possible implementation, the feature point matching based on the first video data and the second video data includes: The first video data is frame extracted to obtain several usable frame images. The original intrinsic parameter data corresponding to the reference camera is obtained. The usable frame images are preprocessed based on the original intrinsic parameter data to obtain a reference image, which is recorded as the reference frame image. The second video data is frame extracted to obtain several usable frame images. The original intrinsic parameter data corresponding to the camera to be calibrated is obtained. The usable frame images are preprocessed based on the original intrinsic parameter data to obtain a reference image, which is recorded as the query frame image. The reference frame image and the query frame image that have overlapping regions are recorded as a pair of matched images; A pair of matching images is obtained, and several feature points of the reference frame image and several feature points of the query frame image in the overlapping area are extracted from the matching images. Based on the several feature points in the reference frame image and the query frame image, several pairs of matching feature points are obtained. The matching feature points include the reference feature points and the query feature points.

[0012] In conjunction with the first aspect above, in one possible implementation, the matching of several feature points in the reference frame image and the query frame image to obtain several pairs of matching feature points includes: Obtain the first feature descriptor of each feature point corresponding to the query frame image; and the second feature descriptor of each feature point corresponding to the reference frame image; Calculate the Euclidean distance between each first feature descriptor and each second feature descriptor. Record the two feature points corresponding to the first feature descriptor and the second feature descriptor whose Euclidean distance is less than a set similarity Euclidean distance threshold as a pair of matching feature points. The feature point corresponding to the first feature descriptor is recorded as the query feature point, and the feature point corresponding to the second feature descriptor is recorded as the reference feature point.

[0013] In conjunction with the first aspect above, in one possible implementation, the frame extraction includes: The video data is processed to extract frame images to obtain a sequence of initial images. The initial image sequence is then input into a dynamic object recognition model to obtain a sequence of labeled images, which includes a number of labeled images. When the label of a marked image is completely static, the marked image is recorded as a usable frame image; When the label of the labeled image is not fully static, the dynamic recognition box of the labeled image is extracted, a mask image is constructed based on the dynamic recognition box, and the mask image and the labeled image are superimposed to obtain a usable frame image. Extract several feature points from each available frame image and mark the feature points in the available frame images.

[0014] Specifically, a binary mask image is constructed based on the dynamic recognition box, that is, the pixel value corresponding to each pixel point inside the dynamic recognition box is set to 255, and the pixel value corresponding to each pixel point outside the dynamic recognition box is set to 0; the mask image and the marker image are superimposed to obtain a usable frame image with the corresponding area of ​​the dynamic recognition box removed; thus avoiding the influence of dynamic objects in the image on subsequent 3D modeling.

[0015] In conjunction with the first aspect above, in one possible implementation, the dynamic object recognition model includes: A series of initial image sequences and corresponding labeled image sequences are obtained. The labeled image sequences include several labeled images. The labeled images, which are not entirely static, contain several dynamic bounding boxes. The labels on the labeled images are set by experts. When a labeled image contains a moving object, such as an animal or a falling leaf, the corresponding label is set to non-fully static; when a labeled image does not contain a moving object, the corresponding label is set to fully static. The dynamic bounding boxes are manually calibrated by experts and are the borders corresponding to the positions of dynamic objects in each initial image sequence. The initial images and their corresponding dynamic bounding boxes in the initial image sequences are integrated into several training and validation datasets. The AI ​​model is trained using training data and tested using validation data, ultimately resulting in a dynamic object recognition model that takes an initial image sequence as input and outputs a labeled image sequence.

[0016] In conjunction with the first aspect above, in one possible implementation, the step of correcting the original extrinsic parameter data of the camera to be calibrated based on matching feature points and a reference 3D map to obtain the current extrinsic parameter data includes: Extract reference feature points and query feature points from each matching feature point; obtain the 3D point coordinates of the reference feature points in the reference 3D map, and denot them as reference 3D point coordinates; The theoretical extrinsic data is obtained by solving the extrinsic data of the camera to be calibrated based on the coordinates of several reference 3D points and their corresponding feature points. The original extrinsic data is then corrected using the theoretical extrinsic data to obtain the current extrinsic data.

[0017] Secondly, this application provides a camera calibration system based on deep learning, including: a data acquisition module, a camera group division module, and an extrinsic parameter calibration module; The data acquisition module includes a video data acquisition unit and a camera data acquisition unit; The video data acquisition unit is used to acquire video data from various peripheral cameras. The camera data acquisition unit is used to acquire the camera's field of view, raw intrinsic parameter data, and raw extrinsic parameter data; The camera group division module is used to obtain the field of view of several peripheral cameras; based on the field of view, the several peripheral cameras are divided into several external parameter calibration groups; each external parameter calibration group includes a reference camera and several cameras to be calibrated. The extrinsic parameter calibration module includes a matching feature point generation unit and an extrinsic parameter revision unit; The matching feature point generation unit is used to acquire the original intrinsic parameter data and original extrinsic parameter data of the reference camera, as well as the corresponding first video data; generate a reference 3D map of the peripheral camera based on the original intrinsic parameter data, original extrinsic parameter data and the first video data; acquire the second video data of the corresponding camera to be calibrated; and perform feature point matching based on the first video data and the second video data to obtain several pairs of matching feature points. The extrinsic parameter revision unit is used to correct the original extrinsic parameter data of the camera to be calibrated based on matching feature points and a reference 3D map, so as to obtain the current extrinsic parameter data.

[0018] This application provides a camera calibration method and system based on deep learning, which can acquire the field of view of several peripheral cameras; divide the peripheral cameras into several extrinsic parameter calibration groups based on the field of view; each extrinsic parameter calibration group includes a reference camera and several cameras to be calibrated; acquire the original intrinsic parameter data and original extrinsic parameter data of the reference camera, as well as the corresponding first video data; generate a reference 3D map of the peripheral camera based on the original intrinsic parameter data, original extrinsic parameter data, and first video data; acquire the second video data of the corresponding camera to be calibrated; perform feature point matching based on the first video data and the second video data to obtain several pairs of matching feature points; and correct the original extrinsic parameter data of the camera to be calibrated based on the matching feature points and the reference 3D map to obtain the current extrinsic parameter data; by revising the extrinsic parameters of the camera, that is, correcting the errors in the relative pose between the cameras caused by the turning and movement wear of each camera, the accuracy of subsequent related data processing is ensured.

[0019] It should be understood that the descriptions of technical features, technical solutions, beneficial effects, or similar language in this application do not imply that all features and advantages can be achieved in any single embodiment. Rather, it is understood that the description of a feature or beneficial effect means that a specific technical feature, technical solution, or beneficial effect is included in at least one embodiment. Therefore, the descriptions of technical features, technical solutions, or beneficial effects in this specification do not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions, and beneficial effects described in this embodiment can be combined in any suitable manner. Those skilled in the art will understand that embodiments can be implemented without one or more specific technical features, technical solutions, or beneficial effects of a particular embodiment. In other embodiments, additional technical features and beneficial effects may be identified in specific embodiments that do not embody all embodiments. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a schematic diagram illustrating the steps of the camera calibration method in this application; Figure 2 This is a schematic diagram of the module connections of the camera calibration system in this application. Detailed Implementation

[0022] The technical solutions of this application will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0023] Please see Figure 1 The first aspect of this application provides a camera calibration method based on deep learning, including: Obtain the field of view of several peripheral cameras; the field of view is the entire field of view that the peripheral cameras can obtain during rotation and translation; it is worth noting that the relative positions of each peripheral camera in this embodiment are fixed, but the field of view angle of the peripheral cameras will change; based on the field of view, the several peripheral cameras are divided into several external parameter calibration groups; the external parameter calibration group includes a reference camera and several cameras to be calibrated. Acquire the original intrinsic and extrinsic parameter data of the reference camera, as well as the corresponding first video data; generate a reference 3D map of the peripheral camera based on the original intrinsic and extrinsic parameter data and the first video data; Acquire the second video data corresponding to the camera to be calibrated, perform feature point matching based on the first and second video data to obtain several pairs of matching feature points, and correct the original extrinsic parameter data of the camera to be calibrated based on the matching feature points and the reference 3D map to obtain the current extrinsic parameter data.

[0024] Based on the above technical solution, in the camera calibration method and system based on deep learning provided in this application, the following steps are taken: First, the field of view of several peripheral cameras is obtained. Then, the peripheral cameras are divided into several extrinsic parameter calibration groups based on the field of view. Each extrinsic parameter calibration group includes a reference camera and several cameras to be calibrated. The original intrinsic parameter data and original extrinsic parameter data of the reference camera, as well as the corresponding first video data, are obtained. A reference 3D map of the peripheral camera is generated based on the original intrinsic parameter data, the original extrinsic parameter data, and the first video data. Second video data of the corresponding camera to be calibrated is obtained. Feature point matching is performed based on the first video data and the second video data to obtain several pairs of matching feature points. The original extrinsic parameter data of the camera to be calibrated is corrected based on the matching feature points and the reference 3D map to obtain the current extrinsic parameter data. By revising the camera's extrinsic parameters, i.e., correcting the errors in the relative pose between cameras caused by camera rotation and movement wear, the accuracy of subsequent related data processing is ensured.

[0025] Meanwhile, the above-mentioned method for correcting extrinsic parameters can adapt to most environments, avoiding the limitation of calibration only in specific laboratory environments; the ability to correct the extrinsic parameters of each vision camera during robot movement greatly improves the convenience of camera calibration.

[0026] In one possible implementation, several peripheral cameras are divided into several external parameter calibration groups based on their field of view, including: S1: Obtain the field of view of each peripheral camera, select any peripheral camera as the reference camera, and calculate the number of peripheral cameras whose field of view overlaps with that of the reference camera. It is understood that the relative positions of different cameras are fixed, but the shooting angles of different cameras can be rotated. The field of view can be understood as the range of angles that a peripheral camera can shoot at. An overlapping field of view refers to an overlapping range of angles between two peripheral cameras. For example, if peripheral camera A and peripheral camera B can both shoot the reference object by rotating their camera angles, then peripheral camera A and peripheral camera B have overlapping field of view. In this embodiment, the reference object is a calibration board, such as a checkerboard calibration board. It is understood that the reference object can be any object other than the calibration board. This embodiment only uses a calibration board as an example. S2: Select the reference camera with the largest number of peripheral cameras as the reference camera, and set the peripheral cameras whose field of view overlaps with the reference camera as the cameras to be calibrated corresponding to the reference camera; and divide the reference camera and the cameras to be calibrated into an external parameter calibration group; determine whether there are any unassigned peripheral cameras; if yes, proceed to S3; if no, proceed to S4; it can be understood that when there are multiple reference cameras with the largest number and the same number, any one of the reference cameras can be selected as the reference camera; S3: Remove the reference camera, select any other camera to be calibrated or an unassigned peripheral camera as the proposed reference camera, and calculate the number of unassigned peripheral cameras whose field of view overlaps with the field of view of the proposed reference camera; proceed to S2; it can be understood that the peripheral cameras here do not include the camera to be calibrated. For example, if a camera A has 4 cameras whose field of view overlaps with it, and 3 of these cameras have been marked as cameras to be calibrated, then the number of peripheral cameras whose field of view overlaps with the field of view of the proposed reference camera is 1. S4: Output each external parameter calibration group.

[0027] In this embodiment, several peripheral cameras are divided into several external parameter calibration groups in the manner described above, so as to minimize the number of external parameter calibration groups and thus reduce the amount of data in subsequent data processing.

[0028] In one possible implementation, the method further includes setting the number of the extrinsic calibration group based on the extrinsic calibration group; and extracting the reference camera in each extrinsic calibration group. When the reference camera is not the camera to be calibrated in other external parameter calibration groups, the number of the external parameter calibration group is set to 1; When the reference camera is a camera to be calibrated in another external parameter calibration group, the number of the external parameter calibration group is set to N+1; where N is the number of the external parameter calibration group corresponding to the camera to be calibrated. The corresponding number in the external parameter calibration group indicates the execution priority of the external parameter calibration group. External parameter calibration groups with the same number have the same execution priority, and the external parameter calibration group with the number 1 has the highest execution priority.

[0029] This embodiment sets the execution priority of the peripheral calibration group, so that the process of correcting the extrinsic parameters of the camera to be calibrated can be carried out in an orderly manner. For example, the extrinsic calibration group numbered 1 is executed first to correct the extrinsic parameters of the camera to be calibrated in it. Then the extrinsic parameters of the camera to be calibrated in the extrinsic calibration group numbered 2 are corrected. At this time, the extrinsic parameters of the reference camera in the extrinsic calibration group numbered 2 have been corrected. Therefore, this method can also increase the accuracy of the subsequent extrinsic parameter revision of the peripheral camera.

[0030] In one possible implementation, generating a reference 3D map of the peripheral camera based on the original intrinsic parameter data, the original extrinsic parameter data, and the first video data includes: The first video data is frame extracted to obtain several usable frame images, each containing several feature points. The usable frame images are then integrated into the first inter-frame motion data according to the extraction order. Extract several usable frame images from the first inter-frame motion data; preprocess the usable frame images based on the original intrinsic parameter data to obtain a reference image; Extract several feature points from the reference image, and convert these feature points into three-dimensional point coordinates based on the original extrinsic data; sequentially obtain the three-dimensional point coordinates corresponding to each feature point in each reference image, and construct a reference three-dimensional map based on the three-dimensional point coordinates.

[0031] Specifically, the reference image is obtained by preprocessing the available frame image based on the original intrinsic parameter data, including: The raw intrinsic parameter data is obtained from the camera's factory calibration or prior calibration, such as the Zhang Zhengyou calibration method. The core parameters include the intrinsic parameter matrix K and the distortion coefficients. , , , and ; ; in, , ; Physical focal length and Pixel size; and Principal point coordinates, i.e., the intersection of the optical axis and the image plane, which is the ideal image center; This is the pixel tilt coefficient, which is usually set to 0; Distortion correction maps the pixel coordinates of the distorted image back to undistorted coordinates through radial and tangential distortion correction; normalization is then applied to each pixel coordinate to obtain normalized pixel coordinates. Distortion correction is performed using normalized pixel coordinates; Radial distortion correction: ; ; ; in, The distance from the origin to the normalized coordinates. , and The radial distortion coefficient is... Dominant first-order distortion, and Compensation for higher-order distortions; These are the pixel coordinates after radial distortion correction. Tangential distortion correction: ; ; in, These are the pixel coordinates after tangential distortion correction. and The tangential distortion coefficient; The corrected pixel coordinates are restored to their original pixel coordinates to obtain the coordinates. This is the inverse normalization operation; This embodiment eliminates lens optical distortion and unifies image coordinate mapping rules through the above steps, providing distortion-free and standardized image data for feature point extraction and avoiding feature point positioning deviations caused by differences in intrinsic parameters. Based on the original extrinsic parameter data, several feature points are converted into three-dimensional point coordinates, including: The raw extrinsic parameter data describes the camera's pose in the world coordinate system, including the rotation matrix. Translation vector The following transformation relationships can be used to obtain the 3D point coordinates corresponding to each pixel coordinate; ; ; in, As a scaling factor, For the camera's 3D point coordinates, The final obtained world 3D point coordinates are the 3D point coordinates used in this embodiment to construct the reference 3D map; In one possible implementation, feature point matching based on the first video data and the second video data includes: extracting frames from the first video data to obtain several usable frame images, obtaining the original intrinsic parameter data corresponding to the reference camera, preprocessing the usable frame images based on the original intrinsic parameter data to obtain a reference image, and recording it as the reference frame image. The second video data is frame extracted to obtain several usable frame images. The original intrinsic parameter data corresponding to the camera to be calibrated is obtained. The usable frame images are preprocessed based on the original intrinsic parameter data to obtain a reference image, which is recorded as the query frame image. The reference frame image and the query frame image with overlapping areas are recorded as a pair of matched images. It can be understood that the overlapping area refers to the image in which part of the area is captured from the same reference object. In this embodiment, the overlapping area refers to the area in which both the reference frame image and the query frame image contain the calibration plate. A pair of matching images is obtained, and several feature points of the reference frame image and several feature points of the query frame image in the overlapping area are extracted from the matching images. Based on the several feature points in the reference frame image and the query frame image, several pairs of matching feature points are obtained. The matching feature points include the reference feature points and the query feature points.

[0032] In one possible implementation, several feature points in the reference frame image and the query frame image are matched to obtain several pairs of matched feature points, including: obtaining a first feature descriptor corresponding to each feature point in the query frame image; and a second feature descriptor corresponding to each feature point in the reference frame image; Calculate the Euclidean distance between each first feature descriptor and each second feature descriptor. Record the two feature points corresponding to the first feature descriptor and the second feature descriptor whose Euclidean distance is less than a set similarity Euclidean distance threshold as a pair of matching feature points. The feature point corresponding to the first feature descriptor is recorded as the query feature point, and the feature point corresponding to the second feature descriptor is recorded as the reference feature point.

[0033] In one possible implementation, frame extraction includes: extracting frame images from video data to obtain a sequence of initial frames; inputting the initial image sequence into a dynamic object recognition model to obtain a sequence of labeled images, wherein the sequence of labeled images includes a number of labeled images. When the label of a marked image is completely static, the marked image is recorded as a usable frame image; When the label of the labeled image is not fully static, the dynamic recognition box of the labeled image is extracted, a mask image is constructed based on the dynamic recognition box, and the mask image and the labeled image are superimposed to obtain a usable frame image. Extract several feature points from each available frame image and mark the feature points in the available frame images.

[0034] Specifically, a binary mask image is constructed based on the dynamic recognition box, that is, the pixel value corresponding to each pixel point inside the dynamic recognition box is set to 255, and the pixel value corresponding to each pixel point outside the dynamic recognition box is set to 0; the mask image and the marker image are superimposed to obtain a usable frame image with the corresponding region of the dynamic recognition box removed; thus avoiding the influence of dynamic objects in the image on subsequent 3D modeling. It is understandable that feature point extraction of images is a relatively existing technology that can be performed by training a model. A feature point is a point that represents the features of an image, and an image can have several feature points. These feature points can be the contour points of objects in the image, etc. The feature points also include their corresponding feature descriptors, which are used to identify the local features corresponding to the feature points. The feature descriptors are represented in binary form, such as the gray-level gradient corresponding to a local area of ​​a feature point; they can be understood as the identity card of the feature points. This embodiment processes the original image using a binary mask image to remove dynamic objects and prevent them from affecting subsequent calculations.

[0035] In one possible implementation, the dynamic object recognition model includes: acquiring several initial image sequences and corresponding labeled image sequences; the labeled image sequences include several labeled images; the labeled images are non-fully static and contain several dynamic recognition boxes; the labels of the labeled images are set by experts, and when there are moving objects in the labeled images, such as animals or falling leaves, the corresponding labels are set to non-fully static; when there are no moving objects in the labeled images, the corresponding labels are set to fully static; the dynamic recognition boxes are manually calibrated by experts, and the dynamic recognition boxes are the bounding boxes corresponding to the positions of dynamic objects in each initial image in the initial image sequence; and the initial images and corresponding dynamic recognition boxes in the initial image sequence are integrated into several training data and validation data. The AI ​​model is trained using training data and tested using validation data, ultimately resulting in a dynamic object recognition model that takes an initial image sequence as input and outputs a sequence of labeled images. It is understood that the models used to recognize dynamic objects and to annotate them are already relatively existing technologies; these will not be elaborated upon here. This embodiment uses other models or methods that can label moving objects in each image of an image sequence to replace the function of the dynamic object recognition model.

[0036] In one possible implementation, the original extrinsic data of the camera to be calibrated is corrected based on the matching feature points and the reference 3D map to obtain the current extrinsic data, including: extracting reference feature points and query feature points from each matching feature point; obtaining the 3D point coordinates corresponding to the reference feature points in the reference 3D map, denoted as the reference 3D point coordinates; The theoretical extrinsic data is obtained by solving the extrinsic data of the camera to be calibrated based on the coordinates of several reference 3D points and their corresponding feature points. The original extrinsic data is then corrected using the theoretical extrinsic data to obtain the current extrinsic data.

[0037] Specifically, the query feature points are extracted from several pairs of matching feature points, and the pixel coordinates corresponding to the query feature points are converted into camera 3D point coordinates: ; in, To query the pixel coordinates corresponding to the feature point; To query the camera's 3D point coordinates after feature point transformation; Obtain the camera 3D point coordinates and reference 3D point coordinates of the query feature point from a set of matching feature points, and substitute them into the following formula: ; The rotation matrix in the theoretical extrinsic parameter data is obtained by solving the problem. Translation vector The number of comprehensible feature matching points should be at least 4 pairs. It is worth noting that when there are many matching feature points, the matching feature points can be arbitrarily matched to obtain several groups of solution feature points consisting of 4 matching feature points. The theoretical extrinsic parameters of the camera are then solved based on each solution feature point group to obtain several sets of theoretical extrinsic parameter data. Any set of theoretical extrinsic parameter data is selected as the extrinsic parameter data. The coordinates of the reference 3D points in the remaining matching feature points are substituted into the above formula to obtain the reverse camera 3D coordinate points. The Euclidean distance between these camera 3D coordinates and the pixel coordinates corresponding to the queried feature points, converted to camera 3D point coordinates, is calculated and denoted as the transformation difference distance corresponding to the reference 3D point coordinates. The average of the transformation difference distances corresponding to each reference 3D point coordinate is calculated to obtain the total transformation difference distance corresponding to the solution feature point group. The theoretical extrinsic parameter data corresponding to the solution feature point group with the smallest total transformation difference distance is taken as the final theoretical extrinsic parameter data.

[0038] Please see Figure 2 Secondly, this application provides a camera calibration system based on deep learning, including: a data acquisition module, a camera group division module, and an extrinsic parameter calibration module; The data acquisition module includes a video data acquisition unit and a camera data acquisition unit; The video data acquisition unit is used to acquire video data captured by various peripheral cameras; The camera data acquisition unit is used to acquire the camera's field of view, raw intrinsic parameter data, and raw extrinsic parameter data; The camera group division module is used to obtain the field of view of several peripheral cameras; based on the field of view, the several peripheral cameras are divided into several external parameter calibration groups; each external parameter calibration group includes a reference camera and several cameras to be calibrated. The extrinsic calibration module includes a matching feature point generation unit and an extrinsic revision unit; The matching feature point generation unit is used to acquire the original intrinsic parameter data and original extrinsic parameter data of the reference camera, as well as the corresponding first video data; generate a reference 3D map of the peripheral camera based on the original intrinsic parameter data, original extrinsic parameter data and the first video data; acquire the second video data of the corresponding camera to be calibrated; and perform feature point matching based on the first video data and the second video data to obtain several pairs of matching feature points. The extrinsic parameter revision unit is used to correct the original extrinsic parameter data of the camera to be calibrated based on the matching feature points and the reference 3D map, so as to obtain the current extrinsic parameter data.

[0039] Some of the data in the above formula are calculated by removing dimensions and taking their numerical values. The formula is the closest to the real situation obtained by software simulation of a large amount of collected data. The preset parameters and preset thresholds in the formula are set by those skilled in the art according to the actual situation or obtained through simulation of a large amount of data.

[0040] The above embodiments are only used to illustrate the technical methods of this application and are not intended to limit it. Although this application has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of this application without departing from the spirit and scope of the technical methods of this application.

Claims

1. A camera calibration method based on deep learning, characterized in that, include: Obtain the field of view of several peripheral cameras; Based on the field of view, several peripheral cameras are divided into several external parameter calibration groups; including: S1: Obtain the field of view of each peripheral camera, select any peripheral camera as the reference camera, and calculate the number of peripheral cameras whose field of view overlaps with the field of view of the reference camera. S2: Select the reference camera with the most peripheral cameras as the reference camera, and set the peripheral cameras whose field of view overlaps with the reference camera as the cameras to be calibrated corresponding to the reference camera; and divide the reference camera and the cameras to be calibrated into an external parameter calibration group; determine whether there are any unassigned peripheral cameras; if yes, proceed to S3; if no, proceed to S4. S3: Remove the reference camera, select any other camera to be calibrated or an unassigned peripheral camera as the proposed reference camera, and calculate the number of unassigned peripheral cameras whose field of view overlaps with the field of view of the proposed reference camera; proceed to S2. S4: Output the calibration groups of each external parameter; The external parameter calibration group is numbered based on the external parameter calibration group; Extract the reference camera from each external parameter calibration group; When the reference camera is not the camera to be calibrated in other external parameter calibration groups, the number of the external parameter calibration group is set to 1; When the reference camera is a camera to be calibrated in another external parameter calibration group, the number of the external parameter calibration group is set to N+1; where N is the number of the external parameter calibration group corresponding to the camera to be calibrated. The corresponding number in the external parameter calibration group indicates the execution priority of the external parameter calibration group. External parameter calibration groups with the same number have the same execution priority, and the external parameter calibration group with the number 1 has the highest execution priority. Acquire the original intrinsic and extrinsic parameter data of the reference camera, as well as the corresponding first video data; generate a reference 3D map of the peripheral camera based on the original intrinsic and extrinsic parameter data and the first video data; Acquire the second video data corresponding to the camera to be calibrated, and perform feature point matching based on the first and second video data to obtain several pairs of matched feature points; including: The first video data is frame extracted to obtain several usable frame images. The original intrinsic parameter data corresponding to the reference camera is obtained. The usable frame images are preprocessed based on the original intrinsic parameter data to obtain a reference image, which is recorded as the reference frame image. The second video data is frame extracted to obtain several usable frame images. The original intrinsic parameter data corresponding to the camera to be calibrated is obtained. The usable frame images are preprocessed based on the original intrinsic parameter data to obtain a reference image, which is recorded as the query frame image. The reference frame image and the query frame image that have overlapping regions are recorded as a pair of matched images; A pair of matching images is obtained, and several feature points of the reference frame image and several feature points of the query frame image in the overlapping area are extracted from the matching images. Based on the several feature points in the reference frame image and the query frame image, several pairs of matching feature points are obtained. The matching feature points include the reference feature points and the query feature points. The frame extraction includes: The video data is processed to extract frame images to obtain a sequence of initial images. The initial image sequence is then input into a dynamic object recognition model to obtain a sequence of labeled images, which includes a number of labeled images. When the label of a marked image is completely static, the marked image is recorded as a usable frame image; When the label of the labeled image is not fully static, the dynamic recognition box of the labeled image is extracted, a mask image is constructed based on the dynamic recognition box, and the mask image and the labeled image are superimposed to obtain a usable frame image. Extract several feature points from each available frame image and mark the feature points in the available frame images; Based on the matching feature points and the reference 3D map, the original extrinsic parameter data of the camera to be calibrated is corrected to obtain the current extrinsic parameter data, including: Extract reference feature points and query feature points from each matching feature point; obtain the 3D point coordinates of the reference feature points in the reference 3D map, and denot them as reference 3D point coordinates; The theoretical extrinsic data is obtained by solving the extrinsic data of the camera to be calibrated based on the coordinates of several reference 3D points and their corresponding feature points. The original extrinsic data is then corrected using the theoretical extrinsic data to obtain the current extrinsic data. The solution process includes: arbitrarily matching the matching feature points to obtain several sets of solution feature point groups consisting of four matching feature points; solving the theoretical extrinsic parameter data of the camera based on each solution feature point group to obtain several sets of theoretical extrinsic parameter data; calculating the total transformation difference distance corresponding to each solution feature point group; taking the theoretical extrinsic parameter data corresponding to the solution feature point group with the smallest total transformation difference distance as the final theoretical extrinsic parameter data; and using the final theoretical extrinsic parameter data to correct the original extrinsic parameter data to obtain the current extrinsic parameter data.

2. The camera calibration method based on deep learning according to claim 1, characterized in that, The process of generating a reference 3D map of the peripheral camera based on the original intrinsic parameter data, the original extrinsic parameter data, and the first video data includes: The first video data is frame extracted to obtain several usable frame images, each containing several feature points. The usable frame images are then integrated into the first inter-frame motion data according to the extraction order. Extract several usable frame images from the first inter-frame motion data; preprocess the usable frame images based on the original intrinsic parameter data to obtain a reference image; Extract several feature points from the reference image, and convert these feature points into three-dimensional point coordinates based on the original extrinsic data; sequentially obtain the three-dimensional point coordinates corresponding to each feature point in each reference image, and construct a reference three-dimensional map based on the three-dimensional point coordinates.

3. The camera calibration method based on deep learning according to claim 1, characterized in that, The matching of several feature points in the reference frame image and the query frame image yields several pairs of matching feature points, including: Obtain the first feature descriptor of each feature point corresponding to the query frame image; and the second feature descriptor of each feature point corresponding to the reference frame image; Calculate the Euclidean distance between each first feature descriptor and each second feature descriptor. Record the two feature points corresponding to the first feature descriptor and the second feature descriptor whose Euclidean distance is less than a set similarity Euclidean distance threshold as a pair of matching feature points. The feature point corresponding to the first feature descriptor is recorded as the query feature point, and the feature point corresponding to the second feature descriptor is recorded as the reference feature point.

4. The camera calibration method based on deep learning according to claim 1, characterized in that, The dynamic object recognition model includes: Acquire several initial image sequences and corresponding labeled image sequences; the labeled image sequences include several labeled images; the labeled images are not entirely static and contain several dynamic bounding boxes; integrate each initial image and its corresponding dynamic bounding box in the initial image sequences into several training data and validation data; The AI ​​model is trained using training data and tested using validation data, ultimately resulting in a dynamic object recognition model that takes an initial image sequence as input and outputs a labeled image sequence.

5. A camera calibration system based on deep learning, based on the application of a camera calibration method based on deep learning as described in any one of claims 1 to 4; characterized in that, include: Data acquisition module, camera group division module, and extrinsic parameter calibration module; The data acquisition module includes a video data acquisition unit and a camera data acquisition unit; The video data acquisition unit is used to acquire video data from various peripheral cameras. The camera data acquisition unit is used to acquire the camera's field of view, raw intrinsic parameter data, and raw extrinsic parameter data; The camera group division module is used to obtain the field of view of several peripheral cameras; based on the field of view, the several peripheral cameras are divided into several external parameter calibration groups; each external parameter calibration group includes a reference camera and several cameras to be calibrated. The extrinsic parameter calibration module includes a matching feature point generation unit and an extrinsic parameter revision unit; The matching feature point generation unit is used to obtain the original intrinsic parameter data and original extrinsic parameter data of the reference camera, as well as the corresponding first video data; A reference 3D map of the peripheral camera is generated based on the original intrinsic parameter data, the original extrinsic parameter data, and the first video data; Acquire the second video data of the corresponding camera to be calibrated, and perform feature point matching based on the first video data and the second video data to obtain several pairs of matching feature points; The extrinsic parameter revision unit is used to correct the original extrinsic parameter data of the camera to be calibrated based on matching feature points and a reference 3D map, so as to obtain the current extrinsic parameter data.