Dual-camera calibration method, system and related device

By automatically selecting the image with the highest matching reliability among the dual cameras, and calculating and converting the image offset value, automatic calibration of the dual cameras is achieved, solving the problem of low manual calibration accuracy and improving calibration accuracy.

CN119942158APending Publication Date: 2025-05-06BEIJING LIZHENG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510152353.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing dual camera calibration method relies on manual adjustment, and there are human errors, resulting in low calibration accuracy.

Method used

By acquiring the captured images of the first camera and the second camera, selecting the target image with the highest matching reliability, calculating the image offset value, and converting it into the camera angle offset value, the second camera is automatically adjusted.

Benefits of technology

Improves the accuracy of dual camera calibration, reduces human error, and realizes automated calibration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942158A_ABST
    Figure CN119942158A_ABST
Patent Text Reader

Abstract

The invention discloses a dual-camera calibration method and system and a related device, and relates to the field of unmanned aerial vehicles, and the method comprises the steps: obtaining a first shot image of a first camera and a second shot image set of a second camera; selecting a matched target second shot image from the second shot image set; an image offset value between the first shot image and the target second shot image is calculated according to a feature point offset set, the feature point offset set comprises an offset value of at least one feature point pair, and each feature point pair comprises a feature point in the first shot image and a corresponding feature point in the target second shot image; and converting the image deviation value into a camera angle deviation value, taking the first camera as a reference object, and adjusting the second camera according to the camera angle deviation value. According to the method, the feature points of the shot image are taken as the calibration reference object, errors of the calibration area are avoided, the camera is automatically adjusted, errors of manual adjustment are avoided, and therefore the calibration accuracy is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of UAV technology, and in particular to a dual-camera calibration method, system and related devices. Background Art

[0002] In order to improve the real-time and accuracy of drone monitoring, a dual-camera system with two cameras, one above and one below, is often used for collaborative monitoring. By combining and analyzing the images taken by the two cameras, more comprehensive and accurate information can be obtained.

[0003] Before the dual cameras are put into use, they need to be calibrated so that the center points of the images taken by the two cameras can overlap, to prevent the images taken by the two cameras from being deviated and affecting the analysis results of the images taken. The current common dual camera calibration method relies on manual adjustment. Since the manual adjustment process often involves human errors, such as the demarcation of the calibration area and the adjustment of the camera, the calibration accuracy of manual adjustment is not high. Summary of the invention

[0004] In view of the above problems, the present application provides a dual-camera calibration method, system and related devices to achieve the purpose of improving the accuracy of dual-camera calibration. The specific scheme is as follows:

[0005] A first aspect of the present application provides a dual-camera calibration method, the dual-camera calibration method comprising:

[0006] Acquire a first image captured by the first camera and a second image set captured by the second camera;

[0007] Selecting a target second captured image from the second captured image set, wherein the matching confidences of the target second captured image and the first captured image are all greater than the matching confidences of other second captured images and the first captured image;

[0008] Calculating an image offset value between the first captured image and the target second captured image according to a feature point offset set, wherein the feature point offset set includes an offset value of at least one feature point pair, each feature point pair including a feature point in the first captured image and a corresponding feature point in the target second captured image;

[0009] The image offset value is converted into a camera angle offset value, and the second camera is adjusted according to the camera angle offset value with the first camera as a reference.

[0010] In a possible implementation, the selecting a target second captured image from the second captured image set includes:

[0011] Performing feature detection on the first captured image to obtain feature points and feature point descriptors in the first captured image;

[0012] Performing feature detection on at least one second captured image in the second captured image set to obtain feature points and feature point descriptors of each second captured image in the second captured image set;

[0013] For the first captured image and one second captured image, calculating the matching confidence between the first captured image and the second captured image according to the similarity between the feature points and the feature point descriptors in the first captured image and the feature points and the feature point descriptors in the second captured image, so as to obtain the matching confidence between the first captured image and each second captured image in the set of the second captured images;

[0014] The second captured image with the maximum matching confidence is selected as the target second captured image.

[0015] In a possible implementation, calculating the image offset value between the first captured image and the target second captured image according to the feature point offset set includes:

[0016] Determine a first long-distance image region in the first captured image and obtain at least one feature point in the first long-distance image region, determine a second long-distance image region in a second captured image of the target and obtain at least one feature point in the second long-distance image region;

[0017] Acquire at least one feature point pair, and calculate the horizontal coordinate offset value and the vertical coordinate offset value in each feature point pair to obtain a feature point offset set, wherein each feature point pair includes a feature point in the first captured image and a corresponding feature point in the second captured image of the target;

[0018] An average value of the plurality of horizontal coordinate offset values ​​and an average value of the plurality of vertical coordinate offset values ​​in the feature point offset set are used as the image offset value.

[0019] In a possible implementation, determining a first long-distance image region in the first captured image and a second long-distance image region in the target second captured image includes:

[0020] Performing depth prediction on the first captured image and the target second captured image respectively to obtain a first depth map of the first captured image and a second depth map of the target second captured image;

[0021] Marking long-distance areas from the first depth map and the second depth map respectively according to a preset pixel threshold;

[0022] A first long-distance image region in the first captured image is determined according to the long-distance region marked in the first depth map, and a second long-distance image region in the target second captured image is determined according to the long-distance region marked in the second depth map.

[0023] In a possible implementation, determining a first long-distance image region in the first captured image according to the long-distance region marked in the first depth map, and determining a second long-distance image region in the target second captured image according to the long-distance region marked in the second depth map, includes:

[0024] generating a first binary mask map of the long-distance area marked by the first depth map, and performing a bitwise operation on the first binary mask map and the first captured image to obtain a first long-distance image area in the first captured image;

[0025] A second binary mask map of the long-distance area marked by the second depth map is generated, and a bitwise operation is performed on the second binary mask map and the target second captured image to obtain a second long-distance image area in the target second captured image.

[0026] In a possible implementation, the image offset value includes an average value of abscissa offset and an average value of ordinate offset, and converting the image offset value into a camera angle offset value includes:

[0027] The average value of the horizontal coordinate offset is converted into a camera angle offset value in the horizontal direction, and the average value of the vertical coordinate offset is converted into a camera angle offset value in the vertical direction.

[0028] A second aspect of the present application provides a dual-camera calibration system, the dual-camera calibration system comprising:

[0029] An acquisition unit, configured to acquire a first image captured by the first camera and a second image captured by the second camera;

[0030] a matching unit, configured to select a target second captured image from the set of second captured images, wherein a matching confidence level between the target second captured image and the first captured image is greater than a matching confidence level between each other second captured image and the first captured image;

[0031] a calculation unit, configured to calculate an image offset value between the first captured image and the target second captured image according to a feature point offset set, wherein the feature point offset set includes an offset value of at least one feature point pair, each feature point pair including a feature point in the first captured image and a corresponding feature point in the target second captured image;

[0032] The calibration unit is used to convert the image offset value into a camera angle offset value, and adjust the second camera according to the camera angle offset value with the first camera as a reference.

[0033] In a possible implementation, the matching unit includes:

[0034] a feature detection subunit, configured to perform feature detection on the first captured image to obtain feature points and feature point descriptors in the first captured image, and perform feature detection on at least one second captured image in the second captured image set to obtain feature points and feature point descriptors of each second captured image in the second captured image set;

[0035] a similarity calculation subunit, configured to calculate, for the first captured image and one second captured image, a matching confidence between the first captured image and the second captured image according to similarities between feature points and feature point descriptors in the first captured image and feature points and feature point descriptors in the second captured image, so as to obtain a matching confidence between the first captured image and each second captured image in a set of second captured images;

[0036] The confidence screening subunit is used to select the second captured image with the maximum matching confidence as the target second captured image.

[0037] A third aspect of the present application provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein:

[0038] The memory is used to store computer programs;

[0039] The processor is used to execute the computer program so that the electronic device can implement the dual-camera calibration method of the above-mentioned first aspect or any implementation manner of the first aspect.

[0040] In a fourth aspect, the present application provides a computer program product, comprising computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements the dual-camera calibration method of the first aspect or any implementation of the first aspect.

[0041] By means of the above technical scheme, the present application provides a dual camera calibration method, system and related device. The method uses the first camera as a reference for adjustment, obtains the first captured image taken by the first camera and the second captured image set of the second camera respectively, and takes the first captured image of the first camera as the main object, selects the target second captured image with the largest matching confidence from the second captured image set, and obtains the feature point offset set by comparing the feature points in the first captured image and the target second captured image. According to the feature point offset set, the image offset value between the first captured image and the target second captured image can be determined, and after the image offset value is converted into the camera angle offset value, the second camera is adjusted with the first camera as a reference. The method uses the feature points of the captured image as the calibration reference to avoid the error caused by artificial demarcation of the calibration area, and selects to determine the angle offset between the two cameras by comparing the image offset between the two captured images with the largest matching confidence of the two cameras, and automatically adjusts the other camera according to the angle offset to avoid the error caused by artificial adjustment. Therefore, the method can effectively improve the calibration accuracy of the dual camera. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale.

[0043] Figure 1 A schematic diagram of a dual-camera calibration method provided in an embodiment of the present application;

[0044] Figure 2 A schematic diagram of a camera field of view at a long distance provided in an embodiment of the present application;

[0045] Figure 3 A schematic diagram of a camera field of view at a close distance provided in an embodiment of the present application;

[0046] Figure 4 A schematic diagram of the structure of a dual-camera calibration system provided in an embodiment of the present application;

[0047] Figure 5 A hardware structure block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0048] The following describes the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. The terms used in the implementation method section of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.

[0049] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0050] The terms "first", "second" etc. in the specification of the application and the above-mentioned drawings are used to distinguish similar objects, and need not be used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable in appropriate circumstances, and this is only to describe the distinction mode adopted by the objects of the same attributes when describing in the embodiments of the application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0051] In order to improve the real-time and accuracy of drone monitoring, a dual camera system with two cameras, one above and one below, is often used for collaborative monitoring. Due to the differences in equipment between the dual cameras, the installation angle errors of the dual cameras, and external interference factors, there will be differences in the perspectives of the dual cameras, so that when one camera captures the target object, the other camera often fails to capture the target object due to the deviation in perspective. For example, the high-magnification camera and the low-magnification camera have different focal lengths, resulting in a different field of view angle between the high-magnification camera and the low-magnification camera, and there is also a large difference in the image coverage of the two. If the high-magnification camera and the low-magnification camera are not calibrated, some areas may be difficult to completely overlap, and the target detected by the low-magnification camera may be missed in the image captured by the high-magnification camera, especially when the target is in the edge area. The deviation between the two can easily cause the target to miss the target. Therefore, before the dual camera is put into use, the two cameras need to be calibrated.

[0052] At present, the calibration of dual cameras relies on manual work, and the process is as follows: manually select the calibration area on the chessboard, and use the dual cameras to capture the chessboard image. According to the offset value between the center point of the captured chessboard image and the center point of the calibration area, manually adjust the PTZ parameters of the camera so that the center point of the chessboard image is aligned with the center point of the calibration area, thereby completing the calibration of the dual cameras. Since manual errors are introduced into the camera calibration process when the calibration area is manually selected or the camera is manually adjusted, the calibration accuracy of the current dual camera calibration is not high.

[0053] The PTZ of the camera includes: Pan, Tilt and Zoom. Pan refers to the horizontal rotation angle of the camera left and right; Tilt refers to the vertical rotation angle of the camera up and down; Zoom refers to the magnification and reduction of the camera's viewing angle.

[0054] In order to solve the above problems, an embodiment of the present application provides a dual camera calibration method. The dual camera calibration method of the embodiment of the present application is described in detail below with reference to the accompanying drawings.

[0055] Reference Figure 1 , Figure 1 A schematic diagram of a dual-camera calibration method provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, a dual-camera calibration method provided in an embodiment of the present application may include steps S10 to S14, and these steps are described in detail below.

[0056] S10: Acquire a first image captured by a first camera and a set of second images captured by a second camera.

[0057] Among them, the combination of the first camera and the second camera in this embodiment can be a top-bottom combination. When calibrating, the lower camera can be adjusted with the upper camera as a calibration reference, or the upper camera can be adjusted with the lower camera as a calibration reference. Therefore, the first camera in this embodiment can be an upper camera or a lower camera. When the first camera is the upper camera, the second camera is the lower camera, and when the first camera is the lower camera, the second camera is the upper camera. Furthermore, there is a difference between the focal length of the first camera and the focal length of the second camera. Specifically, in this embodiment, the upper camera can be a high-magnification camera, and the lower camera can be a low-magnification camera.

[0058] Since the first camera serves as a calibration reference for the second camera, when acquiring captured images, the first camera can capture an image of an outdoor environment under the current PTZ parameters, that is, the first captured image, as a calibration reference image, and the second camera can capture images of the outdoor environment under multiple sets of PTZ parameters as a second captured image set. Specifically, the second camera can gradually switch the third parameter when two of the PTZ parameters are fixed to capture captured images under multiple sets of PTZ parameters. For example, when the P parameter and the Z parameter are fixed, starting from the T parameter of the initial position of the second camera, a second captured image is captured for each increase of 1° in the T parameter until the maximum T parameter of the second camera is reached to obtain a second captured image set. Of course, in this embodiment, when one of the PTZ parameters is fixed, the other two parameters can be switched step by step for shooting to obtain a second captured image set, or the three parameters can be switched step by step at the same time for shooting to obtain a second captured image set.

[0059] In order to facilitate the subsequent calculation of the offset between cameras, in this embodiment, when the first camera is shooting and obtaining the first captured image, the PTZ parameters when the first captured image is captured can be recorded in real time and marked on the first captured image. Similarly, when the second camera is shooting and obtaining the second captured image, the PTZ parameters when each second captured image is captured can also be recorded in real time, and the PTZ parameters of each second captured image can be marked on the second captured image.

[0060] Furthermore, after the present embodiment obtains the first captured image and the second captured image set, the captured images may be preprocessed to improve the image quality and the effect of subsequent image processing. Specifically, the present embodiment may introduce an adaptive image denoising ratio and image contrast enhancement strategy, and dynamically adjust the image denoising ratio and image contrast enhancement strategy according to different outdoor environmental conditions, so that captured images with good image quality can be obtained under various outdoor environmental conditions to adapt to various complex outdoor environments.

[0061] S11. Select a target second captured image from the second captured image set, wherein the matching confidence between the target second captured image and the first captured image is greater than the matching confidence between any other second captured image and the first captured image.

[0062] In order to reduce the adjustment amount of the second camera, the present embodiment screens the second captured image set to obtain the target second captured image with the highest matching confidence with the first captured image. Specifically, the specific process of obtaining the target second captured image in the present embodiment can be as shown in steps 1 to 4:

[0063] Step 1: Perform feature detection on the first captured image to obtain feature points and feature point descriptors in the first captured image;

[0064] Step 2: performing feature detection on at least one second captured image in the second captured image set to obtain feature points and feature point descriptors of each second captured image in the second captured image set;

[0065] Step 3: for the first captured image and one second captured image, according to the similarity between the feature points and feature point descriptors in the first captured image and the feature points and feature point descriptors in the second captured image, calculate the matching confidence of the first captured image and the second captured image to obtain the matching confidence of the first captured image and each second captured image in the set of second captured images;

[0066] Step 4: Select the second captured image with the maximum matching confidence as the target second captured image.

[0067] Among them, this embodiment uses the SuperPoint model to perform feature detection on all second captured images in the first captured image and the second captured image set. The SuperPoint model is a deep learning model for image feature point detection and descriptor generation, which can automatically detect feature points in an image and generate corresponding feature point descriptors. Feature points, also known as key points, are pixels or pixel areas with unique properties in an image, and feature points can still maintain stability and distinguishability under various image transformations (such as scale changes, light changes, etc.). The feature point in this embodiment is a pixel point, that is, a pixel block. A feature point descriptor refers to a numerical representation used to describe image information (such as texture, shape, gradient, etc.) around a feature point, which can usually be a vector or a set of numerical values.

[0068] After completing the feature detection, this embodiment uses the SuperGlue network model to match the first captured image with the second captured image set. Specifically, after obtaining the feature points and feature point descriptors in the first captured image and the feature points and feature point descriptors of each second captured image in the second captured image set output by the SuperPoint model, the output result is input into the SuperGlue network model. For the first captured image and a second captured image, the SuperGlue network model first encodes the feature points and feature point descriptors of the first captured image and the feature points and feature point descriptors in the second captured image into feature vectors, then uses the attention mechanism to increase the global information of the feature vectors, and finally uses the optimal transmission algorithm to calculate the inner product between the feature vectors of the first captured image and the feature vectors of the second captured image, obtains the matching score matrix and solves the matching score matrix, and finally uses the solution result as the matching confidence. Thus, the SuperGlue network model can output the matching feature points and matching confidence between the first captured image and each second captured image in the second captured image set, and selects the second captured image with the highest matching confidence as the target second captured image matching the first captured image.

[0069] S12. Calculate the image offset value between the first captured image and the target second captured image according to the feature point offset set, wherein the feature point offset set includes the offset value of at least one feature point pair, and each feature point pair includes a feature point in the first captured image and a corresponding feature point in the target second captured image.

[0070] The feature point pair is a combination of a feature point in the first captured image and a feature point in the second captured image of the target, and the two feature points have a corresponding relationship, and the positions in the respective captured images match each other to a certain extent. After obtaining at least one feature point pair, this embodiment can calculate the coordinate offset value between the two feature points in the feature point pair, and can obtain multiple coordinate offset values, thereby obtaining a feature point offset set. The image offset value is the offset value between the first captured image and the second captured image of the target.

[0071] Specifically, the specific process of calculating the image offset value between the first captured image and the target second captured image according to the feature point offset set may be as shown in steps 5 to 7:

[0072] Step 5: determining a first long-distance image region in the first captured image and obtaining at least one feature point in the first long-distance image region, determining a second long-distance image region in the target second captured image and obtaining at least one feature point in the second long-distance image region;

[0073] Step 6: Obtain at least one feature point pair, and calculate the horizontal coordinate offset value and the vertical coordinate offset value in each feature point pair to obtain a feature point offset set, where each feature point pair includes a feature point in the first captured image and a corresponding feature point in the second captured image of the target;

[0074] Step 7: Taking the average value of multiple horizontal coordinate offset values ​​and the average value of multiple vertical coordinate offset values ​​in the feature point offset set as the image offset value.

[0075] Among them, since the application scenario of this embodiment is to detect long-distance targets (such as drones), the application scenario is a long-distance scenario and the longer the distance, the smaller the field of view error. Figure 2 and Figure 3 As shown in , when the distance between camera 1 and camera 2 and their respective field of view angles are known, the calibration point distance of camera 1 can be obtained. Figure 2 As shown in , the calibration point distance of the camera is far away, and the real target point just falls within the field of view of camera 1. Figure 3 As shown, the camera calibration point is at a close distance, and the field of view convergence point of camera 1 is not far from the real target point, so the real target point is not within the field of view of camera 1. Of course, when the application scenario of this embodiment is a close-range scene, this embodiment can only select feature points in the close-range image area of ​​the first captured image and the target second captured image to form at least one feature point pair. This embodiment can also freely select a fixed area used in the captured image according to actual conditions.

[0076] Since the application scenario of this embodiment is a long-distance scenario, this embodiment mainly uses the feature points in the long-distance image area of ​​the first captured image and the long-distance image area of ​​the target second captured image to calculate the offset value, and the feature point pair is composed of the feature points in the long-distance image area of ​​the first captured image and the corresponding feature points in the long-distance image area of ​​the target second captured image. Among them, this embodiment can directly determine the feature points in the long-distance area of ​​the first captured image that match the feature points in the long-distance area of ​​the target second captured image based on the output results of the SuperGlue network model (the feature points that match between the first captured image and the target second captured image). Of course, this embodiment can also perform feature detection on the long-distance image area again after demarcating the long-distance image area of ​​the first captured image and the long-distance image area of ​​the target second captured image, so as to obtain the feature points in the long-distance image area of ​​the first captured image that match the feature points in the long-distance image area of ​​the target second captured image.

[0077] Specifically, in this embodiment, the depth map of the first captured image and the depth map of the second captured image of the target may be obtained by performing depth prediction on the first captured image and the second captured image of the target, thereby determining the long-distance image area of ​​the first captured image and the long-distance image area of ​​the second captured image of the target. The specific process may be shown in steps eight to ten:

[0078] Step 8: Perform depth prediction on the first captured image and the target second captured image respectively to obtain a first depth map of the first captured image and a second depth map of the target second captured image;

[0079] Step nine: marking the long-distance area from the first depth map and the second depth map respectively according to a preset pixel threshold;

[0080] Step ten: determining a first long-distance image region in the first captured image according to the long-distance region marked in the first depth map, and determining a second long-distance image region in the target second captured image according to the long-distance region marked in the second depth map.

[0081] Among them, the depth map is a special type of image, and each pixel value in the map represents the distance from a certain point in the scene to the shooting device. Specifically, this embodiment uses the Depth Anything model to perform depth prediction on the first captured image and the target second captured image respectively. The Depth Anything model is a monocular depth estimation model that can estimate depth information from a single image.

[0082] Since each pixel value of the depth map represents the distance of an object contained in the outdoor environment from the first camera or the second camera, the present embodiment can filter the pixel values ​​of the depth map by a preset pixel threshold, and filter out a plurality of pixel blocks representing the long-distance image area in the first captured image and a plurality of pixel blocks representing the long-distance image area in the second captured image of the target in the depth map as the long-distance area in the depth map.

[0083] After obtaining the long-distance area of ​​the depth map, the long-distance image area in the captured image can be determined. The specific process can be shown in steps 11 and 12:

[0084] Step 11: generating a first binary mask map of the long-distance area marked by the first depth map, and performing a bitwise operation on the first binary mask map and the first captured image to obtain a first long-distance image area in the first captured image;

[0085] Step 12: Generate a second binary mask map of the long-distance area marked by the second depth map, and perform a bitwise operation on the second binary mask map and the target second captured image to obtain a second long-distance image area in the target second captured image.

[0086] Among them, the binary mask map is a special type of image, and each pixel value on the mask map contains only two numerical values, and the two numerical values ​​represent two completely different states or categories. For example, the two numerical values ​​are 0 or 1, 0 represents the background or the "uninteresting" area, and 1 represents the foreground or the "interesting" area. Therefore, in the binary mask map of this embodiment, the pixel values ​​of multiple pixel blocks representing distant areas in the depth map are all the same numerical value (for example, the pixel values ​​are all 1), indicating that the distant area in the depth map is the "interesting" area, and the pixel values ​​of multiple pixel blocks representing other areas in the depth map are all another numerical value (for example, the pixel values ​​are all 0), indicating that other areas in the depth map are "uninteresting" areas.

[0087] Bitwise operation means: the pixel values ​​of each pixel point of the captured image are processed bit by bit. After obtaining the first binary mask map of the first depth map and the second binary mask map of the second depth map, the first binary mask map is bitwise operated with the first captured image, and the second binary mask map is bitwise operated with the target second captured image. Taking the first binary mask map and the first captured image as an example, in the first binary mask map, the pixel values ​​of the long-distance area are all 1, while the pixel values ​​of other areas are all 0. The pixel values ​​of the pixel blocks at the corresponding positions in the first binary mask map and the first captured image are multiplied respectively. In the first captured image, the pixel values ​​of the pixel blocks in the long-distance image area remain unchanged (all multiplied by 1), while the pixel values ​​of the pixel blocks in other areas are 0 (all multiplied by 0). Therefore, in the first captured image, only the long-distance image area is displayed normally with the original pixel value (such as color), while other areas are displayed abnormally (such as black), so that the first long-distance image area in the first captured image is obtained. Similarly, only the long-distance image area in the second captured image of the target is displayed normally with original pixel values, while other areas are displayed abnormally, so that the second long-distance image area in the second captured image of the target is obtained.

[0088] Of course, in this embodiment, the long-distance image area can also be directly defined in the captured image according to the pixel coordinate range of the long-distance area in the depth map and the pixel coordinate range.

[0089] After obtaining the first long-distance image region in the first captured image and the second long-distance image region in the second captured image of the target, multiple feature point pairs can be determined from the long-distance image region and the horizontal coordinate offset values ​​and the vertical coordinate offset values ​​in the feature pairs can be calculated to obtain a feature point offset set ( ).

[0090] When calculating the image offset value according to the feature point offset set, the present embodiment can directly calculate the average value of multiple horizontal coordinate offset values ​​and the average value of multiple vertical coordinate offset values ​​in the feature point offset set, and use the horizontal coordinate offset average value and the vertical coordinate offset average value as the image offset value. Of course, the present embodiment can also use each group of offset values ​​in the feature point offset set as a point, use the clustering algorithm K-Means to cluster and group all points in the feature point offset set, obtain the center point of all points in the feature point offset set through iterative calculation, and use the coordinate offset value represented by the center point as the image offset value.

[0091] S13: Convert the image offset value into a camera angle offset value, take the first camera as a reference, and adjust the second camera according to the camera angle offset value.

[0092] Among them, the camera angle offset value is the actual offset value of the camera in the three PTZ directions. In this embodiment, the offset value mainly involves the P parameter and the T parameter. Since the image offset value is the offset value in the pixel space, including the average horizontal coordinate offset and the average vertical coordinate offset, it is necessary to convert the image offset value to the corresponding PTZ parameter to facilitate the adjustment of the camera. Specifically, this embodiment can convert the average horizontal coordinate offset into the camera angle offset value in the horizontal direction, and convert the average vertical coordinate offset into the camera angle offset value in the vertical direction, so that the second camera can be calibrated with the first camera by adjusting the P parameter and the T parameter.

[0093] Specifically, the formula for calculating the horizontal camera angle offset value can be as follows:

[0094]

[0095] in, It can represent the camera angle offset value in the horizontal direction; It can represent the P parameter deviation when the first camera captures the first captured image and when the second camera captures the second captured image of the target. , It refers to the P parameter moved by the first camera compared to the initial position when the first camera captures the first captured image; It refers to the P parameter of the movement of the second camera compared to the initial position when capturing the second image of the target; It can represent the average value of the horizontal coordinate offset in the feature point offset set; the image width is the length of the first captured image or the second captured image of the target, and the lengths of the two are the same; the horizontal field of view angle is the viewing angle range of the second camera in the horizontal direction.

[0096] Further, for the and , since the first camera and the second camera both have an initial position PTZ parameter when starting to shoot, and this embodiment marks the PTZ parameter of each shot image in real time when acquiring the shot image, when the first camera moves from its initial position to a certain position, it shoots the first shot image, and according to the P parameter of its initial position and the P parameter marked by the first shot image, is the change value of the P parameter (in the horizontal direction, the horizontal movement angle from its initial position to a certain position). For example, the initial position of the first camera is a P parameter of 10 degrees, and when the first camera takes the first shot, the P parameter is 50 degrees, then is 40 degrees (50 degrees minus 10 degrees). Similarly, when the second camera moves from its initial position to a certain position, it captures the second captured image of the target. According to the P parameter of its initial position and the P parameter of the target second captured image tag, is the change value of the P parameter (in the horizontal direction, the horizontal movement angle from its initial position to a certain position).

[0097] The formula for calculating the vertical camera angle offset value can be as follows:

[0098]

[0099] in, It can represent the camera angle offset value in the vertical direction of water; It can represent the T parameter deviation when the first camera captures the first captured image and when the second camera captures the second captured image of the target. , It refers to the T parameter of the first camera moving compared to the initial position when capturing the first captured image; It refers to the T parameter of the movement of the second camera compared to the initial position when capturing the second image of the target; The image height of the average value of the vertical coordinate offset in the feature point offset set can be represented as the width of the first captured image or the second captured image of the target, and the widths of the two are the same; the vertical field of view angle is the viewing angle range of the second camera in the vertical direction.

[0100] against and Similarly, the first camera moves from its initial position to a certain position to obtain a first captured image. is the change value of the T parameter (in the vertical direction, the vertical movement angle from its initial position to a certain position); the second camera moves from its initial position to a certain position to obtain the second captured image of the target, is the change value of the T parameter (in the vertical direction, the vertical movement angle from its initial position to a certain position).

[0101] Furthermore, the calculation formula for the offset angles of the above two cameras is designed based on PTZ parameters. The initial position and PTZ parameters recorded when taking the image are used to calculate the preliminary angle offset. The image offset value obtained according to the taken image is converted into an actual camera angle offset value to compensate for the errors in camera installation and hardware parameters. The final calibration angle is calculated by the sum of the two offset values.

[0102] In this embodiment, the camera angle offset value in the horizontal direction is calculated ( ) and the vertical camera angle offset value ( ) and then, according to and The positive and negative values ​​of are used to adjust the second camera. , it means that the center of the second camera's field of view is biased to the right, so the second camera needs to be adjusted to the left. The adjustment angle is The absolute value of , it means that the center of the second camera's field of view is biased to the left, and the second camera needs to be adjusted to the right. The adjustment angle is The absolute value of . Similarly, when , it means that the center of the second camera's field of view is biased upward, and the second camera needs to be adjusted downward. The adjustment angle is The absolute value of , it means that the center of the second camera's field of view is biased downward, and the second camera needs to be adjusted upward. The adjustment angle is The absolute value of .

[0103] After the initial calibration, a loop optimization can be performed to further refine the offset value and improve the alignment accuracy of the first camera and the second camera. Finally, the center of the field of view of the first camera coincides with the center of the field of view of the second camera.

[0104] The embodiment of the present application provides a dual camera calibration method, which uses the first camera as a reference for adjustment, obtains a first captured image taken by the first camera and a second captured image set of the second camera respectively, takes the first captured image of the first camera as the main object, selects the target second captured image with the largest matching confidence from the second captured image set, and obtains a feature point offset set by comparing the feature points in the first captured image and the target second captured image. The image offset value between the first captured image and the target second captured image can be determined according to the feature point offset set, and after converting the image offset value into a camera angle offset value, the second camera is adjusted with the first camera as a reference. The method uses the feature points of the captured image as a calibration reference to avoid errors caused by artificial demarcation of the calibration area, and selects to determine the angle offset between the two cameras by comparing the image offset between the two captured images with the largest matching confidence of the two cameras, and automatically adjusts the other camera according to the angle offset to avoid errors caused by artificial adjustment. Therefore, the method can effectively improve the calibration accuracy of the dual camera.

[0105] A dual-camera calibration method provided in an embodiment of the present application is introduced above, and a system applying the dual-camera calibration method will be introduced below.

[0106] See also Figure 4 , Figure 4 This is a schematic diagram of the structure of a dual-camera calibration system provided in an embodiment of the present application. Figure 4 As shown, the dual camera calibration system includes:

[0107] An acquisition unit 100 is used to acquire a first captured image of a first camera and a second captured image set of a second camera;

[0108] A matching unit 110, configured to select a target second captured image from the set of second captured images, wherein a matching confidence level between the target second captured image and the first captured image is greater than a matching confidence level between each of the other second captured images and the first captured image;

[0109] A calculation unit 120, configured to calculate an image offset value between the first captured image and the target second captured image according to a feature point offset set, wherein the feature point offset set includes an offset value of at least one feature point pair, each feature point pair including a feature point in the first captured image and a corresponding feature point in the target second captured image;

[0110] The calibration unit 130 is used to convert the image offset value into a camera angle offset value, and adjust the second camera according to the camera angle offset value by taking the first camera as a reference.

[0111] In a possible implementation, the matching unit 110 may include:

[0112] a feature detection subunit, configured to perform feature detection on the first captured image to obtain feature points and feature point descriptors in the first captured image, and perform feature detection on at least one second captured image in the second captured image set to obtain feature points and feature point descriptors of each second captured image in the second captured image set;

[0113] a similarity calculation subunit, for calculating, for a first captured image and a second captured image, a matching confidence of the first captured image and the second captured image according to similarities between feature points and feature point descriptors in the first captured image and feature points and feature point descriptors in the second captured image, so as to obtain a matching confidence of the first captured image and each second captured image in a set of second captured images;

[0114] The confidence screening subunit is used to select the second captured image with the maximum matching confidence as the target second captured image.

[0115] In a possible implementation, the calculation unit 120 may include:

[0116] A region delineation subunit, configured to determine a first long-distance image region in a first captured image and obtain at least one feature point in the first long-distance image region, determine a second long-distance image region in a second captured image of the target and obtain at least one feature point in the second long-distance image region;

[0117] an offset calculation subunit, configured to obtain at least one feature point pair, and calculate a horizontal coordinate offset value and a vertical coordinate offset value in each feature point pair to obtain a feature point offset set, wherein each feature point pair includes a feature point in the first captured image and a corresponding feature point in the second captured image of the target;

[0118] The average value calculation subunit is used to take the average value of multiple horizontal coordinate offset values ​​and the average value of multiple vertical coordinate offset values ​​in the feature point offset set as the image offset value.

[0119] In a possible implementation, determining the first long-distance image region in the first captured image and the second long-distance image region in the target second captured image in the region demarcation subunit may include:

[0120] A depth prediction subunit, configured to perform depth prediction on the first captured image and the target second captured image respectively, and obtain a first depth map of the first captured image and a second depth map of the target second captured image;

[0121] a region marking subunit, for marking a long-distance region from the first depth map and the second depth map respectively according to a preset pixel threshold;

[0122] The area determination subunit is used to determine a first long-distance image area in the first captured image according to the long-distance area marked in the first depth map, and to determine a second long-distance image area in the target second captured image according to the long-distance area marked in the second depth map.

[0123] In a possible implementation, the above region determination subunit may be specifically configured as follows:

[0124] A first binary mask map of a long-distance area marked by a first depth map is generated, and a bitwise operation is performed on the first binary mask map and the first captured image to obtain a first long-distance image area in the first captured image; a second binary mask map of a long-distance area marked by a second depth map is generated, and a bitwise operation is performed on the second binary mask map and the target second captured image to obtain a second long-distance image area in the target second captured image.

[0125] In a possible implementation, the image offset value may include an average value of abscissa offset and an average value of ordinate offset. The calibration unit 130 converts the image offset value into a camera angle offset value, which may be specifically configured as follows:

[0126] The average value of the horizontal axis offset is converted into a camera angle offset value in the horizontal direction, and the average value of the vertical axis offset is converted into a camera angle offset value in the vertical direction.

[0127] The present application also provides an electronic device in an embodiment. Figure 5 As shown, it shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiment of the present application. The electronic device in the embodiment of the present application may include but is not limited to fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 5 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0128] like Figure 5 As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 to a random access memory (RAM) 503. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 503. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0129] Typically, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a memory card, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 5 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0130] An embodiment of the present application also provides a computer program product, including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any dual-camera calibration method provided in the embodiment of the present application.

[0131] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any dual-camera calibration method provided in the embodiment of the present application.

[0132] It should also be noted that the system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the system embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.

[0133] Through the description of the above implementation mode, the technicians in the field can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. In general, all functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better implementation mode in more cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, a U disk, a mobile hard disk, a ROM, a RAM, a disk or an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0134] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0135] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a training device, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, training device, or data center. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.

[0136] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0137] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0138] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the protection scope of the present application.

Claims

1. A dual-camera calibration method, characterized in that: The dual-camera calibration method comprises: Acquire a first image captured by the first camera and a second image set captured by the second camera; Selecting a target second captured image from the second captured image set, wherein the matching confidences of the target second captured image and the first captured image are greater than the matching confidences of other second captured images and the first captured image; Calculating an image offset value between the first captured image and the target second captured image according to a feature point offset set, wherein the feature point offset set includes an offset value of at least one feature point pair, each feature point pair including a feature point in the first captured image and a corresponding feature point in the target second captured image; The image offset value is converted into a camera angle offset value, and the second camera is adjusted according to the camera angle offset value with the first camera as a reference.

2. The dual camera calibration method according to claim 1, characterized in that: The selecting a target second captured image from the second captured image set comprises: Performing feature detection on the first captured image to obtain feature points and feature point descriptors in the first captured image; Performing feature detection on at least one second captured image in the second captured image set to obtain feature points and feature point descriptors of each second captured image in the second captured image set; For the first captured image and one second captured image, calculating the matching confidence between the first captured image and the second captured image according to the similarity between the feature points and the feature point descriptors in the first captured image and the feature points and the feature point descriptors in the second captured image, so as to obtain the matching confidence between the first captured image and each second captured image in the set of the second captured images; The second captured image with the maximum matching confidence is selected as the target second captured image.

3. The dual camera calibration method according to claim 1 or 2, characterized in that: The calculating the image offset value between the first captured image and the target second captured image according to the feature point offset set includes: Determine a first long-distance image region in the first captured image and obtain at least one feature point in the first long-distance image region, determine a second long-distance image region in a second captured image of the target and obtain at least one feature point in the second long-distance image region; Acquire at least one feature point pair, and calculate the horizontal coordinate offset value and the vertical coordinate offset value in each feature point pair to obtain a feature point offset set, wherein each feature point pair includes a feature point in the first captured image and a corresponding feature point in the second captured image of the target; An average value of the plurality of horizontal coordinate offset values ​​and an average value of the plurality of vertical coordinate offset values ​​in the feature point offset set are used as the image offset value.

4. The dual camera calibration method according to claim 3, characterized in that: Determining a first long-distance image region in the first captured image and a second long-distance image region in the target second captured image includes: Performing depth prediction on the first captured image and the target second captured image respectively to obtain a first depth map of the first captured image and a second depth map of the target second captured image; Marking long-distance areas from the first depth map and the second depth map respectively according to a preset pixel threshold; A first long-distance image region in the first captured image is determined according to the long-distance region marked in the first depth map, and a second long-distance image region in the target second captured image is determined according to the long-distance region marked in the second depth map.

5. The dual camera calibration method according to claim 4, characterized in that: The method of determining a first long-distance image region in the first captured image according to the long-distance region marked in the first depth map, and determining a second long-distance image region in the second captured image of the target according to the long-distance region marked in the second depth map, comprises: generating a first binary mask map of the long-distance area marked by the first depth map, and performing a bitwise operation on the first binary mask map and the first captured image to obtain a first long-distance image area in the first captured image; A second binary mask map of the long-distance area marked by the second depth map is generated, and a bitwise operation is performed on the second binary mask map and the target second captured image to obtain a second long-distance image area in the target second captured image.

6. The dual-camera calibration method according to claim 1, characterized in that: The image offset value includes an abscissa offset average value and an ordinate offset average value, and the step of converting the image offset value into a camera angle offset value includes: The average value of the horizontal coordinate offset is converted into a camera angle offset value in the horizontal direction, and the average value of the vertical coordinate offset is converted into a camera angle offset value in the vertical direction.

7. A dual camera calibration system, characterized in that: The dual-camera calibration system comprises: An acquisition unit, configured to acquire a first image captured by the first camera and a second image captured by the second camera; a matching unit, configured to select a target second captured image from the set of second captured images, wherein a matching confidence level between the target second captured image and the first captured image is greater than a matching confidence level between each other second captured image and the first captured image; a calculation unit, configured to calculate an image offset value between the first captured image and the target second captured image according to a feature point offset set, wherein the feature point offset set includes an offset value of at least one feature point pair, each feature point pair including a feature point in the first captured image and a corresponding feature point in the target second captured image; The calibration unit is used to convert the image offset value into a camera angle offset value, and adjust the second camera according to the camera angle offset value with the first camera as a reference.

8. The dual camera calibration system according to claim 7, characterized in that: The matching unit comprises: a feature detection subunit, configured to perform feature detection on the first captured image to obtain feature points and feature point descriptors in the first captured image, and perform feature detection on at least one second captured image in the second captured image set to obtain feature points and feature point descriptors of each second captured image in the second captured image set; a similarity calculation subunit, configured to calculate, for the first captured image and one second captured image, a matching confidence between the first captured image and the second captured image according to similarities between feature points and feature point descriptors in the first captured image and feature points and feature point descriptors in the second captured image, so as to obtain a matching confidence between the first captured image and each second captured image in a set of second captured images; The confidence screening subunit is used to select the second captured image with the maximum matching confidence as the target second captured image.

9. An electronic device, characterized in that: The method comprises at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the electronic device can implement the dual camera calibration method as described in any one of claims 1 to 6.

10. A computer program product, characterized in that The method comprises computer-readable instructions, and when the computer-readable instructions are executed on an electronic device, the electronic device implements the dual-camera calibration method as described in any one of claims 1 to 6.