Calibration, 3D Vision, and Depth Point Cloud Computing Methods for Autofocus Binocular Cameras

By using a production line calibration method for autofocus binocular cameras, the optimal focus position relationship curve is obtained and the relative position is calibrated, which solves the blurring problem of binocular cameras when the scene depth changes, and realizes fast and clear focusing and high-precision dense depth point cloud computing.

CN114359406BActive Publication Date: 2025-11-14XIANGCHANG (SHENZHEN) TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111654221.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2025-11-14
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

Existing binocular cameras are prone to blurring when shooting scenes with varying depths, and their accuracy in calculating dense depth point clouds is poor. Especially in applications such as smartphones, drones, automobiles, and virtual reality, existing focusing algorithms cannot adjust in time, resulting in unclear videos or images.

Method used

The production line calibration method of autofocus binocular cameras is adopted. By configuring a calibration board, the optimal focus position relationship curve of the left and right cameras is obtained. Combined with the parameters of the monocular camera and the relative position calibration, fast focusing and dense depth point cloud computing are achieved.

Benefits of technology

It achieves fast and clear focusing of the binocular camera when the scene depth changes, generating high-quality 3D images and dense depth point clouds, improving the clarity and accuracy of captured videos and images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114359406B_ABST
    Figure CN114359406B_ABST
Patent Text Reader

Abstract

This invention discloses a production line calibration method for an autofocus binocular camera, comprising: A1: configuring a first calibration board at a scene depth position of the binocular camera; A2: moving the lenses of the left and right cameras to a new position to acquire an image of the first calibration board; A3: repeating step A2; A4: determining the optimal focus position of each of the left and right cameras for the current scene depth; A5: repeating steps A1 to A4 to obtain the optimal focus positions of each of the left and right cameras for each scene depth within the entire depth range, obtaining curves showing the relationship between the scene depth and the optimal focus position of each of the left and right cameras, and obtaining the associated focus positions corresponding to multiple scene depths within the entire depth range; A6: calibrating the left and right cameras for monocular and binocular cameras respectively based on the associated focus positions obtained in step A5. This invention also discloses a 3D stereo vision shooting method for an autofocus binocular camera and a method for calculating dense depth point clouds.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a production line calibration method for an autofocus binocular camera, a 3D stereo vision shooting method, and a method for calculating dense depth point clouds. Background Technology

[0002] Autofocus refers to a function built into a camera that automatically focuses on the subject and achieves a sharp image through electronic and mechanical devices. Due to its accurate focusing and ease of operation, cameras with autofocus are increasingly used in various industries, including smartphones, drones, and video surveillance.

[0003] Camera applications are transitioning from 2D to a 3D era based on binocular and multi-view cameras. In consumer electronics, 3D stereoscopic video allows virtual objects to be better integrated into real-world scenes in Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR) applications. In the 5G era, the cultural and entertainment industry is seeing VR live streams of performances and sporting events providing viewers with an immersive experience. In high-risk industries such as tunnel excavation, mining operations, explosive ordnance disposal, and pipeline inspection, remote 3D stereoscopic vision systems serve as a means of "full informatization," protecting the safety of construction workers. Remote medical care and high-definition binocular endoscopes and exoscopes allow surgeons to experience realistic 3D effects, improving diagnostic and surgical outcomes. The left and right images captured by binocular cameras are the foundation for these 3D stereoscopic vision applications. Currently, commonly used binocular cameras typically use fixed-focus lenses, and the focusing distance of the shooting device is fixed. When the shooting scene changes, if the depth of the new scene exceeds the camera's depth of field, the captured video will be blurry. In consumer electronics such as smartphones, when shooting videos with a monocular autofocus camera, the currently common contrast-based focusing algorithms take a long time to calculate in certain scenarios, resulting in blurry video content. Furthermore, some current mobile phone focusing algorithms cannot effectively handle minor changes in the shooting scene while maintaining the same depth. In such cases, the phone continuously moves the lens to search for the optimal focus point, causing the video to be sometimes clear and sometimes blurry.

[0004] Furthermore, in industries such as mobile phones, infrared structured light camera modules are used to capture 3D point clouds of faces for facial recognition. In the drone industry, binocular cameras enable 3D environmental perception to support obstacle avoidance during flight. In the automotive industry, binocular cameras are used for ranging and 3D environmental perception to assist driving and support autonomous driving. In the robotic vacuum cleaner industry, binocular cameras are used for 3D environmental perception to achieve path planning, obstacle avoidance, and object recognition. In the virtual reality (VR) industry, binocular cameras are used for VR headset positioning and tracking, controller positioning, and 3D content acquisition. The left and right images captured by binocular cameras are the foundation for these 3D measurement applications. Currently, commonly used binocular cameras typically use fixed-focus lenses, and the focusing distance of the shooting device is fixed. When the shooting scene changes, if the depth of the new scene exceeds the camera's depth of field, the captured image may be blurry, resulting in poor stereo matching between the left and right camera images and low accuracy in the calculated dense depth point cloud.

[0005] The above background information is provided only to aid in understanding the concept and technical solution of this invention. It does not necessarily belong to the prior art of this patent application. In the absence of clear evidence that the above information was disclosed on the filing date of this patent application, the above background information should not be used to evaluate the novelty and inventiveness of this application. Summary of the Invention

[0006] To address the aforementioned technical problems, this invention proposes a production line calibration method for an autofocus binocular camera, a 3D stereo vision shooting method, and a method for calculating dense depth point clouds. These methods enable the captured video to remain clear at all times and also allow for the calculation of dense depth point clouds.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] This invention discloses a production line calibration method for an autofocus binocular camera, comprising:

[0009] A1: Configure the first calibration board at a scene depth location of the binocular camera;

[0010] A2: After moving the lenses of the left and right cameras of the binocular camera to a new position according to a preset step size, the left and right cameras respectively take pictures of the first calibration board to obtain the image of the first calibration board;

[0011] A3: Repeat step A2 to obtain first calibration board images captured from multiple positions of the left and right camera lenses, so that each of the left and right cameras acquires a set of first calibration board images.

[0012] A4: Calculate the contrast of the preset area in a set of first calibration board images acquired by the left and right cameras respectively, and determine the position of the lens corresponding to the maximum contrast in each set as the best focus position of the left and right cameras for the current scene depth.

[0013] A5: Repeat steps A1 to A4 to obtain the optimal focus positions of the left and right cameras for each scene depth within the entire depth range. Based on the optimal focus positions of the left and right cameras for each scene depth, obtain the curves showing the relationship between the scene depth and the optimal focus position of the left and right cameras. Then, based on the curves showing the relationship between the scene depth and the optimal focus position of the left and right cameras, obtain the associated focus positions corresponding to the multiple scene depths within the entire depth range. Each associated focus position is composed of the scene depth and the corresponding optimal focus position of the left camera and the optimal focus position of the right camera.

[0014] A6: Based on the associated focus positions obtained in step A5, at the scene depth of each associated focus position, perform single-target calibration on the left and right cameras to obtain the monocular camera parameters of the left and right cameras, and combine the monocular camera parameters of the left and right cameras to calibrate the relative position of the binocular cameras.

[0015] In some embodiments, the present invention further includes the following technical features:

[0016] The first calibration board is a solid dot array calibration board, and its size, dot diameter and spacing are positively correlated with the scene depth.

[0017] The preset area in step A4 includes one area or five areas. One area refers to the square area at the center of the image, and the five areas refer to the square area at the center of the image and four square areas at a field of view of 0.5 or 0.8 in the image.

[0018] The formula used in step A4 Calculate the contrast of a preset region in the first calibration plate image, where L max and L min These are the largest and smallest pixel grayscale values ​​in the preset area, respectively.

[0019] In step A1, the first calibration plate is perpendicular to the optical axis of the binocular camera, and the entire depth range is divided into multiple depth intervals, each using a calibration template of a different size. When the binocular camera enters the next depth interval for calibration, the calibration template adapted to that depth interval is replaced. In step A2, which is executed before each step A3, the lenses of the left and right cameras are moved to the closest focusing distance. When step A2 is repeated in step A3, the lenses of the left and right cameras are moved sequentially by a preset step size until the lenses of the left and right cameras are at the furthest focusing distance, so as to traverse all positions of the lenses of the left and right cameras.

[0020] Step A6 specifically includes:

[0021] A61: Based on the scene depth in each of the associated focus positions obtained in step A5, the entire depth range is divided into multiple depth intervals, and a second calibration plate is configured in each depth interval.

[0022] A62: The left and right cameras take pictures of the second calibration board from different angles to obtain multiple images of the second calibration board. Based on minimizing the reprojection error, the intrinsic parameters, extrinsic parameters and distortion parameters of the left and right cameras are obtained. The relative position of the binocular cameras is calibrated by combining the extrinsic parameters of the left and right cameras.

[0023] A63: Repeat steps A61 to A62 to obtain the intrinsic, extrinsic, and distortion parameters of the left and right cameras, as well as the relative positions of the binocular cameras, at the scene depth in each of the associated focus positions.

[0024] The second calibration board uses a black and white checkerboard calibration board, the size of which, the size of the checkerboard and the spacing of the grid are positively correlated with the scene depth. The entire depth range is divided into multiple depth intervals, and each interval uses a calibration template of a different size. When the binocular camera enters the next depth interval for calibration, the calibration template adapted to that depth interval is replaced.

[0025] This invention discloses a production line calibration device for an autofocus binocular camera, which includes a computer program for implementing the above-mentioned production line calibration method.

[0026] This invention discloses a 3D stereoscopic vision shooting method using an autofocus binocular camera, comprising: performing production line calibration on the binocular camera using the aforementioned production line calibration method, and:

[0027] B1: The lenses of the left and right cameras are respectively located at the initial position of focusing on the same scene depth, the scene depth is saved as the contrast depth value, and the left and right cameras respectively capture the first frame image of the current scene;

[0028] B2: Calculate the depth of the current scene based on the monocular camera parameters and the relative position of the binocular camera at the associated focus position corresponding to the contrast depth value;

[0029] B3: Compare the calculated depth of the current scene with the comparison depth value. If they are consistent, proceed to step B6; otherwise, proceed to step B4.

[0030] B4: Based on the calculated depth of the current scene, according to the multiple associated focus positions obtained in step A5, push the lenses of the left and right cameras to the optimal focus positions corresponding to the depth of the current scene, and save the depth of the current scene as a comparison depth value.

[0031] B5: The left and right cameras respectively capture a new frame of the current scene;

[0032] B6: Based on the monocular camera parameters and the relative position of the binocular camera at the associated focus position corresponding to the contrast depth value, stereo correction is performed on the images captured by the left and right cameras respectively to obtain the left image and the right image. The left image refers to the stereo-corrected image captured by the left camera, and the right image refers to the stereo-corrected image captured by the right camera.

[0033] B7: Combine the left and right images and output them to the screen of the 3D display device;

[0034] B8: The left and right cameras capture a new frame of the current scene and return to step B2.

[0035] In some embodiments, the present invention further includes the following technical features:

[0036] In step B2, the following method is used: based on the accurately matched feature point pairs and the relative positions of the monocular camera parameters and binocular camera of the left and right cameras corresponding to the contrast depth value, the three-dimensional spatial position of the object point corresponding to the feature point is calculated using triangulation to calculate the depth of the current scene.

[0037] Step B6 specifically includes: eliminating the distortion of the images captured by the left and right cameras respectively according to the monocular camera parameters of the left and right cameras corresponding to the contrast depth value; then, according to the monocular camera parameters of the left and right cameras corresponding to the contrast depth value and the relative position of the binocular cameras, projecting the distortion-eliminated images captured by the left and right cameras onto the same plane and performing row alignment, so that any point on one image is in the same row as the corresponding point on the other image, so as to obtain the left and right images through stereo correction.

[0038] Step B7 involves stitching the left and right images together, adjusting the resolution of the stitched image to match the resolution of the 3D display device, and then outputting the stitched and resolution-adjusted image to the screen of the 3D display device.

[0039] The present invention discloses a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are invoked and executed by a processor, the computer-executable instructions cause the processor to implement the steps of the above-described 3D stereoscopic vision imaging method.

[0040] This invention discloses a method for calculating depth point clouds based on an autofocus binocular camera, comprising: performing production line calibration on the binocular camera using the aforementioned production line calibration method, and:

[0041] C1: The lenses of the left and right cameras are respectively located at the initial position of focusing on the same scene depth, the scene depth is saved as the contrast depth value, and the left and right cameras respectively capture the first frame image of the current scene;

[0042] C2: Calculate the depth of the current scene based on the monocular camera parameters and the relative position of the binocular camera at the associated focus position corresponding to the contrast depth value;

[0043] C3: Compare the calculated depth of the current scene with the comparison depth value. If they are the same, proceed to step C6; otherwise, proceed to step C4.

[0044] C4: Based on the calculated depth of the current scene, according to the multiple associated focus positions obtained in step A5, push the lenses of the left and right cameras to the optimal focus positions corresponding to the depth of the current scene, and save the depth of the current scene as a comparison depth value.

[0045] C5: The left and right cameras each capture a new frame of the current scene;

[0046] C6: Based on the calibrated monocular camera parameters and the relative position of the binocular camera at the associated focus position corresponding to the contrast depth value, stereo correction is performed on the images captured by the left and right cameras respectively to obtain the left image and the right image. The left image refers to the stereo-corrected image captured by the left camera, and the right image refers to the stereo-corrected image captured by the right camera.

[0047] C7: Perform stereo matching between the left and right images to generate a disparity map;

[0048] C8: Calculate the dense depth point cloud of a binocular camera based on the disparity map;

[0049] C9: The left and right cameras capture a new frame of the current scene and return to step C2.

[0050] In some embodiments, the present invention further includes the following technical features:

[0051] In step C2, the following method is used: based on the accurately matched feature point pairs and the relative positions of the monocular camera parameters and binocular camera of the left and right cameras corresponding to the contrast depth value, the three-dimensional spatial position of the object point corresponding to the feature point is calculated using triangulation to calculate the depth of the current scene.

[0052] Step C6 specifically includes: eliminating the distortion of the images captured by the left and right cameras respectively according to the monocular camera parameters of the left and right cameras corresponding to the contrast depth value; then, according to the monocular camera parameters of the left and right cameras corresponding to the contrast depth value and the relative position of the binocular cameras, projecting the distortion-eliminated images captured by the left and right cameras onto the same plane and performing row alignment, so that any point on one image is in the same row as the corresponding point on the other image, so as to obtain the left and right images through stereo correction.

[0053] Step C7 specifically includes:

[0054] Stereo matching is performed on any pixel in the left image along the same row in the right image. After matching, the disparity Δδ of the pixel in the left and right images is calculated based on the horizontal coordinates of the two corresponding pixels in the image: Δδ = x1 - x2. The disparity is calculated for each pixel to obtain a disparity map; where x1 and x2 are the horizontal coordinates of the pixel in the left and right images, respectively.

[0055] Step C8 specifically includes:

[0056] C81: Based on the disparity map, according to the formula Calculate the depth Z of the object point corresponding to the pixel in the image captured by the right camera after stereo correction, so as to convert the disparity map into a depth map; where B is the distance between the optical centers of the left and right cameras, f is the focal length of the left or right camera, and Δδ is the disparity of the pixel in the images captured by the left and right cameras respectively after stereo correction.

[0057] C82: Obtain the third-dimensional depth coordinate Z of each pixel in the right camera coordinate system based on the depth map, and calculate the X and Y coordinates of each pixel in the right camera coordinate system:

[0058]

[0059] In the formula, (c x c y) is the principal point of the right camera, f x f y y represents the horizontal and vertical focal lengths of the right camera, respectively, and (x, y) represents the pixel coordinates of the corresponding pixel in the right camera.

[0060] C83: Based on step C82, obtain the three-dimensional coordinates of each pixel in the right image in the right camera coordinate system. Combine the RGB color information of each pixel in the right image to obtain the point cloud of each pixel as P = {X,Y,Z,R,G,B}. Calculate the three-dimensional coordinates and corresponding RGB color information of all pixels in the depth map to obtain the dense depth point cloud of the binocular camera.

[0061] The present invention discloses a computer-readable storage medium storing computer-executable instructions, which, when invoked and executed by a processor, cause the processor to implement the steps of the above-described method for calculating depth point clouds.

[0062] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0063] The present invention discloses a production line calibration method for an autofocus binocular camera, which calibrates the scene depth and optimal focus position relationship curves of the left and right cameras respectively. Based on the curves, a series of associated focus positions can be determined within the entire depth range, allowing the lens to be directly pushed to the optimal focus position corresponding to the depth, completing focus in one step, eliminating the frequent back-and-forth search process, improving focus speed, and thus ensuring that the binocular camera maintains clear focus at all times. Furthermore, at each associated focus position, the left and right cameras are calibrated individually, and the relative positions of the binocular cameras are calibrated, laying the foundation for subsequent generation of clear, accurately aligned, and high-quality 3D images and calculation of dense depth point clouds.

[0064] The 3D stereoscopic vision shooting method of the autofocus binocular camera disclosed in this invention involves, after production line calibration of the autofocus binocular camera, capturing images of the current scene from the left and right cameras respectively at the initial associated focus position, and calculating the depth of the current scene. Based on the current scene depth, and according to the calibrated relationship curve between the scene depth and the optimal focus position of each of the left and right cameras, the lenses of the left and right cameras are pushed to the optimal focus position of the current scene, and the left and right cameras capture clear images of the current scene respectively. Then, distortion removal and stereoscopic correction are performed on the images from the left and right cameras, and the stereoscopically corrected left and right images are stitched together and output in full screen on a 3D display device to generate a clear, accurately aligned, and high-quality 3D image.

[0065] The method for calculating dense depth point clouds using an autofocus binocular camera disclosed in this invention involves, after production line calibration of the autofocus binocular camera, capturing images of the current scene from the left and right cameras respectively at the initial associated focus position, and calculating the depth of the current scene. Based on the current scene depth, and according to the calibrated relationship curve between the scene depth and the optimal focus position of each of the left and right cameras, the lenses of the left and right cameras are pushed to the optimal focus position of the current scene, and the left and right cameras capture clear images of the current scene respectively. Then, distortion removal and stereo correction are performed on the images of the left and right cameras. Stereo matching is performed on the stereo-corrected left and right images to generate a disparity map, thereby enabling the calculation of a high-precision dense depth point cloud. Attached Figure Description

[0066] Figure 1 This is a flowchart of the production line calibration method for an autofocus binocular camera according to Embodiment 1 of the present invention;

[0067] Figure 2 This is a schematic diagram of the first calibration plate in Embodiment 1 of the present invention;

[0068] Figure 3 yes Figure 1 The detailed flowchart of step A6;

[0069] Figure 4 This is a schematic diagram of the second calibration plate in Embodiment 1 of the present invention;

[0070] Figure 5 It is a curve obtained by calculating contrast during the autofocus process;

[0071] Figure 6 It is a graph showing the relationship between the scene depth and the optimal focus position of each of the left and right cameras, as well as a schematic diagram of the associated focus position.

[0072] Figure 7 This is a flowchart of the 3D stereoscopic vision shooting method of the autofocus binocular camera according to Embodiment 2 of the present invention;

[0073] Figure 8 This is a diagram illustrating depth calculation;

[0074] Figure 9a This is a schematic diagram of the original left and right images from a binocular camera;

[0075] Figure 9b This is a schematic diagram of the left and right images after stereo correction by a binocular camera;

[0076] Figure 10 This is a flowchart of the method for calculating dense depth point clouds using an autofocus binocular camera according to Embodiment 3 of the present invention. Detailed Implementation

[0077] The embodiments of the present invention will be described in detail below. It should be emphasized that the following description is merely exemplary and not intended to limit the scope and application of the present invention.

[0078] It should be noted that when a component is referred to as "fixed to" or "set on" another component, it can be directly on or indirectly on that other component. When a component is referred to as "connected to" another component, it can be directly connected to or indirectly connected to that other component. Furthermore, a connection can be used for both fixing and circuit / signal connectivity.

[0079] It should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", and "outer" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing the embodiments of the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the present invention.

[0080] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of the present invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0081] Example 1:

[0082] like Figure 1 As shown, Embodiment 1 of the present invention discloses a production line calibration method for an autofocus binocular camera, comprising the following steps:

[0083] A1: Configure the first calibration board at a scene depth location of the binocular camera;

[0084] In this embodiment, the first calibration plate is a solid dot array calibration plate. Specifically, the pattern of the solid dot array calibration plate is as follows: Figure 2 As shown, the solid circle array consists of C rows × L columns of solid circles, where C and L are natural numbers (e.g., C = 8, L = 11). Each solid circle is the same size, has the same radius, and the center distance is the same in both the horizontal and vertical directions. During calibration, the first calibration board must be perpendicular to the optical axis of the binocular camera.

[0085] The field of view of a camera is directly proportional to the scene depth. For example, when a camera captures a scene at depth g, its field of view (horizontal and vertical directions) is twice that of a camera capturing a scene at depth g / 2. Therefore, in this embodiment, the size of the solid circular array template, the diameter of the dots, and the spacing of the dots used for calibration should increase proportionally with the scene depth. Furthermore, considering production line efficiency and production costs, the entire depth range is divided into several depth intervals, and each depth interval uses a calibration template of a specific size. When the binocular camera enters the next depth interval for calibration, a calibration template adapted to that depth interval needs to be replaced, with its size, dot diameter, and spacing proportional to the scene depth.

[0086] A2: After moving the lenses of the left and right cameras of the binocular camera to a new position according to a preset step size, the left and right cameras respectively take pictures of the first calibration board to obtain the image of the first calibration board.

[0087] A3: Repeat step A2 to obtain the first calibration images captured at multiple positions of the left and right camera lenses, so that each of the left and right cameras acquires a set of first calibration board images.

[0088] Step A3 further involves repeating step A2 to traverse all positions of the left and right camera lenses, with each camera acquiring a set of first calibration board images taken at different positions. In a preferred embodiment, in step A2 performed before each execution of step A3, the lenses of the left and right cameras are moved to their closest focusing distance. When repeating step A2 in step A3, the lenses of the left and right cameras are moved sequentially by a preset step size until the lenses of the left and right cameras are at their furthest focusing distance, thus traversing all positions of the left and right camera lenses.

[0089] A4: Calculate the contrast of the preset area in the first calibration board images captured by the left and right cameras at different positions respectively, and determine the position of the lens corresponding to the maximum contrast in each group as the best focus position of the left and right cameras for the current scene depth (the current scene depth refers to the distance from the first calibration board to the left and right cameras).

[0090] In the embodiments, the preset area includes one area or five areas, where one area refers to the square area at the center of the image, and the five areas refer to the square area at the center of the image and four square areas at a field of view of 0.5 or 0.8 in the image.

[0091] Specifically, the formula is used to calculate the contrast of a preset region in a set of first calibration board images captured by the left and right cameras at different positions. Perform the calculation, where L max and L minThese represent the highest and lowest pixel grayscale values ​​in the preset area, respectively. The higher the contrast value, the clearer the image in that area.

[0092] When an autofocus camera focuses, it iterates through lens positions according to a certain pattern to find the lens position with the clearest image. In this embodiment, the strategy is that the camera's optical axis is perpendicular to the calibration template. The lens moves from the closest focusing distance position in steps until it reaches the lens position with the farthest focusing distance. At each position, an image of the calibration template is captured, and the contrast of a specified area of ​​that image is calculated, such as the center area of ​​the image. After capturing images and calculating contrast at all positions, the camera lens position corresponding to the contrast peak is found. The lens is then moved back to that position, which is the optimal focusing position, and the camera focusing is complete.

[0093] A5: Repeat steps A1 to A4 to obtain the optimal focus positions of the left and right cameras for each scene depth within the entire depth range. Based on the optimal focus positions of the left and right cameras for each scene depth, obtain the curves showing the relationship between the scene depth (in the calibration process, the scene depth refers to the distance from the first calibration board to the corresponding camera) and the optimal focus position for each of the left and right cameras. Then, based on the curves showing the relationship between the scene depth and the optimal focus position for each of the left and right cameras, obtain the associated focus positions corresponding to the multiple scene depths within the entire depth range. Each associated focus position is composed of the scene depth and the corresponding optimal focus position of the left camera and the optimal focus position of the right camera.

[0094] Specifically, the step of obtaining the optimal focus position of the left and right cameras for each scene depth within the entire depth range is as follows: traverse each scene depth within the entire depth range and obtain the optimal focus position of the left and right cameras for each scene depth.

[0095] In this embodiment, at the calibration station on the production line, a binocular camera is fixed on a precisely positioned motorized slide rail. A solid circular array is placed at the other end of the slide rail as a camera focusing calibration template pattern. The binocular camera is moved to a position g1 closest to the calibration plate and remains stationary. The optical axes of the left and right cameras are perpendicular to the calibration plate, and they focus separately to determine the optimal focusing position Code. L1 and Code R1 The binocular cameras move to the next position at certain intervals, with the left and right cameras stopping at a new distance g2 from the calibration board. The left and right cameras focus separately to determine the optimal focus position Code. L2 and Code R2 By doing this, the optimal focus position of the left and right cameras at each calibrated distance can be determined within the entire depth range.

[0096] The field of view of a camera is directly proportional to the scene depth. For example, when a camera captures a scene at depth g, its field of view (horizontal and vertical directions) is twice that of a camera capturing a scene at depth g / 2. Therefore, in this embodiment, the size of the solid circular array template, the diameter of the dots, and the spacing of the dots used for calibration should increase proportionally with the scene depth. Furthermore, considering production line efficiency and production costs, the entire depth range is divided into several depth intervals, and each depth interval uses a calibration template of a specific size. When the binocular camera enters the next depth interval for calibration, a calibration template adapted to that depth interval needs to be replaced, with its size, dot diameter, and spacing proportional to the scene depth.

[0097] For the left and right cameras, the optimal focus position is searched for at each calibration distance across the entire depth range, thus determining the relationship curve between scene depth and optimal focus position for each of the left and right cameras. Due to errors in the camera assembly process, the optimal focus positions for the same depth may differ between the left and right cameras of each binocular camera system. Therefore, the relationship curve between scene depth and optimal focus position needs to be calibrated separately for each left and right camera.

[0098] Focusing on the same scene depth g i The lens positions of the left and right cameras are recorded as the associated focus positions of the left and right cameras at that depth [g] i Code Li Code Ri Once the relationship curve between scene depth and optimal focus position is calibrated for the entire depth range, a series of associated focus positions are formed. These associated focus positions can be used to achieve fast focusing. Given a known scene depth, the lens can be directly moved to the optimal focus position corresponding to that depth, completing focus in one step and eliminating frequent back-and-forth searching, thus improving focusing speed.

[0099] A6: Based on the scene depths within the entire depth range traversed in step A5, perform monocular calibration on the left and right cameras at each scene depth to obtain the monocular camera parameters of the left and right cameras, and combine the monocular camera parameters of the left and right cameras to calibrate the relative positions of the binocular cameras.

[0100] like Figure 3 As shown, step A6 specifically includes:

[0101] A61: Based on the scene depth in each of the associated focus positions obtained in step A5, the entire depth range is divided into multiple depth intervals, and a second calibration plate is configured in each depth interval.

[0102] In this embodiment, the second calibration board is a black and white checkerboard calibration board. Specifically, the pattern of the black and white checkerboard calibration board is as follows: Figure 4As shown, the black and white checkerboard pattern consists of Q rows × E columns of black and white squares, where Q and E are natural numbers. Each black and white square is the same size and has the same side length, and the black and white squares are arranged alternately in both the horizontal and vertical directions. Since the field of view of the camera is proportional to the scene depth, the size of the checkerboard template, the square size, and the side length used during calibration should be proportionally increased as the scene depth increases. In this embodiment, considering production line efficiency and production costs, the entire depth range is divided into several depth intervals, and each depth interval uses a calibration template of a different size. When the binocular camera enters the next depth interval for calibration, it needs to be replaced with a calibration template adapted to that depth interval, whose size, square size, and side length are proportional to the scene depth.

[0103] A62: The left and right cameras take pictures of the second calibration board from different angles to obtain multiple images of the second calibration board. Based on minimizing the reprojection error, the intrinsic parameters, extrinsic parameters and distortion parameters of the left and right cameras are obtained. The relative position of the binocular cameras is calibrated by combining the extrinsic parameters of the left and right cameras.

[0104] (1) Projection model of monocular camera and its intrinsic and extrinsic parameters

[0105] Here, the homogeneous coordinates of the object point in the reference coordinate system are [XYZ 1], and it is assumed that the homogeneous coordinates of the pixel obtained by the camera are [xy 1]. According to the projection model based on pinhole imaging, the light rays from the object point [XYZ 1] pass through the projection center, i.e., the optical center of the lens, and are projected onto the image along a straight line, resulting in the corresponding imaging pixel point [xy 1].

[0106]

[0107] In the formula, σ is the scale factor. The rotation vector R and translation vector T are the extrinsic parameters of the camera, describing the camera's spatial position in the reference coordinate system. K is the intrinsic parameter of the camera, defined as...

[0108]

[0109] In the formula, f x and f y The focal lengths in the horizontal and vertical directions are c, respectively. x and c y These are the principal points of the image in the horizontal and vertical directions, respectively.

[0110] (2) Monocular camera distortion parameters

[0111] Due to optical distortion inherent in lenses, the actual projected pixels typically exhibit a slight offset in both the radial and tangential directions on the image. Radial distortion refers to inward or outward movement in the radial direction. Tangential distortion refers to the offset of the actual image point in the direction perpendicular to the radial direction, i.e., the tangential direction. The theoretical pixel position [xy] based on the aforementioned projection model is thus affected by distortion and shifts, resulting in a different actual projected position. Simulate using the following relationship

[0112]

[0113]

[0114] In the formula, [k1 k2 k3] are radial distortion parameters, [p1 p2] are tangential distortion parameters, and r 2 =x 2 +y 2 .

[0115] (3) Monocular camera (intrinsic parameters, extrinsic parameters, distortion parameters) calibration

[0116] During calibration, the camera captures images of the calibration template from different angles; here, a black and white checkerboard calibration template is used as an example. Each corner point of a black and white checkerboard generates a corresponding image pixel in the image. Based on the pinhole imaging principle and image distortion projection model, a theoretical imaging position (including the offset caused by distortion) can be calculated for each corner point. The deviation between the actual image point and this theoretical pixel position is called the reprojection error. The parameters of this geometric projection model, i.e., the camera calibration parameters, should minimize the sum of the squares of the reprojection errors of all black and white checkerboard corner points on the calibration template. At this point, the projection model most accurately describes the optical imaging projection process of the camera at this lens position.

[0117]

[0118] Among them, M i To determine the position of the i-th black and white square corner point on the template, where n is the total number of black and white square corner points on the black and white checkerboard template, and m is the pixel position corresponding to that corner point in the image. The reprojection position is the pixel position calculated based on the offset of the grid corner points after central projection due to image distortion. and (Including distortion offset), K is the intrinsic parameter of the camera, R is the rotation vector and T is the translation vector, T is the extrinsic parameter, k = [k1 k2 k3] is the radial distortion parameter, and p = [p1 p2] is the tangential distortion parameter. This equation is optimized using the Levenberg-Marquardt algorithm. After several iterations, the optimization iteration ends when the iteration error is less than a preset threshold. The obtained results, K, R, T, k, and p, represent the intrinsic, extrinsic, and distortion parameters corresponding to the calibration distance, respectively.

[0119] (4) Calibrate the relative positional relationship of the binocular cameras

[0120] After the left and right cameras have completed their monocular camera calibration, the next step in calibrating the binocular cameras is to calculate the relative positional relationship between the left and right cameras.

[0121] When the left and right cameras focus on the same scene depth, the positions of the left and right camera lenses are associated positions corresponding to that scene depth. When calibrating the binocular cameras at their associated lens positions, after the left and right cameras are calibrated as monocular cameras at their respective associated lens positions, the positions of the two cameras relative to the same plane calibration template are known (i.e., the extrinsic parameters of the left and right cameras). The spatial position of the left camera is R. l T l The spatial position of the right camera is R. r T r This allows us to calculate the relative positions of the two cameras. for:

[0122]

[0123] Using the result (or its average value) in equation (6) as the initial value, the single-target positioning parameters of the left and right cameras and the relative positions of the binocular cameras should minimize the sum of squares of the reprojection errors of all black and white grid corner points in the left and right images of the binocular cameras:

[0124]

[0125] Among them, M i To determine the position of the i-th black and white square corner point on the template, where n is the total number of black and white square corner points on the black and white checkerboard template, and m is the pixel position corresponding to that corner point in the image. Let K be the reprojection position, K be the intrinsic parameter of the camera, k be the radial distortion parameter, p be the tangential distortion parameter, l be the subscript for the left camera, and r be the subscript for the right camera. R and T are the rotation and translation vectors of the left camera relative to the right camera, respectively, describing the relative positional relationship of the binocular cameras. Equation (7) is also optimized using the Levenberg-Marquardt algorithm, and the optimized relative positions R and T of the binocular cameras are more accurate.

[0126] A63: Repeat steps A61 to A62 to obtain the intrinsic, extrinsic, and distortion parameters of the left and right cameras, as well as the relative positions of the binocular cameras, at the scene depth in each of the associated focus positions.

[0127] In one specific embodiment, at a shared focus position of the left and right cameras, both cameras simultaneously focus on a scene depth, and the same object appears to be the same size in both camera images. The autofocus binocular camera calibration in this embodiment is performed only when the left and right cameras are at the same shared focus position. This embodiment uses a black and white checkerboard pattern as the calibration template, such as... Figure 4 As shown, the black and white checkerboard is composed of 11 rows × 16 columns of black and white squares; each black and white square is the same size and has the same side length, and the black and white squares are arranged alternately in the horizontal and vertical directions. Based on the number and size of the checkerboard squares in the horizontal and vertical directions in the calibration template, the distribution of the checkerboard squares in the calibration template coordinate system can be determined, and the homogeneous coordinates [XY 1] of the corner points of the black and white squares can be obtained. In this example, at a certain distance between the binocular camera and the calibration board, the lenses of the left and right cameras are pushed to the corresponding depth-related focus position, and about 20 calibration template images are taken from different angles. In addition, the calibration template is placed on the LED panel light, and the calibration pattern and the white background have a strong contrast and are easy to extract. In this example, after the calibration image is taken, the checkerboard area is detected, and the homogeneous coordinates [xy 1] of the projected pixel positions of the corner points of the black and white squares are obtained. According to the projection model based on pinhole imaging, the corner points [XY 1] of the black and white squares of the calibration template are projected onto the image according to the relationship of Equation (1), and the corresponding imaging pixel points [xy 1] are obtained.

[0128] Each black and white grid corner point generates a corresponding image pixel in the image when it is photographed. According to the projection model based on the pinhole imaging principle and the above distortion model, a theoretical imaging position (including the offset caused by distortion) can also be calculated for each grid corner point. The deviation between the actual image point and the theoretical pixel position is called the reprojection error. The parameters of the geometric projection model, that is, the calibration parameters of the camera, should minimize the sum of the squares of the reprojection errors of all black and white grid corner points on the calibration template (Equation (5)). At this time, the projection model most accurately describes the optical imaging projection process of the camera at this lens position. The equation is optimized using the Levenberg-Marquardt algorithm. After several iterations, when the error of the iteration is less than a preset threshold, the optimization iteration ends. The obtained results K, R, T, k, and p are the intrinsic parameters, extrinsic parameters, and distortion parameters corresponding to the calibration distance, respectively.

[0129] When the left and right cameras focus on the same scene depth, the positions of the left and right camera lenses are associated positions corresponding to that scene depth. When the binocular camera is calibrated at the associated lens position, after the left and right cameras are calibrated as monocular cameras at their respective associated lens positions, the positions of the two cameras relative to the same plane calibration template are known. Therefore, the relative positions of the two cameras can be calibrated according to equations (6) and (7). When the binocular camera focuses at different depths, the distance between the lens and the imaging plane changes, and the intrinsic parameters, extrinsic parameters, distortion parameters, and relative positions between the two cameras will change accordingly. Therefore, it is necessary to calibrate the binocular camera at different distances. The above calibration process needs to be repeated once at the associated position at each focusing calibration depth, that is, to calibrate the intrinsic parameters, extrinsic parameters, distortion parameters, and relative positions of the left and right cameras at that depth. The calibration parameters for each focusing distance (i.e., each associated focusing position) are used to perform stereo correction, scene ranging, and depth point cloud generation on the left and right images of the binocular camera at that distance.

[0130] The production line calibration method disclosed in Embodiment 1 of this invention innovatively proposes a calibration step for the relationship curve between scene depth and optimal focus position, and a step for obtaining the associated focus position of the binocular camera.

[0131] The calibration steps for the scene depth vs. optimal focus position curve include: for each scene depth, the autofocus camera lens corresponds to an optimal focus position where the scene image captured is the sharpest. The calibration of the scene depth vs. optimal focus position curve means that, across the entire depth range, the autofocus camera sequentially captures images at each scene depth g. iUsing a calibration template, such as a dot matrix calibration template, to find the optimal focus position corresponding to that depth, the optimal focus position of the lens is determined when the autofocus camera captures each scene depth. i The calibrated scene depth and optimal focus position relationship curve consists of a series of scene depths and their corresponding optimal lens focus positions (e.g., ...). Figure 6 The diagram shows the relationship curves between scene depth and optimal focus position for both the left and right cameras. When the scene depth is known, the lens can be directly moved to the optimal focus position corresponding to that depth based on the calibrated scene depth and optimal focus position curves, achieving focus directly without searching. For autofocus binocular cameras, the scene depth and optimal focus position relationship curves for each camera need to be calibrated separately.

[0132] The steps for establishing the associated focus position for a binocular camera include: after the left and right cameras of the autofocus binocular camera have completed the calibration of the relationship curve between scene depth and optimal focus position, the lens positions of the left and right cameras focusing at the same scene depth can be associated, that is, the associated focus position of the binocular camera at that depth [g]. i Code Li Code Ri ](like Figure 6 As shown, in the same scene at depth g i The following codes are associated with the optimal focus position of the left camera. Li Code for optimal focus position of the right camera Ri The process iterates through different scene depths to obtain the associated focus positions of the dual cameras for each scene depth. When the left and right cameras are at their associated focus positions, both cameras capture the clearest image at that scene depth; the same object appears the same size in both camera images at that moment. During focusing, if the scene depth is known (e.g., calculated based on calibrated dual camera parameters), the lenses of the left and right cameras are pushed to their optimal focus positions according to the calibrated curves showing the relationship between the scene depth and the optimal focus position for each camera, achieving clear focus.

[0133] Example 2:

[0134] Embodiment 2 of this invention discloses a 3D stereoscopic vision shooting method using an autofocus binocular camera. This autofocus binocular camera for recording 3D stereoscopic video requires two calibrations during production. First, the relationship curve between the scene depth and optimal focus position for each of the left and right cameras is calibrated, and a series of associated focus positions are determined based on this curve within the entire depth range. Second, at each associated position, single-target calibration is performed on the left and right cameras, and the relative position of the binocular cameras is calibrated.

[0135] After the autofocus binocular camera completes the above two calibrations on the production line, it can be used online as follows: Figure 7 As shown, the 3D stereoscopic vision imaging method disclosed in Embodiment 2 of the present invention is performed according to the following steps:

[0136] B0: Connect and turn on the autofocus binocular camera;

[0137] B1: The lenses of the left and right cameras are respectively located at the initial position of focusing on the same scene depth, the current scene depth is saved as the contrast depth value, and the left and right cameras respectively capture the first frame image of the current scene;

[0138] B2: Calculate the depth of the current scene based on the monocular camera parameters and the relative position of the binocular camera at the associated focus position corresponding to the depth ratio;

[0139] Specifically, the steps for depth information calculation include: after the binocular cameras are calibrated at a lens-associated position, a set of binocular camera calibration parameters is obtained. Based on this set of calibration parameters, taking a point P on the surface of an object in the scene as an example, the depth information of this point is calculated as follows: Figure 8 As shown, in the left camera coordinate system, its optical center C1 is at the origin, and the coordinates of point P at the corresponding point m1 in the left camera's image are [x1 y1 f1]. In the right camera coordinate system, its optical center C2 is at the origin, and the coordinates of point P at the corresponding point m2 in the right camera's image are [x2 y2 f2]. The coordinates of C1 and m1 in the right camera coordinate system are respectively... and The coordinates of point P are the intersection of the lines connecting C1 and m1 and C2 and m2 in the coordinate system of the right camera. The depth information of point P is the Z-axis coordinate of this intersection.

[0140] In this embodiment, the depth calculation using the binocular camera employs triangulation. Triangulation is a method for determining the 3D depth of an object point based on the triangle formed by the object point and the optical centers of the left and right images, given the binocular camera calibration parameters and the matching relationship between the object point's projection points in the left and right images. Triangulation first requires matching the projection positions of the same object point in the left and right images one-to-one. Methods for finding matching projection points of the same object point in the left and right images include optical flow, feature point matching, and block matching. This system uses the mature and computationally efficient SURF (Speeded Up Robust Features) algorithm for feature point detection. Then, it eliminates inaccurately matched feature point pairs through uniqueness checks (i.e., a feature point in both the left and right images is simultaneously a unique matching point in the other) and RANSAC (RANdom Sampling Consensus, which calculates perspective transformation matrices to filter feature point matching pairs that meet the same perspective conditions). Based on the accurately matched feature point pairs, the binocular camera's internal parameters, and relative positions, the 3D spatial position of the object point corresponding to the feature point is calculated using triangulation.

[0141] B3: Compare the calculated depth of the current scene with the comparison depth value. If they are consistent, proceed to step B6; otherwise, proceed to step B4.

[0142] B4: Based on the calculated depth of the current scene, according to the curve of the relationship between the scene depth and the optimal focus position of the left and right cameras obtained in step A5, push the lenses of the left and right cameras to the optimal focus position corresponding to the depth of the current scene, and save the depth of the current scene as a comparison depth value.

[0143] B5: The left and right cameras respectively capture a new frame of the current scene;

[0144] B6: Based on the monocular camera parameters and the relative position of the binocular camera at the associated focus position corresponding to the contrast depth value, stereo correction is performed on the images captured by the left and right cameras respectively to obtain the left image and the right image. The left image refers to the stereo-corrected image captured by the left camera, and the right image refers to the stereo-corrected image captured by the right camera.

[0145] Ideally, a parallel binocular camera has only a horizontal positional offset, with the optical axes of the two cameras parallel and the object's projection point at the same height in both images. However, due to mechanical errors during camera assembly and lens assembly errors, the optical axes of the left and right cameras are not actually perfectly parallel. Stereo correction primarily performs two tasks: removing image distortion based on distortion parameters calibrated at the associated positions of the binocular cameras at different depths; and aligning the two distortion-free images row-wise based on the intrinsic and extrinsic parameters and relative positions of the left and right cameras at their associated positions, ensuring that any point in one image is on the same row as its corresponding point in the other image. Figure 9a and Figure 9b As shown, in this embodiment, the two left and right images are first projected onto the same plane, and then the images are rotated to align the rows to complete the stereo correction of the binocular images. After stereo correction, the optical axes of the corrected left and right cameras are parallel, the left and right images are distortion-free, coplanar, and aligned in rows. Matching feature points can be found on the same row to calculate depth or stitched into a Side-By-Side format to output a 3D image.

[0146] B7: Stitch the left and right images together into a Side-By-Side format (SBS) and output it in full screen on the 3D display device screen;

[0147] After distortion correction and stereo correction are performed on the left and right images from the binocular cameras, the optical axes of the two cameras are parallel, and the height of the same object point is consistent in both images. The left and right images are distortion-free and row-aligned. At this point, a point in the left image has its corresponding point in the right image in the same row. After row alignment, the projection point of an object point in the left image and its projection point in the right image are on the same row. Matching projection points in the right image only need to be found in that row, not across the entire image.

[0148] After binocular stereo correction, the images from the left and right cameras are arranged side-by-side to generate a stitched image. The left image is on the left side of the stitched image, and the image from the right camera is on the right side. This format is called the side-by-side format of 3D images. When output to 3D display devices, such as VR headsets, the human eye can view the 3D stereoscopic effect of the scene. The stitched left and right images have the following properties: First, the image content of the left and right images is consistent, that is, only the content of the common area of ​​the field of view of the left and right cameras is retained; second, the size of the left and right images is consistent, that is, in addition to the same resolution, the size of the same object is the same in both images; third, the left and right images are captured synchronously, the image acquisition frame rate of the left and right cameras is the same and synchronized, ensuring that the capture time of the left and right images is consistent; fourth, they are row-aligned, that is, after binocular stereo correction, the projection point of the same object point in the left and right images is in the same row. The stitched left and right images are coplanar and row-aligned, with only horizontal parallax, that is, the horizontal position difference of the projection point of the same object point in the left and right images. The closer an object is, the greater the parallax; the farther away an object is, the smaller the parallax.

[0149] To adapt to the display device's resolution, the resolution of the mosaic needs to be enlarged or reduced to make the mosaic the same as the display device's resolution (or have the same horizontal and vertical aspect ratio). Outputting the Side-By-Side format image to AR headsets, VR headsets, MR headsets, and professional 3D displays that support this format allows for full-screen display, enabling both eyes (or those wearing 3D glasses) to observe a clear stereoscopic effect.

[0150] B8: The left and right cameras each capture a new frame of the current scene and return to step B2; until the binocular cameras are turned off and disconnected.

[0151] Based on the 3D stereoscopic vision shooting method of the autofocus binocular camera disclosed in Embodiment 2 above, this 3D stereoscopic vision shooting method can be applied to a 3D stereoscopic vision shooting system of a binocular camera with autofocus function. The lens of the camera with autofocus function is generally fixed on an electronic and mechanical motion mechanism, such as a VCM motor (Voice Coil Motor). The lens is pushed back and forth along the optical axis within the lens barrel by the motion mechanism, changing its position. The change in distance from the lens to the surface of the imaging chip (i.e., the image distance changes) also changes the depth of focus of the camera, thereby enabling focusing for different scene depths. For each scene depth, the lens in the lens barrel of the autofocus camera has a corresponding position, and the parameter values ​​differ at different positions.

[0152] The 3D stereoscopic vision shooting method of the autofocus binocular camera disclosed in Embodiment 2 has the following advantages: (1) Compared with the traditional fixed-focus binocular camera, the method disclosed in Embodiment 2 of this invention utilizes the advantages of the binocular camera's autofocus, and the camera can always adapt to the depth of the scene and maintain focus. Compared with the currently commonly used fixed-focus binocular camera, during the shooting process, it always maintains real-time focus in response to changes in the scene, and the video image is always clear, providing a high-quality 3D stereoscopic vision experience and enhancing product competitiveness. (2) During production, the binocular camera is calibrated at multiple associated lens positions, resulting in multiple sets of binocular camera calibration parameters. Multiple sets of calibration parameters effectively compensate for the parameter differences caused by changes in lens position during the autofocus camera's shooting process, thereby improving the accuracy of stereoscopic correction, optimizing the effect of 3D video, and enhancing the user experience.

[0153] Example 3:

[0154] Embodiment 3 of the present invention discloses a method for calculating dense depth point clouds of an autofocus binocular camera. The autofocus binocular camera also needs to be calibrated in two ways during production: calibration of the scene depth and optimal focus position relationship curves of the left and right cameras respectively, and single-target calibration of the left and right cameras at each associated position and calibration of the relative position of the binocular cameras. The calibration method and steps are the same as those in Embodiment 2, and will not be repeated here.

[0155] After the autofocus binocular camera completes the above two calibrations on the production line, it can be used online as follows: Figure 10 As shown, the method for calculating dense depth point clouds disclosed in Embodiment 3 of the present invention is performed according to the following steps:

[0156] C0: Connect and turn on the autofocus binocular camera;

[0157] C1: The lenses of the left and right cameras are respectively located at the initial position of focusing on the same scene depth, the current scene depth is saved as the contrast depth value, and the left and right cameras respectively capture the first frame image of the current scene;

[0158] C2: Calculate the depth of the current scene based on the monocular camera parameters and the relative position of the binocular camera at the associated focus position corresponding to the contrast depth value;

[0159] The method for calculating depth in this embodiment is the same as in Embodiment 2, and will not be repeated here.

[0160] C3: Compare the calculated depth of the current scene with the comparison depth value. If they are the same, proceed to step C6; otherwise, proceed to step C4.

[0161] C4: Based on the calculated depth of the current scene, according to the curve of the relationship between the scene depth and the optimal focus position of the left and right cameras obtained in step A5, push the lenses of the left and right cameras to the optimal focus position corresponding to the depth of the current scene, and save the depth of the current scene as a comparison depth value.

[0162] C5: The left and right cameras each capture a new frame of the current scene;

[0163] C6: Based on the monocular camera parameters and the relative position of the binocular camera at the associated focus position corresponding to the contrast depth value, stereo correction is performed on the images captured by the left and right cameras respectively to obtain the left image and the right image. The left image refers to the stereo-corrected image captured by the left camera, and the right image refers to the stereo-corrected image captured by the right camera.

[0164] The steps for distortion removal and stereo correction in this embodiment are the same as in Embodiment 2, and will not be repeated here.

[0165] C7: Perform stereo matching between the left and right images to generate a disparity map;

[0166] Stereo matching of left and right camera images from a binocular camera specifically includes: The goal of stereo matching is to find a matching point in the left camera image for each pixel in the right camera image, meaning the two pixels represent the projection of the same object point onto the left and right camera images. After distortion removal and stereo correction, the optical axes of the two cameras are parallel, the left and right images are parallel and coplanar, and the height of the projection point of the same object point is consistent in both images. Since the projection point of the object point in the right image is on the same row as its projection point in the left image, the matching projection point in the left image only needs to be found in that row rather than across the entire image, simplifying the matching of the same point in the left and right images of the binocular camera and making the matching of all pixels in the left and right camera images more efficient. After matching, the disparity Δδ = x1 - x2 of the point in the left and right images can be calculated based on the horizontal x-coordinates of the two corresponding points in the image pixels, where x1 and x2 are the horizontal x-coordinates of the pixel in the right and left camera images, respectively. Based on:

[0167]

[0168] The depth Z of the object point corresponding to the pixel can be calculated, thereby converting the disparity map into a depth map. Here, B is the distance between the optical centers of the left and right cameras, f is the focal length of the camera (here, the left and right cameras are of the same type and have the same focal length), and Δδ is the disparity of the pixel.

[0169] This embodiment uses the Semi-Global Block Matching (SGBM) algorithm to calculate the disparity map based on the corrected left and right camera images. This algorithm has the characteristics of good disparity effect and fast speed. After the matching is completed, the disparity of the point in the left and right images is calculated according to the horizontal position coordinates of the two corresponding points in the pixel of the image. The disparity is calculated for each pixel to obtain the disparity map of the right camera image. According to Equation (8), the depth corresponding to these pixels can be calculated, thereby converting the disparity map into a depth map.

[0170] C8: Calculate the dense depth point cloud of a binocular camera based on the disparity map;

[0171] Dense depth point cloud computing from binocular cameras specifically includes: Point cloud data refers to a set of vectors in a three-dimensional coordinate system. This set of vectors is usually represented in the form of X, Y, and Z three-dimensional coordinates, and is generally used to represent the outer surface shape of an object. In addition to the spatial location information represented by (XYZ), point cloud data can also include the color, grayscale value, segmentation result, etc. of a point. For example, P i ={X i ,Y i Z i If {P1, P2, P3, ...} represents a point in space, then {P1, P2, P3, ...} represents a set of point cloud data. Following the aforementioned method, the third-dimensional depth coordinate Z of each pixel in the right camera coordinate system has been obtained from the depth map of the binocular camera. It is still necessary to calculate the X and Y coordinates of each point in the right camera coordinate system, i.e., generate the three-dimensional coordinates of that point. The method for calculating the X and Y coordinates is as follows:

[0172]

[0173] Among them, c x c y f x f y c represents the calibrated intrinsic parameters of the right camera. x and c y These are the principal points of the right camera image in the horizontal and vertical directions, respectively. x and f y These are the horizontal and vertical focal lengths of the right camera, respectively; x and y are the pixel coordinates of that point in the right image. This calculates the 3D coordinates of a pixel in the right camera image within its coordinate system. Adding the RGB (red, green, blue) color information of the right camera image pixel to each coordinate results in a point cloud description P = {X, Y, Z, R, G, B}. Calculating the 3D coordinates and color information of all pixels in the depth map reconstructs a dense 3D point cloud from a 2D image.

[0174] In this embodiment, after obtaining the third-dimensional depth coordinates of each pixel in the right camera coordinate system from the depth map of the binocular camera, the horizontal and vertical coordinates of the point are calculated according to equation (9), thereby obtaining the three-dimensional coordinates of a pixel in the right camera image in its coordinate system. Each coordinate is then added to the RGB (red, green, and blue) color information of the right camera image pixel to finally obtain the point cloud corresponding to a point. Calculating the three-dimensional coordinates and color information of all pixels in the depth map allows for the reconstruction of dense three-dimensional point cloud information from a two-dimensional image.

[0175] C9: The left and right cameras each capture a new frame of the current scene and return to step C2 until the binocular cameras are turned off and disconnected.

[0176] The method for calculating dense depth point clouds using an autofocus binocular camera disclosed in Embodiment 3 has the following advantages: (1) Compared with traditional fixed-focus binocular cameras, the method disclosed in Embodiment 3 of this invention utilizes the advantages of autofocus in binocular cameras, allowing the camera to adapt to the depth of the scene and maintain focus. Compared with currently used fixed-focus binocular cameras, during the shooting process, it maintains real-time focus in response to changes in the scene, ensuring that the captured image remains clear. This results in more accurate stereo matching of the left and right images, enabling the calculation of high-precision dense depth point clouds and enhancing product competitiveness. (2) During production, the binocular camera is calibrated at multiple associated lens positions, resulting in multiple sets of binocular camera calibration parameters. These multiple sets of calibration parameters effectively compensate for parameter differences caused by changes in lens position during autofocus camera shooting, thereby improving the accuracy of dense depth point clouds, optimizing the effect of 3D applications, and enhancing user experience.

[0177] In the production line calibration method, 3D stereo vision shooting method, and dense depth point cloud calculation method based on an autofocus binocular camera disclosed in Embodiments 1 to 3 of this invention, the calibration, shooting, and calculation are all based on completing the following two calibrations of the autofocus binocular camera on the production line:

[0178] (1) Calibration of the relationship curves between scene depth and optimal focus position for the left and right cameras.

[0179] The basis for calibrating the relationship between scene depth and optimal focus position curve for cameras is that, for each scene depth, the left and right cameras find the lens position that provides the clearest focus at that depth, i.e., the optimal focus position for the left and right camera lenses at that depth. For either the left or right camera, the process of finding the optimal focus position for a given scene depth is as follows: A 2D plane calibration template is placed at that scene depth; in this case, it's a dot matrix calibration template (the size of the calibration template must be proportional to the scene depth. In actual production, considering production line efficiency and cost, several dot matrix calibration templates are placed on the production line, each responsible for a segment of the scene depth). The binocular camera module is fixed in a position directly opposite the calibration template, i.e., the optical axis of the binocular camera is perpendicular to the calibration template. The lenses of the left and right cameras move to a new position in steps, and after reaching that position, each camera takes a static image. After taking the image, the lenses of the left and right cameras move to the next position in steps, and the binocular camera takes another static image. After traversing all the positions of the left and right camera lenses, each camera obtains a series of images of the calibration template taken by the lens at different positions. Calculate the contrast of a specific region in the series of images (this specific region is generally one region, namely the square region in the center of the image; or five regions, namely the square region in the center of the image and four square regions at the 0.5 or 0.8 field of view of the image). The position of the lens with the highest contrast is the optimal focus position of the left or right camera for the depth of the scene.

[0180] For a binocular camera, the process of determining the relationship curve between the scene depth and the optimal focus position of the left and right cameras for the entire depth range is as follows: A 2D plane calibration template is placed at scene depth g. This is a circular dot matrix calibration template (similarly, there are several circular dot matrix calibration templates on the production line, each responsible for a segment of the entire depth range). The binocular camera module is fixed directly opposite the calibration template, meaning the optical axis of the binocular camera is perpendicular to the calibration template. The left and right cameras respectively find the optimal focus position Code for that scene depth. L and Code R For the scene depth, this depth, the optimal focus position of the left camera, and the optimal focus position of the right camera form a related position [g Code]. L Code R The calibration template is moved at certain intervals to a new distance from the binocular cameras to determine the optimal focus positions and associated positions of the left and right camera lenses at the new scene depth. After traversing all scene depths within the entire depth range, the relationship curves between the scene depth and the optimal focus position for each of the left and right cameras, as well as a series of associated focus positions, are generated.

[0181] (2) At each associated location, the left and right cameras perform single-target calibration and calibrate the relative positions of the binocular cameras.

[0182] After calibrating the relationship curves between the scene depth and optimal focus position for each of the left and right cameras, for each scene depth, the depth g and the optimal focus position Code of the left camera are determined. L and the optimal focus position of the right camera (Code) R Form an associated location [g Code] L Code R At this associated position, both the left and right cameras simultaneously focus on the scene depth, and the same object appears at the same size in both camera images. The autofocus binocular camera disclosed in this invention calculates the 3D depth of an object or synthesizes stereoscopic video only when the left and right cameras are at the same associated position.

[0183] After determining the associated focus positions of the autofocus binocular cameras, at each associated focus position, the monocular camera calibration of the left and right cameras and the relative position calibration of the binocular cameras are performed. During calibration, the cameras photograph the calibration template from different angles; here, a black and white checkerboard calibration template is used as an example (similarly, there are several black and white checkerboard calibration templates on the production line, each responsible for a segment of the entire depth range), taking approximately 20 images. Based on the aforementioned method, through mathematical optimization, the camera calibration parameters minimize the sum of squares of the reprojection errors of all black and white checkerboard corner points on the calibration template. At this point, the projection model most accurately describes the optical imaging projection process of the camera at this lens position. The obtained results K, R, T, k, and p are the intrinsic, extrinsic, and distortion parameters of the left and right cameras corresponding to the calibration distance, respectively. After the left and right cameras complete the monocular camera calibration, the relative position relationship of the binocular cameras is determined by combining the extrinsic parameters of the left and right cameras.

[0184] By performing the above two calibrations on the production line, the binocular camera with autofocus can achieve fast focusing, maintain clear focus at all times, and further generate clear, accurately aligned, and high-quality 3D images as well as accurately calculated dense depth point clouds.

[0185] In addition, corresponding to Embodiment 2, Embodiment 4 of the present invention also discloses a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the steps of the 3D stereoscopic vision imaging method of Embodiment 2.

[0186] Corresponding to Embodiment 3, Embodiment 5 of the present invention also discloses a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the steps of the method for calculating dense depth point clouds in Embodiment 3.

[0187] The background section of this invention may include background information about the problems or circumstances surrounding the invention, rather than a description of prior art by others. Therefore, the content included in the background section is not an admission of prior art by the applicant.

[0188] The above description provides a further detailed explanation of the present invention in conjunction with specific / preferred embodiments, and it should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various substitutions or modifications can be made to these described embodiments without departing from the concept of the present invention, and all such substitutions or modifications should be considered within the scope of protection of the present invention. In the description of this specification, the reference to terms such as "an embodiment," "some embodiments," "preferred embodiment," "example," "specific example," or "some examples," etc., indicates that the specific features, structures, materials, or characteristics described in connection with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples. Furthermore, those skilled in the art can combine and integrate different embodiments or examples and features of different embodiments or examples described in this specification without contradiction. Although the embodiments of the present invention and their advantages have been described in detail, it should be understood that various changes, substitutions, and modifications can be made herein without departing from the scope defined by the appended claims.

Claims

1. A method for calculating depth point clouds based on an autofocus binocular camera, characterized in that, include: The binocular camera is calibrated on the production line using steps A1 to A6: A1: A first calibration board is configured at a scene depth position of the binocular camera. The first calibration board is a solid dot array calibration board, and its size, dot diameter and spacing are positively correlated with the scene depth. A2: After moving the lenses of the left and right cameras of the binocular camera to a new position according to a preset step size, the left and right cameras respectively take pictures of the first calibration board to obtain the image of the first calibration board. A3: Repeat step A2 to obtain first calibration board images captured from multiple positions of the left and right camera lenses, so that each of the left and right cameras acquires a set of first calibration board images. A4: Calculate the contrast of the preset area in a set of first calibration board images acquired by the left and right cameras respectively, and determine the position of the lens corresponding to the maximum contrast in each set as the best focus position of the left and right cameras for the current scene depth. A5: Repeat steps A1 to A4 to obtain the optimal focus positions of the left and right cameras for each scene depth within the entire depth range. Based on the optimal focus positions of the left and right cameras for each scene depth, obtain the curves showing the relationship between the scene depth and the optimal focus position of the left and right cameras. Then, based on the curves showing the relationship between the scene depth and the optimal focus position of the left and right cameras, obtain the corresponding associated focus positions for multiple scene depths within the entire depth range. Each associated focus position is composed of the scene depth and the corresponding optimal focus position of the left and right cameras, respectively. This allows the lens to be directly pushed to the optimal focus position corresponding to the scene depth, completing the focusing in one step. A6: Based on the associated focus positions obtained in step A5, at the scene depth of each associated focus position, perform single-target calibration on the left and right cameras to obtain the single-camera parameters of the left and right cameras, and combine the single-camera parameters of the left and right cameras to calibrate the relative position of the binocular cameras, so as to generate a clear, accurately aligned and high-quality 3D image, feature point ranging, and calculate a dense depth point cloud. Step A6 specifically includes: A61: Based on the scene depth in each of the associated focus positions obtained in step A5, the entire depth range is divided into multiple depth intervals. A second calibration board is configured in each depth interval. The second calibration board adopts a black and white checkerboard calibration board, and its size, checkerboard size and spacing are positively correlated with the scene depth. A62: The left and right cameras take pictures of the second calibration board from different angles to obtain multiple images of the second calibration board. Based on minimizing the reprojection error, the intrinsic parameters, extrinsic parameters and distortion parameters of the left and right cameras are obtained. The relative position of the binocular cameras is calibrated by combining the extrinsic parameters of the left and right cameras. A63: Repeat steps A61 to A62 to obtain the intrinsic parameters, extrinsic parameters, and distortion parameters of the left and right cameras, as well as the relative positions of the binocular cameras, at the scene depth in each of the associated focus positions. And perform the following deep point cloud computing steps: C1: The lenses of the left and right cameras are respectively located at the initial position of focusing on the same scene depth, the scene depth is saved as the contrast depth value, and the left and right cameras respectively capture the first frame image of the current scene; C2: Calculate the depth of the current scene based on the monocular camera parameters and the relative position of the binocular camera at the associated focus position corresponding to the contrast depth value; C3: Compare the calculated depth of the current scene with the comparison depth value. If they are the same, proceed to step C6; otherwise, proceed to step C4. C4: Based on the calculated depth of the current scene, according to the multiple associated focus positions obtained in step A5, push the lenses of the left and right cameras to the optimal focus positions corresponding to the depth of the current scene, and save the depth of the current scene as a comparison depth value. C5: The left and right cameras each capture a new frame of the current scene; C6: Based on the calibrated monocular camera parameters and the relative position of the binocular camera at the associated focus position corresponding to the contrast depth value, stereo correction is performed on the images captured by the left and right cameras respectively to obtain the left image and the right image. The left image refers to the stereo-corrected image captured by the left camera, and the right image refers to the stereo-corrected image captured by the right camera. C7: Perform stereo matching between the left and right images to generate a disparity map; C8: Calculate the dense depth point cloud of a binocular camera based on the disparity map; C9: The left and right cameras capture a new frame of the current scene and return to step C2.

2. The method for calculating depth point clouds according to claim 1, characterized in that, The preset area in step A4 includes one area or five areas. One area refers to the square area at the center of the image, and the five areas refer to the square area at the center of the image and four square areas at a field of view of 0.5 or 0.8 in the image.

3. The method for calculating depth point clouds according to claim 1, characterized in that, The formula used in step A4 Calculate the contrast of a preset region in the first calibration plate image, where L max and L min These are the largest and smallest pixel grayscale values ​​in the preset area, respectively.

4. The method for calculating depth point clouds according to claim 1, characterized in that, In step A1, the first calibration plate is perpendicular to the optical axis of the binocular camera, and the entire depth range is divided into multiple depth intervals, each using a calibration template of a different size. When the binocular camera enters the next depth interval for calibration, the calibration template adapted to that depth interval is replaced. In step A2, which is executed before each step A3, the lenses of the left and right cameras are moved to the closest focusing distance. When step A2 is repeated in step A3, the lenses of the left and right cameras are moved sequentially by a preset step size until the lenses of the left and right cameras are at the furthest focusing distance, so as to traverse all positions of the lenses of the left and right cameras.

5. The method for calculating depth point clouds according to claim 1, characterized in that, The entire depth range is divided into multiple depth intervals, each using a calibration template of a different size. When the binocular camera enters the next depth interval for calibration, the calibration template is changed to one that is compatible with that depth interval.

6. The method for calculating depth point clouds according to claim 1, characterized in that, In step C2, the following method is used: based on the accurately matched feature point pairs and the relative positions of the monocular camera parameters and binocular camera of the left and right cameras corresponding to the contrast depth value, the three-dimensional spatial position of the object point corresponding to the feature point is calculated using triangulation to calculate the depth of the current scene.

7. The method for calculating depth point clouds according to claim 1, characterized in that, Step C6 specifically includes: eliminating the distortion of the images captured by the left and right cameras respectively according to the monocular camera parameters of the left and right cameras corresponding to the contrast depth value; then, according to the monocular camera parameters of the left and right cameras corresponding to the contrast depth value and the relative position of the binocular cameras, projecting the distortion-eliminated images captured by the left and right cameras onto the same plane and performing row alignment, so that any point on one image is in the same row as the corresponding point on the other image, so as to obtain the left and right images through stereo correction.

8. The method for calculating depth point clouds according to claim 1, characterized in that, Step C7 specifically includes: Stereo matching is performed on any pixel in the left image along the same row in the right image. After matching, the disparity Δδ of the pixel in the left and right images is calculated based on the horizontal coordinates of the two corresponding pixels in the image: Δδ = x1 - x2. The disparity is calculated for each pixel to obtain a disparity map; where x1 and x2 are the horizontal coordinates of the pixel in the left and right images, respectively.

9. The method for calculating depth point clouds according to claim 1, characterized in that, Step C8 specifically includes: C81: Based on the disparity map, according to the formula Calculate the depth Z of the object point corresponding to the pixel in the image captured by the right camera after stereo correction, so as to convert the disparity map into a depth map; where B is the distance between the optical centers of the left and right cameras, f is the focal length of the left or right camera, and Δδ is the disparity of the pixel in the images captured by the left and right cameras respectively after stereo correction. C82: Obtain the third-dimensional depth coordinate Z of each pixel in the right camera coordinate system based on the depth map, and calculate the X and Y coordinates of each pixel in the right camera coordinate system: In the formula, (c x c y ) is the principal point of the right camera, f x f y y represents the horizontal and vertical focal lengths of the right camera, respectively, and (x, y) represents the pixel coordinates of the corresponding pixel in the right camera. C83: Based on step C82, obtain the three-dimensional coordinates of each pixel in the right image in the right camera coordinate system. Combine the RGB color information of each pixel in the right image to obtain the point cloud of each pixel as P = {X,Y,Z,R,G,B}. Calculate the three-dimensional coordinates and corresponding RGB color information of all pixels in the depth map to obtain the dense depth point cloud of the binocular camera.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when invoked and executed by a processor, cause the processor to perform the steps of the method for calculating a depth point cloud as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Three-dimensional imaging system and three-dimensional image construction method

    CN107872664A

  • An asphalt pavement structure depth detection method based on binocular vision

    CN109919856A

  • Automatic focusing binocular camera calibration method and device

    CN111080705A

  • Autofocus device

    JP2010175696A