Method for determining depth information and electronic equipment

By keeping one image unchanged in a binocular camera, only the other image is corrected, and the optical axis angle and feature matching direction are adjusted, the problems of large calculation and high power consumption caused by the optical axis tilt setting are solved, and the accuracy and efficiency of depth information are improved.

CN120279073AActive Publication Date: 2025-07-08HONOR DEVICE CO LTD

Patent Information

Application Number
CN202311873094.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-08
Estimated Expiration
2043-12-29

AI Technical Summary

Technical Problem

The tilt setting of the optical axis of the binocular camera results in a long time to determine the depth information, high power consumption of the equipment, and large image distortion angle when the polar line is corrected, which increases the calculation amount and power consumption, affecting the accuracy of feature matching and depth calculation.

Method used

By keeping one image unchanged, only the other image is corrected in polar lines, the camera internal reference and rotation matrix are used to adjust the angle of the optical axis, reduce image distortion, feature matching in a specific direction, and depth information is calculated.

Benefits of technology

It reduces the calculation amount and time of feature matching, reduces the power consumption of equipment, improves the calculation accuracy and efficiency of depth information, and avoids the increase in image resolution and invalid areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279073A_ABST
    Figure CN120279073A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method for determining depth information and electronic equipment, the method is executed by the electronic equipment, the electronic equipment comprises a first camera and a second camera, and the method comprises the following steps: obtaining a first image collected by the first camera and a second image collected by the second camera; obtaining camera internal parameters of the first camera and the second camera, and a first rotation matrix and a first translation parameter between the second camera; the first translation parameter comprises a first horizontal translation amount and a first vertical translation amount; correcting the second image according to the internal reference of the camera and the first rotation matrix to obtain a third image; performing feature matching operation based on the first image and the third image, and calculating depth information of each pixel point in the first image; wherein the search direction when the feature matching operation is executed is determined according to the first horizontal translation amount and the first vertical translation amount. The method can reduce the calculation amount of depth calculation and reduce the power consumption of equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electronic technologies, and particularly to a method for determining depth information and an electronic device. Background Art

[0002] With the development of electronic technologies, the shooting effects of both digital cameras and cameras in terminal devices are getting better and better. For example, many cameras are binocular stereo vision depth cameras (referred to as binocular cameras or depth cameras for short). A binocular camera includes two cameras (also known as binoculars), similar to the two eyes of a human. A binocular camera can not only obtain an image of an object to be photographed, but also obtain the depth information of each point of the object to be photographed. Based on the depth information, the image can be further processed, such as defocusing processing, etc., to improve the visual effect of the image.

[0003] Generally, when setting the binoculars of a binocular camera, the optical axes should be as parallel as possible. However, a binocular camera with an inclined optical axis setting has advantages such as high robustness, and thus has important significance in practical applications.

[0004] The inventor found that when a binocular camera has an inclined layout of binocular optical axes, that is, the included angle between the binocular optical axes is relatively large (for example, 45°), the time for the binocular camera to determine the depth information is relatively long, and the device power consumption is relatively high. Summary of the Invention

[0005] This application provides a method for determining depth information and an electronic device, which can reduce the distortion angle of the image during the image correction process, thereby reducing the calculation amount of depth calculation and reducing the device power consumption.

[0006] In a first aspect, this application provides a method for determining depth information, which is executed by an electronic device. The electronic device includes a first camera and a second camera. The method includes: obtaining a first image collected by the first camera and a second image collected by the second camera; obtaining the camera internal parameters of the first camera and the second camera, as well as a first rotation matrix and a first translation parameter between the second cameras; the first translation parameter includes a first horizontal translation amount and a first vertical translation amount; correcting the second image according to the camera internal parameters and the first rotation matrix to obtain a third image; wherein, the third image is consistent with the imaging of the second camera under the target binocular structure, and the included angle between the optical axes of the first camera and the second camera in the target binocular structure is determined according to the ratio of the first horizontal translation amount to the first vertical translation amount; performing a feature matching operation based on the first image and the third image, and calculating the depth information of each pixel point in the first image; wherein, the search direction during the feature matching operation is determined according to the first horizontal translation amount and the first vertical translation amount.

[0007] For the method for determining depth information provided in the first aspect of the present application, the first image remains unchanged, and only the second image is rectified along epipolar lines. In this way, the first image will not be distorted, the resolution of the first image will not increase, and the invalid area of the first image will not increase. Therefore, the computational amount and computational time consumed for feature matching can be reduced, and the device power consumption can be saved. At the same time, the features of the pixel points in the first image will not be damaged, the accuracy of feature matching can be improved, and further the accuracy of depth information calculation can be improved. Moreover, in this method, in the target binocular structure for epipolar rectification, the included angle between the optical axes of the first camera and the second camera is determined according to the ratio of the first horizontal translation amount to the first vertical translation amount. Therefore, the included angle between the two optical axes can be basically kept unchanged, so the adjustment angle for the second image is very small, the distortion angle is very small, the increase in the resolution of the second image is very small, and the increase in the invalid area is also very small. This further reduces the computational amount and computational time consumed during feature matching, and saves the device power consumption. At the same time, since the distortion angle of the second image is small, it is not easy to damage the features of the pixel points in the second image, the accuracy of feature matching is improved, and further the accuracy of depth information calculation is improved.

[0008] In a possible implementation manner, the included angle between the optical axes of the first camera and the second camera in the target binocular structure is equal to the arctangent value of the ratio of the first horizontal translation amount to the first vertical translation amount.

[0009] That is to say, the goal of epipolar rectification is that the included angle between the optical axis of the first camera and the optical axis of the second camera is equal to the arctangent value of the ratio of the first horizontal translation amount to the first vertical translation amount. For example, when the ratio of the first horizontal translation amount to the first vertical translation amount is 1, in the target binocular structure, the included angle between the optical axis of the first camera and the optical axis of the second camera is 45°. Rectifying the second image according to this included angle results in the smallest distortion angle for the second image, further reducing the computational amount and computational time consumed during feature matching, saving the device power consumption, and improving the accuracy of depth calculation.

[0010] In a possible implementation manner, the camera internal parameters include the first camera internal parameters of the first camera and the second camera internal parameters of the second camera; rectifying the second image according to the camera internal parameters and the first rotation matrix to obtain a third image includes: determining a first rectification matrix according to the product of the first camera internal parameters, the first rotation matrix, and the second camera internal parameters; and rectifying the second image based on the first rectification matrix and the first translation parameter to obtain the third image.

[0011] Optionally, the first rectification matrix is obtained by left-multiplying the first rotation matrix by the first camera internal parameters and then right-multiplying by the second camera internal parameters.

[0012] In a possible implementation, based on the first image and the third image, a feature matching operation is performed, including: extracting features from the first image to obtain a plurality of first features; extracting features from the third image to obtain a plurality of third features; for a second feature, searching in the third image along a first search direction with a first search step to determine a fifth feature that matches the second feature among the plurality of third features; the second feature is any one of the plurality of first features, and the first search direction is a straight line with a slope equal to the ratio of the first horizontal translation amount to the first vertical translation amount; determining the second feature and the fifth feature as a pair of features.

[0013] In this implementation, the first search direction is a straight line with a slope equal to the ratio of the first horizontal translation amount to the first vertical translation amount, realizing one-dimensional search and not increasing the computational complexity during feature matching.

[0014] In a possible implementation, the first search step is f x1 *T x ,or f y1 *T y ,or where f x1 represents the value of the focal length of the first camera along the horizontal direction in the first camera intrinsics, and f y1 represents the value of the focal length of the first camera along the vertical direction in the first camera intrinsics, and T x represents the first horizontal translation amount, and T y represents the first vertical translation amount.

[0015] In a possible implementation, calculating the depth information of each pixel point in the first image includes: calculating the correlation information of each pair of features according to the camera intrinsics, the first rotation matrix, and the first translation parameter; determining the disparity corresponding to each pixel point in the first image according to the correlation information of each pair of features; based on the first baseline distance, determining the depth information corresponding to each pixel point in the first image according to the disparity corresponding to each pixel point; the first baseline distance is determined according to the first camera intrinsics, the second camera intrinsics, the first horizontal translation amount, and the first vertical translation amount.

[0016] In a possible implementation, the first baseline distance is determined according to the following formula: where B1 represents the first baseline distance.

[0017] The first baseline distance is the distance of the baseline in the target binocular structure. The first baseline distance matches the above first search direction, and based on the first baseline distance, the depth information corresponding to each pixel point can be accurately calculated.

[0018] In a possible implementation, based on the first image and the third image, a feature matching operation is performed, and the depth information of each pixel point in the first image is calculated, including: correcting the first rotation matrix and the first translation parameter respectively according to the first image and the third image to obtain a second rotation matrix and a second translation parameter; correcting the third image based on the camera internal parameters, the second rotation matrix and the second translation parameter to obtain a fourth image; performing a feature matching operation based on the first image and the fourth image, and calculating the depth information of each pixel point in the first image.

[0019] It can be understood that during the process of a user using an electronic device, the position of one or both cameras in the binocular camera may change due to reasons such as collision or dropping, resulting in a change in the relative position of the two cameras. In this implementation, by correcting the first rotation matrix and the first translation parameter, the accuracy of the calibration data is improved. Subsequently, based on the corrected second rotation matrix and second translation parameter for correction, the accuracy of image correction can be improved, thereby improving the accuracy of feature matching and depth information calculation, and improving the image shooting effect.

[0020] In a possible implementation, correcting the first rotation matrix and the first translation parameter respectively according to the first image and the third image to obtain a second rotation matrix and a second translation parameter includes: calculating an essential matrix according to the first image, the third image, the first rotation matrix and the first translation parameter; performing singular value decomposition on the essential matrix to solve for the second rotation matrix and a third translation parameter; performing scale correction on the third translation parameter according to the first translation parameter to obtain the second translation parameter.

[0021] In this implementation, by calculating the essential matrix and performing singular value decomposition on the essential matrix, an accurate rotation matrix and translation parameter can be determined, realizing accurate correction of the first rotation matrix and the first translation parameter. At the same time, in this implementation, by performing scale correction on the first translation parameter, the accuracy of the translation parameter is further improved, thereby improving the accuracy of subsequent epipolar correction.

[0022] In a possible implementation, calculating the essential matrix according to the first image, the third image, the first rotation matrix and the first translation parameter includes: extracting feature points from the first image to obtain the feature points of the first image; extracting feature points from the third image to obtain the feature points of the third image; performing feature point matching on the feature points of the first image and the feature points of the third image to obtain a feature point matching result; calculating the essential matrix according to the feature point matching result, the first rotation matrix and the first translation parameter.

[0023] In a possible implementation, the second translation parameter includes a second horizontal translation amount and a second vertical translation amount, and the third translation parameter includes a third horizontal translation amount and a third vertical translation amount; performing scale correction on the third translation parameter according to the first translation parameter to obtain the second translation parameter, including: calculating the ratio of the first horizontal translation amount to the third horizontal translation amount, or calculating the ratio of the first vertical translation amount to the third vertical translation amount to obtain a correction coefficient; calculating the product of the correction coefficient and the third horizontal translation amount to obtain the second horizontal translation amount; calculating the product of the correction coefficient and the third vertical translation amount to obtain the second vertical translation amount.

[0024] In this implementation, by the ratio of the translation amounts in the horizontal or vertical direction, the correction coefficient can be simply and quickly determined, improving the accuracy of scale correction.

[0025] In a possible implementation, the camera internal parameters include the first camera internal parameters of the first camera and the second camera internal parameters of the second camera; correcting the third image based on the camera internal parameters, the second rotation matrix, and the second translation parameter to obtain the fourth image, including: determining a second correction matrix according to the product of the first camera internal parameters, the second rotation matrix, and the second camera internal parameters; correcting the third image based on the second correction matrix and the second translation parameter to obtain the fourth image.

[0026] In a possible implementation, the second translation parameter includes a second horizontal translation amount and a second vertical translation amount, and performing a feature matching operation based on the first image and the fourth image, including: extracting features from the first image to obtain a plurality of first features; extracting features from the fourth image to obtain a plurality of fourth features; for a second feature, searching in the third image along a second search direction according to a second search step to determine a sixth feature among the plurality of third features that matches the second feature; the second feature is any one of the plurality of first features, and the second search direction is a straight line with a slope equal to the ratio of the second horizontal translation amount to the second horizontal translation amount; determining the second feature and the sixth feature as a pair of features.

[0027] In a possible implementation, the camera internal parameters include the first camera internal parameters of the first camera and the second camera internal parameters of the second camera; the second search step is f x1 *T x ″, or f y1 *T y ″, or where f x1 represents the value of the focal length of the first camera in the first camera internal parameters along the horizontal direction, f y1 represents the value of the focal length of the first camera in the first camera internal parameters along the vertical direction, T x ″ represents the second horizontal translation amount, and T y ″ represents the second vertical translation amount.

[0028] That is to say, after correcting the first rotation matrix and the first translation parameter, the first search direction and the search step size are further corrected to make the parameters of epipolar rectification consistent with those during feature matching search, improving the accuracy of feature search and further improving the accuracy of depth information calculation.

[0029] In a possible implementation, calculating the depth information of each pixel point in the first image includes: calculating the correlation information of each group of feature pairs according to the camera internal parameters, the second rotation matrix, and the second translation parameter; determining the disparity corresponding to each pixel point in the first image according to the correlation information of each group of feature pairs; determining the depth information corresponding to each pixel point in the first image based on the second baseline distance according to the disparity corresponding to each pixel point; the second baseline distance is determined according to the first camera internal parameters, the second camera internal parameters, the second horizontal translation amount, and the second vertical translation amount.

[0030] In a possible implementation, the second baseline distance is determined according to the following formula: where B2 represents the second baseline distance.

[0031] In this implementation, after the translation parameter is corrected, the second baseline distance is further adjusted synchronously to make the obtained second baseline distance match the corrected translation parameter, improving the accuracy of depth information calculation.

[0032] In a possible implementation, the camera internal parameters include the first camera internal parameters of the first camera and the second camera internal parameters of the second camera; correcting the second image according to the camera internal parameters and the first rotation matrix to obtain a third image includes: determining a first correction matrix according to the product of the first camera internal parameters, the first rotation matrix, and the second camera internal parameters; correcting the second image based on the first correction matrix to obtain a third image.

[0033] That is to say, if further dynamic correction is performed on the third image, translation correction may not be performed during static correction, simplifying the correction process and improving the algorithm operation efficiency.

[0034] In a possible implementation, obtaining the first image collected by the first camera and the second image collected by the second camera includes: collecting a first original image through the first camera; collecting a second original image through the second camera; obtaining the first distortion parameter of the first camera and the second distortion parameter of the second camera; performing distortion correction on the first original image according to the first distortion parameter to obtain the first image; performing distortion correction on the second original image according to the second distortion parameter to obtain the second image.

[0035] In this implementation manner, before performing epipolar rectification, the image is first subjected to distortion correction to eliminate the deviation introduced by the manufacturing precision or assembly process of the camera, improve the accuracy of the image, and thus improve the image display effect.

[0036] In a second aspect, the present application provides a device, which is included in an electronic device and has the function of implementing the behavior of the electronic device in the first aspect and the possible implementation manners of the first aspect. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above functions. For example, a receiving module or unit, a processing module or unit, etc.

[0037] In a third aspect, the present application provides an electronic device, which includes: a processor, a memory, and an interface; the processor, the memory, and the interface cooperate with each other to enable the electronic device to execute any one of the methods in the technical solution of the first aspect.

[0038] Optionally, the electronic device may be an imaging device such as a digital camera or a monitoring device, or may also be a terminal device with a camera such as a mobile phone or a tablet computer. The embodiments of the present application do not make any limitation thereto.

[0039] In a fourth aspect, the present application provides a chip, which includes a processor. The processor is used to read and execute a computer program stored in the memory to execute the methods in the first aspect and any possible implementation manners thereof.

[0040] Optionally, the chip further includes a memory, and the memory is connected to the processor through a circuit or a wire.

[0041] Further optionally, the chip further includes a communication interface.

[0042] In a fifth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the processor is enabled to execute any one of the methods in the technical solution of the first aspect.

[0043] In a sixth aspect, the present application provides a computer program product, which includes: computer program code. When the computer program code runs on an electronic device, the electronic device is enabled to execute any one of the methods in the technical solution of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is a schematic diagram of a camera imaging coordinate system provided by an embodiment of the present application;

[0045] Figure 2 is a schematic diagram of a binocular stereo vision geometric structure before epipolar rectification provided by an embodiment of the present application;

[0046] Figure 3 It is a schematic diagram of image feature matching provided by an embodiment of the present application;

[0047] Figure 4 It is a schematic diagram of the standard geometric structure of a head-up binocular provided by an embodiment of the present application;

[0048] Figure 5 It is another schematic diagram of image feature matching provided by an embodiment of the present application;

[0049] Figure 6 It is a schematic diagram of the principle of calculating depth information based on the standard geometric structure of a head-up binocular provided by an embodiment of the present application;

[0050] Figure 7 It is a schematic diagram of the position of a camera provided by an embodiment of the present application;

[0051] Figure 8 It is another schematic diagram of the position of a camera provided by an embodiment of the present application;

[0052] Figure 9 It is a schematic diagram of the comparison of images before and after epipolar rectification provided by an embodiment of the present application;

[0053] Figure 10 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application;

[0054] Figure 11 It is a software structure block diagram of an electronic device provided by an embodiment of the present application;

[0055] Figure 12 It is a schematic diagram of an application scenario for determining depth information provided by an embodiment of the present application;

[0056] Figure 13 It is a schematic diagram of the flow of a method for determining depth information provided by an embodiment of the present application;

[0057] Figure 14 It is a schematic diagram of the principle of determining depth information provided by an embodiment of the present application;

[0058] Figure 15 It is a schematic diagram of the principle of feature search provided by an embodiment of the present application;

[0059] Figure 16 It is a schematic diagram of the principle of calculating depth information after epipolar rectification provided by an embodiment of the present application;

[0060] Figure 17 It is another schematic diagram of the comparison of images before and after epipolar rectification provided by an embodiment of the present application;

[0061] Figure 18It is a schematic diagram of a search direction provided by an embodiment of the present application;

[0062] Figure 19 It is a schematic diagram of the principle for determining depth information provided by another embodiment of the present application;

[0063] Figure 20 It is a flowchart of a method for determining depth information provided by another embodiment of the present application. Detailed implementation manners

[0064] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. Among them, in the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B; herein, "and / or" is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, "a plurality of" means two or more than two.

[0065] Hereinafter, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", and "third" may explicitly or implicitly include one or more of such features.

[0066] The reference to "an embodiment" or "some embodiments" etc. described in the specification of the present application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present application. Thus, the statements "in an embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in the specification of the present application are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0067] To better understand the embodiments of the present application, the following explanations are provided for the terms or concepts that may be involved in the embodiments.

[0068] 1. Binocular stereo vision

[0069] Binocular stereo vision is an important form of machine vision and a method for obtaining three-dimensional geometric information of an object. This method uses two cameras to acquire two images of the object to be measured from different positions, and calculates the position deviation between corresponding points of the two images based on the parallax principle to obtain the three-dimensional geometric information (including depth information) of the object.

[0070] 2. Camera Intrinsic Parameters, Extrinsic Parameters, and Distortion Parameters

[0071] See Figure 1 , the camera imaging mainly involves four coordinate systems: the world coordinate system, the camera coordinate system, the image coordinate system (also known as the image physical coordinate system), and the pixel coordinate system (also known as the image pixel coordinate system, the image coordinate system, etc.). The imaging process requires several coordinate system conversions. First, the points in space are converted from the world coordinate system to the camera coordinate system, and this process is also called a rigid body transformation. Among them, the points in the world coordinate system can be expressed as (X w , Y w , X w ), and the points in the camera coordinate system can be expressed as (X c , Y c , X c ). Then, the points are converted from the camera coordinate system to the image coordinate system, and this process is also called perspective projection. The points in the image coordinate system can be expressed as (x, y). Then, the points are converted from the image coordinate system to the pixel coordinate system, and this process is also called a secondary transformation. The points in the pixel coordinate system can be expressed as (u, v).

[0072] During the conversion process of coordinate points between various coordinate system conversions, relevant parameters need to be used.

[0073] The internal parameters of the camera, abbreviated as camera intrinsic parameters, refer to the parameters that describe the internal attributes of the camera and are the parameters used to convert coordinate points from the camera coordinate system to the pixel coordinate system. Optionally, the camera intrinsic parameters can include the focal length of the camera, the coordinates of the principal point (optical center, abbreviated as the optical center), etc.

[0074] The external parameters of the camera, abbreviated as camera extrinsic parameters, refer to the parameters that describe the external attributes of the camera and are the parameters used to convert coordinate points from the world coordinate system to the camera coordinate system. The camera extrinsic parameters can include the rotation matrix and translation parameters, etc.

[0075] Due to deviations in manufacturing accuracy and assembly process of the camera lens, distortion will be introduced, resulting in distortion of the original image. In order to correct the distortion, distortion parameters are introduced. Optionally, the distortion parameters can include radial distortion parameters and tangential distortion parameters. It can be understood that the distortion parameters can also be classified as camera intrinsic parameters. In the embodiments of the present application, the camera intrinsic parameters do not include distortion parameters.

[0076] 3. Epipolar Rectification

[0077] The process of a binocular camera determining depth information generally includes: 1) The left and right cameras respectively acquire images, obtaining two images; 2) Respectively match the projection points of each point in space in the two images, that is, perform feature matching on the two images; 3) Based on the feature matching results, calculate the disparity between each pixel point according to the triangulation principle; 4) Determine the depth information according to the disparity.

[0078] In step 2), during feature matching, if the projection points of the same point in the two images are not in the same pixel row, a search needs to be performed in the entire image, that is, a two-dimensional matching search needs to be performed, which has a large amount of calculation, takes a long time, and is prone to false matching points, with poor accuracy. Therefore, it is necessary to perform epipolar rectification on the two images to simplify the two-dimensional matching search to a one-dimensional matching search, so as to save the amount of calculation, shorten the calculation time, eliminate false matching points, and improve the accuracy.

[0079] Epipolar rectification is also called parallel rectification, or stereo rectification, etc. Its essence is to correct the images obtained by a binocular camera with a non-standard binocular stereo vision geometric structure to the binocular standard geometric structure in a frontal view through a series of transformations. In the binocular standard geometric structure in a frontal view, the imaging planes of the left and right cameras are on the same plane and are aligned horizontally.

[0080] Exemplarily, Figure 2 FIG. is a schematic diagram of the binocular stereo vision geometric structure before epipolar rectification provided by an embodiment of the present application. As shown in the figure, the camera coordinate system of the left camera is O C1 -X c1 Y c1 Z c1 , and its corresponding image coordinate system is o1-x1y1, and the pixel coordinate system is u1v1. Among them, the origin O of the camera coordinate system of the left camera C1 is the optical center of the left camera, Z c1 is the optical axis direction of the left camera, and x1 and y1 are respectively parallel to the horizontal (u1) and vertical (v1) directions of the imaging plane of the left camera (hereinafter referred to as the left imaging plane). The origin o1 of the image coordinate system of the left camera is the intersection of the optical axis of the left camera and the left imaging plane. The camera coordinate system of the right camera is O C2 -X c2 Y c2 Z c2 , and its corresponding image coordinate system is u2v2, and the pixel coordinate system is o2-x2y2. Among them, the origin O of the camera coordinate system of the right camera C2 is the optical center of the right camera, Z c2is the optical axis direction of the right camera. x2 and y2 are respectively perpendicular to the horizontal and vertical directions of the imaging plane of the right camera (hereinafter referred to as the right imaging plane). The coordinate origin o2 of the image coordinate system of the right camera is the intersection of the optical axis and the right imaging plane. O C1 The distance between O and o1 is the focal length f1 of the left camera. O C2 The distance between O and o2 is the focal length f2 of the left camera. The optical center O of the left camera C1 and the optical center O of the right camera C2 The line segment between them is the baseline, and the baseline distance can be expressed as B.

[0081] It can be seen that in the non-standard binocular stereo vision geometry structure, the optical axes of the left and right cameras are not parallel, and the left imaging plane and the right imaging plane are not in the same plane. In this case, the same point in space cannot be horizontally aligned in the image obtained by the left camera (referred to as the left image) and the image obtained by the right camera (referred to as the right image). Specifically, see Figure 3 , for example, points a, b, and c in the left image and the right image cannot be horizontally aligned.

[0082] Therefore, it is necessary to perform epipolar rectification on the images collected by the cameras of the binocular camera so that the images are located in the Figure 4 shown frontal binocular standard geometry structure. It can be seen that in the frontal binocular standard geometry structure, the optical axes of the left and right cameras are parallel, the left imaging plane and the right imaging plane are in the same plane and horizontally aligned. In this case, the same point in space is horizontally aligned in the left image and the right image. Specifically, see Figure 5 , in the rectified left image and the rectified right image, all points including points a, b, and c can be horizontally aligned. Therefore, when performing feature matching, only a one-dimensional search needs to be performed in the vertical direction, saving computational effort.

[0083] The specific implementation process of epipolar rectification is to determine the rotation matrix and translation parameters respectively according to the actual positions of the two cameras and the frontal binocular standard geometry structure. Then, the images obtained by the two cameras are rectified according to the rotation matrix and translation parameters.

[0084] All in all, generally, epipolar rectification takes the frontal binocular standard geometry structure as the rectification standard, and rectifies the two images obtained by the binocular respectively so that the two rectified images are located in the frontal binocular standard geometry structure. In this way, subsequent feature matching can be achieved through one-dimensional search based on the frontal binocular standard geometry structure, and depth information can be calculated based on the frontal binocular standard geometry structure.

[0085] Exemplarily, Figure 6Schematic diagram of the principle for calculating depth information based on the standard geometric structure of a head-up binocular in an embodiment of the present application. Refer to Figure 6 , for the standard geometric structure of a head-up binocular, the focal lengths of the left camera and the right camera are the same, and here they are both denoted as f, and the baseline distance is B. According to the basic principle of similar triangles, for any point P in space, the abscissa of its perspective projection point p1 on the left imaging plane is x1, and the abscissa of its perspective projection point p2 on the right imaging plane is x2. Therefore, the depth Z can be derived based on formula (1):

[0086]

[0087] That is, the depth Z can be calculated through formula (2):

[0088]

[0089] where d represents the disparity between point p1 and point p2.

[0090] Next, the technical problems of the present application will be described.

[0091] As introduced above, if the structure of a binocular camera is the standard geometric structure of a head-up binocular, it can save the computational amount of feature matching. Therefore, when setting the cameras of a general binocular camera, the optical axes of the two eyes are arranged as close to parallel as possible. As Figure 7 shown in figure (a) in Figure 7 , taking the binocular camera in a mobile phone as an example, the camera 701 and the camera 702 can be arranged along the x-axis, or, as

[0092] shown in figure (b) in Figure 8 , the camera 701 and the camera 702 are arranged along the y-axis, so that their optical axes are close to parallel.

[0093] However, in practical applications, a binocular camera with the optical axes of the two eyes tilted also has important significance and wide applications. The optical axes of the two eyes being tilted means that there is a certain angle between the optical axes of the two eyes. Exemplarily, as Figure 9 shown in Figure 9 Figure 5 Figure 9 shown in figure (a1) in Figure 9As shown in Figure (b1), the rectified right image obtained after epipolar rectification is as follows Figure 9 shown in Figure (b2). It can be seen that epipolar rectification distorts the image by a large angle. The problems brought about by this are as follows: 1) The larger the distortion angle, the higher the resolution of the rectified image, and the computational complexity of subsequent calculations increases. In extreme cases, the rectified image needs to be split into two parts for separate matching and then stitched together for depth calculation, which is extremely time-consuming and consumes a large amount of device power. 2) The larger the distortion angle, the larger the invalid area. As shown in Figure (b1) and Figure (b2), the grid-filled parts are the invalid areas. The larger the invalid area, the greater the amount of invalid calculations, further increasing the calculation time and device power consumption. 3) The larger the distortion angle, the more likely it is to damage the features of the pixels in the image, resulting in inaccurate subsequent feature matching and inaccurate determination of the final depth information. Figure 9

[0094] In view of this, the present application provides a method for determining depth information. For two images obtained by binocular vision, one of the images (e.g., the left image) is kept unchanged, and the other image (e.g., the right image) is rectified. The goal of rectification is to satisfy the angle determined by the translation parameter (T x / T y ) in the external camera parameters, where T x is the translation amount along the x-axis (also known as the horizontal translation amount), and T y is the translation amount along the y-axis (also known as the vertical translation amount). Then, feature search and matching are performed along the direction of T x / T y , and the depth is calculated. In this way, based on the calibration data, only one image is fine-tuned, greatly reducing the image distortion angle, the resolution of the rectified image, and the size of the invalid area, thereby saving calculation time, reducing device power consumption, and not easily damaging the features of the pixels in the image, improving the accuracy of depth calculation.

[0095] The method for determining depth information provided by the embodiments of the present application can be applied to any electronic device with a binocular camera, including but not limited to digital cameras, mobile phones, tablet computers, wearable devices, vehicle-mounted devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), etc. The embodiments of the present application do not impose any restrictions on the specific type of the electronic device.

[0096] Exemplarily, Figure 10It is a schematic structural diagram of an electronic device 100 provided by an embodiment of the present application. The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0097] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0098] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0099] Among them, the controller may be the nerve center and command center of the electronic device 100. The controller may generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching instructions and executing instructions.

[0100] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can be directly called from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0101] The electronic device 100 can implement the shooting function through the ISP, camera 193, video codec, GPU, display screen 194, application processor, etc.

[0102] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and the light passes through the lens and is transmitted to the camera sensor. The light signal is converted into an electrical signal, and the camera sensor transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be provided in the camera 193.

[0103] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the sensor. The sensor can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The sensor converts the light signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard RGB, YUV, etc. formats. In some embodiments, the electronic device 100 may include at least two cameras 193.

[0104] The software system of the electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservices architecture, or cloud architecture. In the embodiments of this application, the Android system with a layered architecture is taken as an example to exemplarily illustrate the software structure of the electronic device 100.

[0105] In some embodiments, as Figure 11 shown, the system architecture of the electronic device 100 includes an application layer, an application framework layer, a hardware abstraction layer (HAL), a driver layer, and a hardware layer (Hardwork).

[0106] It can be understood thatFigure 11 As an example only, the layers divided in the electronic device 100 are not limited to Figure 11 the layers shown. For example, between the application framework layer and the HAL layer, there may also be included layers such as the Android runtime and system libraries.

[0107] The application layer may include a series of application packages. Such as Figure 11 shown, the application packages may include a camera and other applications. The other applications include, but are not limited to: applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message.

[0108] Generally speaking, the application is developed using the Java language and completed by calling the application programming interface (API) and programming framework provided by the application framework layer.

[0109] The application framework layer provides the application programming interface (API) and programming framework for the applications in the application layer. The application framework layer includes some predefined functions.

[0110] For example, the application framework layer may include a camera access interface. The camera access interface may include a camera service and a camera management. Among them, the camera service can be used to provide an interface for accessing the camera, and the camera management can be used to provide an interface for managing the access to the camera.

[0111] In addition, the application framework layer may also include a content provider, a resource manager, a notification manager, a window manager, a view system, a phone manager, etc. Similarly, the camera application can also call the content provider, the resource manager, the notification manager, the window manager, the view system, etc. according to actual business needs. The embodiments of the present application do not make any restrictions on this.

[0112] The hardware abstraction layer is used to abstract the hardware. For example, it can encapsulate the driver programs in the driver layer and provide an interface for the application framework layer to call, shielding the implementation details of the lower-level hardware.

[0113] For example, the hardware abstraction layer may include a camera hardware abstraction layer (Camera HAL), and of course, may also include other hardware device abstraction layers. The Camera HAL is the core software framework of the camera. The Camera HAL includes an interface module, a camera abstract device, an image processing module, a calibration file, etc. Among them, the interface module, the camera abstract device, and the image processing module are components in the transmission pipeline of image data and control instructions in the Camera HAL. Of course, different components also correspond to different functions. For example, the interface module may be a software interface facing the application framework layer, used for data interaction with the application framework layer. Of course, the interface module can also perform data interaction with other modules in the HAL (such as the camera abstract device and the image processing module). For another example, the camera abstract device may be a software interface facing the driver layer, used for data interaction with the driver layer, such as calling the camera device driver in the driver layer. For another example, the image processing module can process the raw image data returned by the camera device.

[0114] The calibration file stores the original calibration data of each camera of the electronic device. The original calibration data is the parameters of the binocular camera pre-calibrated and stored in the calibration file when the electronic device leaves the factory. The original calibration data may include original camera internal parameters, original camera external parameters, original distortion parameters, etc. Exemplarily, the calibration file may be a binary (Bin) file.

[0115] Among them, the image processing module may include software algorithms, such as an algorithm for determining depth information. This algorithm is used to calculate the depth information of each pixel point according to the images obtained by two cameras according to the method provided in the embodiments of the present application. Optionally, the depth information may be output in the form of a depth map. In a specific embodiment, the image processing module may include a static correction module, a dynamic correction module, a depth estimation module, etc. The static correction module is used to correct the images obtained by two cameras according to the calibration data in the calibration file. The dynamic correction module is used to correct the calibration data in the calibration file according to the images obtained by two cameras, and then further correct the images obtained by static correction based on the corrected calibration data. The depth estimation module is used to estimate the depth of each pixel point in the image based on the image obtained by static correction or the image obtained by dynamic correction to obtain the depth information corresponding to each pixel point.

[0116] In addition, the image processing module may also include nodes or algorithms with other image processing capabilities, which can be specifically referred to the related art and will not be elaborated here.

[0117] The driver layer is used to provide drivers for different hardware devices. For example, the driver layer may include a camera device driver, and of course, may also include other hardware device drivers.

[0118] In addition, the hardware layer includes hardware modules that can be driven, such as camera devices. For example, a camera device includes multiple cameras, such as Camera 1, Camera 2, ..., Camera n, where n is an integer greater than 2. Additionally, the camera device may also include multispectral sensors, etc., which are not limited in the embodiments of this application.

[0119] In this application, by invoking the hardware abstraction layer interface in the hardware abstraction layer, the connection between the application layer and the application framework layer above and the driver layer and the hardware layer below can be achieved, enabling camera data transmission and function control.

[0120] Next, in combination with the capture and photo-taking scenario, the working processes of the software and hardware of the electronic device 100 will be exemplarily described.

[0121] The camera application in the application layer can be displayed on the screen of the electronic device 100 in the form of an icon. When the icon of the camera application is clicked by the user to be triggered, the electronic device 100 starts running the camera application. When the camera application is running on the electronic device 100, the camera application invokes the corresponding interface of the camera application in the application framework layer. Then, by invoking the hardware abstraction layer, the camera device driver is started, and multiple cameras on the electronic device 100 are turned on to capture images through the multiple cameras. Among them, the captured images are stored in the photo gallery.

[0122] During the process of capturing images through multiple cameras 193, the camera application can invoke the image processing module in the application framework layer to process the images obtained by binocular vision to determine the depth information corresponding to each pixel point in the image. Among them, the determined depth information can be used for subsequent blurring processing or other scenarios, which are not limited in the embodiments of this application.

[0123] For ease of understanding, in the following embodiments of this application, an electronic device having the Figure 10 and Figure 11 shown structure will be taken as an example. In combination with the accompanying drawings and application scenarios, the method for determining depth information provided in the embodiments of this application will be specifically described. For ease of understanding, in the following embodiments of this application, among the multiple cameras configured in the electronic device, Camera 1 (also referred to as the first camera) and Camera 2 (also referred to as the second camera) are included, and the determination of depth information based on the images captured by Camera 1 and Camera 2 will be taken as an example for illustration. In other words, Camera 1 and Camera 2 are the two eyes of a binocular camera. Camera 1 and Camera 2 are disposed obliquely, where the translation parameter between Camera 1 and Camera 2 is T = (T x , T y ). T x and T y are both not zero.

[0124] Optionally, camera 1 can be the main camera and camera 2 can be the auxiliary camera. The main camera is the one that sends the display to the display screen. The auxiliary camera is the one that is used to assist the main camera in calculating the depth and does not send the display to the display screen. It should be understood that in some other embodiments, camera 1 can also be the auxiliary camera and camera 2 can be the main camera, and the embodiments of the present application do not make any limitations in this regard.

[0125] First, the application scenario will be described.

[0126] It can be understood that in some scenarios or shooting modes, the camera needs to perform depth calculation to determine the depth information. In this case, the camera application will call the binocular camera to capture images and perform depth calculation. For example, when the shooting mode selected by the user is the portrait mode, the camera application can call the binocular camera to capture images, and after calculating the depth information, perform background blurring on the background to improve the display effect of the person in the image. For another example, when the shooting mode selected by the user is the large aperture mode, the camera application can also call the binocular camera to capture images, and then calculate the depth information, and perform blurring processing on the image based on the depth information to improve the image effect.

[0127] Exemplarily, Figure 12 FIG. is a schematic diagram of an application scenario for determining depth information provided by an embodiment of the present application. As Figure 12 shown in FIG. (a), taking the electronic device as a mobile phone as an example, the mobile phone desktop includes a camera icon 1201. The user clicks on the camera icon to open the camera application, as Figure 12 shown in FIG. (b). Taking the portrait mode as the default mode when the camera application is launched as an example, the user can click on the shooting control 1202 in the Figure 12 interface shown in FIG. (b). The camera application responds to the user's operation and sends a call instruction to camera 1 and camera 2 to simultaneously call camera 1 and camera 2 for shooting. Then, calculate the depth information according to the method provided by the embodiment of the present application, and perform background blurring processing based on the depth information to obtain the Figure 12 image 1203 shown in FIG. (c).

[0128] It should be noted that the above are only several examples of the application scenarios of the method provided by the embodiments of the present application and are not used as limitations. This method can be applied to any shooting scenario that requires determining depth information, and will not be listed one by one here.

[0129] Next, taking the Figure 12 shown application scenario as an example, the method for determining depth information provided by the present application will be described.

[0130] Embodiment 1:

[0131] In the method for determining depth information provided in this embodiment, after correcting the image acquired by camera 2 according to the original calibration data calibrated when the electronic device leaves the factory (referred to as static correction), depth calculation is performed.

[0132] Exemplarily, Figure 13 is a schematic flowchart of a method for determining depth information provided in an embodiment of this application, Figure 14 and is a schematic diagram of the principle for determining depth information provided in an embodiment of this application. Please refer to Figure 13 and Figure 14 together. This method includes:

[0133] S101. In response to a shooting instruction input by the user in the portrait mode, the camera application sends a call instruction to camera 1 and camera 2.

[0134] As described above, the method provided in this embodiment can also be applied to other application scenarios. Therefore, in this step, the operation of triggering the camera application to send a call instruction to camera 1 and camera 2 can also be other operations. Optionally, the instruction input by the user for triggering the execution of the method of this application can be uniformly referred to as a binocular shooting instruction (also referred to as the first instruction). The binocular shooting instruction is used to instruct to capture images through camera 1 and camera 2 in the binocular camera and determine the depth information of each pixel point in the images. The user can input this binocular shooting instruction by clicking the shooting control when the camera application is in the portrait shooting mode.

[0135] Optionally, in different scenarios, the cameras called by the camera application can be the same or different. For example, the electronic device includes three cameras a, b, and c. In scenario 1, the camera application can call cameras a and b, and in scenario 2, the camera application can call cameras b and c. The method provided in the embodiment of this application can be triggered when two cameras that need to be tilted are to be called. Among them, whether they are tilted can be determined according to T x and T y in the external camera parameters. When both T x and T y are not zero, it is considered that the two cameras are tilted. As a possible implementation manner, it can also be further limited that when T x / T y is greater than a preset threshold, the method of this application is triggered to be executed.

[0136] S102. In response to the call instruction of the camera application, camera 1 captures image 1 (also referred to as the first original image).

[0137] According to the call instruction of the camera application, camera 1 starts the shooting function and captures image 1 through the shooting function.

[0138] Among them, the shooting function can be a photographing function or a recording function. That is to say, Image 1 can be a single-frame image captured by Camera 1 through the photographing function, or an image frame in a video frame recorded by Camera 1 through the recording function.

[0139] S103. Camera 1 sends Image 1 to the static correction module in the image processing module.

[0140] S104. Camera 2 responds to the call instruction of the camera application and captures Image 2 (also referred to as the second original image).

[0141] Camera 2 starts the shooting function according to the call instruction of the camera application and captures Image 2 through the shooting function.

[0142] Among them, the shooting function can be a photographing function or a recording function. That is to say, Image 2 can be a single-frame image captured by Camera 2 through the photographing function, or an image frame in a video frame recorded by Camera 2 through the recording function.

[0143] S105. Camera 2 sends Image 2 to the static correction module in the image processing module.

[0144] S106. The static correction module obtains the original calibration data of Camera 1 and Camera 2 from the calibration file.

[0145] In the embodiments of the present application, the original calibration data may include the original camera internal parameters of Camera 1 (also referred to as the first camera internal parameters), the original camera internal parameters of Camera 2 (also referred to as the second camera internal parameters), the original camera external parameters, the original distortion parameters of Camera 1 (also referred to as the first distortion parameters), and the original distortion parameters of Camera 2 (also referred to as the second distortion parameters), etc.

[0146] In the embodiments of the present application, the original camera internal parameters of Camera 1 are denoted as K1. K1 may include the focal length of Camera 1 (denoted as f1) and the optical center coordinate O C1 in its pixel coordinate system. The focal length f1 is converted to the focal length value in pixels in the pixel coordinate system of Camera 1, which can be denoted as f x1 and f y1 . f x1 represents the value of the focal length f1 in the x direction in the pixel coordinate system of Camera 1 (i.e., the value of the focal length of the first camera along the horizontal direction), and f y1 represents the value of the focal length f1 in the y direction in the pixel coordinate system of Camera 1 (i.e., the value of the focal length of the first camera along the vertical direction). The optical center coordinate O C1 is converted to the pixel coordinate system, and the value in the x direction is denoted as u 01 , and the value in the y direction is denoted as v 01 .

[0147] Optionally, the original camera internal parameter K1 can be expressed in matrix form as the following formula (3):

[0148]

[0149] The original camera internal parameter of camera 2 is expressed as K2. K2 is similar to K1 and can be expressed as the following formula (4):

[0150]

[0151] f x2 represents the value of the focal length f2 of camera 2 in the x - direction in its pixel coordinate system, f y2 represents the value of the focal length f2 in the y - direction in the pixel coordinate system of camera 2. 6. u 02 represents the x - direction value of the optical center coordinate O C2 of camera 2 converted to its pixel coordinate system, v 02 represents the x - direction value of the optical center coordinate O C2 of camera 2 converted to its pixel coordinate system in the y - direction.

[0152] The original camera external parameter can include an original rotation matrix (also known as the first rotation matrix, denoted as R) and an original translation parameter (also known as the first translation parameter, denoted as T). Optionally, the original rotation matrix R represents the translation amount required to adjust camera 2 from its actual position to an angle of arctan(T x / T y ) with the optical axis of camera 1. The original translation parameter T represents the translation amount required to adjust camera 2 from its actual position to the position of camera 1. The original translation parameter T=(T x , T y ), where T x is the translation amount in the x - direction (also known as the first horizontal translation amount), equal to the coordinate difference between camera 1 and camera 2 in the x - axis direction. T y is the translation amount in the vertical direction (also known as the first vertical translation amount), equal to the coordinate difference between camera 1 and camera 2 in the y - axis direction. Specifically, reference can be made to Figure 8 .

[0153] S107. The static correction module uses the original distortion parameters of camera 1 to correct the distortion of image 1, obtaining the distortion - corrected image 1 (also known as the first image).

[0154] S108. The static correction module uses the original distortion parameters of camera 2 to correct the distortion of image 2, obtaining the distortion - corrected image 2 (also known as the second image).

[0155] Perform distortion correction on Image 1 and Image 2 respectively to eliminate the deviation introduced by the manufacturing precision or assembly process of the camera, improve the accuracy of the image, and thus improve the image display effect.

[0156] S109. The static correction module calculates a static correction matrix H (also referred to as the first correction matrix) according to the original calibration data of Camera 1 and Camera 2.

[0157] The correction matrix is also referred to as the distortion matrix. In the embodiments of the present application, the correction matrix determined based on the original calibration data in the calibration file is referred to as the static correction matrix, denoted as H.

[0158] Optionally, the static correction matrix H can be calculated by formula (5):

[0159]

[0160] Wherein, K1 represents the original camera internal parameters of Camera 1. R represents the original rotation matrix. K2 represents the original camera internal parameters of Camera 2.

[0161] S110. The static correction module corrects the distortion-corrected Image 2 according to the static correction matrix H and the original translation parameter T to obtain the static-corrected Image 2 (also referred to as the third image).

[0162] It should be noted that the static correction module keeps the distortion-corrected Image 1 unchanged and only corrects the distortion-corrected Image 2.

[0163] The static correction matrix H contains the original camera internal parameters K1 of Camera 1 and the original camera internal parameters K2 of Camera 2. Therefore, the static correction matrix H can correct the errors caused by the internal structures of Camera 1 and Camera 2.

[0164] As described above, the original rotation matrix R represents the rotation amount required to adjust Camera 2 from the actual position to an angle between the optical axis and the optical axis of Camera 1 of arctan(T x / T y ). The original translation parameter T represents the translation amount required to adjust Camera 2 from the actual position to the position of Camera 1. Then, keeping the distortion-corrected Image 1 unchanged, through the correction of the static correction matrix H containing R and the original translation parameter T, the static-corrected Image 2 and the distortion-corrected Image 1 are equivalent to a binocular stereo vision geometric structure with a binocular optical axis angle of arctan(T x / T y ) after correction, and the two imaging planes are aligned in the x direction and the y direction (this structure is also referred to as the target binocular structure). That is, for Figure 2In the binocular stereo vision geometry shown, the right imaging plane is translated and the angle is finely adjusted. In this way, the image rotation angle adjustment is extremely small, and thus the distortion degree of Image 2 is extremely small.

[0165] S111. The static correction module determines the search direction k1 (also known as the first search direction), the search step size step1 (also known as the first search step size), and the baseline distance B1 (also known as the first baseline distance) according to the original calibration data of Camera 1 and Camera 2.

[0166] Among them, the search direction k1 is a straight line with a slope of T x / T y . T x represents the translation amount in the x direction, that is, the coordinate difference between Camera 1 and Camera 2 in the x direction, and T y represents the translation amount in the y direction, that is, the coordinate difference between Camera 1 and Camera 2 in the y direction.

[0167] Optionally, the search step size step1 can be f x1 *T x . Optionally, the search step size step1 can also be f y1 *T y . Optionally, the search step size step1 can also be

[0168] the baseline distance

[0169] S112. The static correction module sends the search direction k1, the search step size step1, and the baseline distance B1 to the depth estimation module.

[0170] S113. The depth estimation module performs feature matching on the distortion-corrected Image 1 and the static-corrected Image 2 according to the search direction k1 and the search step size step1 to obtain multiple groups of feature pairs.

[0171] Please continue to refer to Figure 14 , optionally, the depth estimation module may include a deep learning feature extraction model, a feature matching model, a correlation calculation model, a disparity calculation model, etc.

[0172] The depth estimation module can input the distortion-corrected Image 1 and the static-corrected Image 2 into the deep learning feature extraction model respectively for feature extraction to obtain features. Among them, the feature extracted from the distortion-corrected Image 1 is called Feature F-1 (also known as the first feature), and the feature extracted from the static-corrected Image 2 is called Feature F-2 (also known as the third feature).

[0173] After that, the depth estimation module inputs each feature F-1, each feature F-2, the search direction k1, and the search step size step1 into the feature matching model to match the feature F-2 corresponding to each feature F-1. Specifically, for each feature F-1, the feature matching model searches for the corresponding feature F-2 in the statically corrected image 2 along the search direction k1 according to the search step size step1. Each feature F-1 and the corresponding F-2 are called a pair of features, and the feature matching model outputs all pairs of features.

[0174] Specifically, please refer to Figure 15 , for example, features can be represented by feature blocks. As Figure 15 shown in Figure (a) of x / T y , for any feature block 1501 in the distortion-corrected image 1, the feature matching model can search for the corresponding feature block in the statically corrected image 2 along the search direction k1. The search direction k1 is a straight line with a slope of T As Figure 15 shown in Figure (b) of

[0175] S114. The depth estimation module calculates the correlation (also called correlation information) of each pair of features based on multiple pairs of features and the original calibration data.

[0176] Please continue to refer to Figure 14 , the depth estimation module inputs each group of features, the original camera intrinsics of camera 1, the original camera extrinsics of camera 2, and the original camera extrinsics into the correlation calculation model. The correlation calculation model calculates the correlation between the feature F-1 and the feature F-2 in each pair of features based on the original camera intrinsics of camera 1, the original camera extrinsics of camera 2, and the original camera extrinsics.

[0177] S115. The depth estimation module calculates the disparity d corresponding to each pixel point based on the correlation of each pair of features.

[0178] Please continue to refer to Figure 14 , the depth estimation module inputs the correlation of each pair of features into the disparity calculation model. The disparity calculation model calculates the disparity d corresponding to each pixel point. The disparity d corresponding to all pixel points can be represented in the form of a disparity map. That is, the output of the disparity calculation model can be a disparity map.

[0179] S116. The depth estimation module calculates the depth information Z corresponding to each pixel point based on the disparity d corresponding to each pixel point and the baseline distance B1.

[0180] For example, Figure 16This is a schematic diagram of the principle for calculating depth information after epipolar rectification provided by an embodiment of the present application. It can be understood that after epipolar rectification, the focal lengths of the two cameras are the same, which is uniformly denoted as f.

[0181] Specifically, for any point p1 in the rectified image 1 (i.e., any point P of the object being photographed), its depth information can be calculated according to the above formula (6):

[0182]

[0183] After calculating the depth information Z of each pixel point, the camera application can, based on the depth information, perform background blurring processing on the rectified image 1, etc., to obtain the final image and send it to the display for display, such as Figure 12 shown in Figure (c). The process of background blurring processing, etc. will not be elaborated here too much.

[0184] In this embodiment, the rectified image 1 remains unchanged, and only the rectified image 2 is rectified. It will not distort the image 1, nor will it increase the resolution of the image 1, nor will it increase the invalid area, reducing the computational amount and computational time during subsequent feature matching, and will not destroy the features of the pixel points in the image 1, improving the accuracy of subsequent feature matching, and thus improving the accuracy of depth information calculation. Moreover, in this embodiment, the standard of epipolar rectification, or rather, the goal of epipolar rectification is to make the included angle between camera 1 and camera 2 the angle determined in the original calibration data as T x / T y maintaining its own included angle basically. Therefore, the adjustment angle of the image 2 during the whole process is very small, the distortion angle is very small, the increase in the resolution of the image 2 is also very small, the increased invalid area is very small, further reducing the computational amount and computational time during subsequent feature matching, not easily destroying the features of the pixel points in the image 2, improving the accuracy of subsequent feature matching, and thus improving the accuracy of depth information calculation.

[0185] Exemplarily, Figure 17 This is a schematic diagram of the image comparison before and after epipolar rectification provided by an embodiment of the present application. The original image 1 obtained by camera 1 is as shown in Figure (a1) in Figure 17 , and the original image 2 obtained by camera 2 is as shown in Figure (a2) in Figure 17 . The rectified image 1 after rectification by the method provided by this embodiment is as shown in Figure (b1) in Figure 17 , and the statically rectified image 2 obtained is as shown in Figure (b2) in Figure 17 . By comparing Figure 9 , it can be clearly seen that the method of the embodiment of the present application does not distort the image 1, and the distortion angle of the image 2 is also very small, and the increased invalid area is as shown in Figure 17The black-filled part in Figure (b2) is very small.

[0186] In addition, referring to Figure 18 , the method provided in the embodiment of the present application performs a search along the T x / T y direction, which is a one-dimensional search and does not increase the computational amount.

[0187] Embodiment 2:

[0188] It can be understood that during the process of a user using an electronic device, the position of one or both cameras in the binocular camera may change due to reasons such as collision or dropping, resulting in a change in the relative position of the two cameras. In this case, if correction continues based on the external camera parameters in the original calibration data, the corrected image obtained is inaccurate, thereby making the depth information calculation inaccurate and the image shooting effect poor.

[0189] In view of this, in the method for determining depth information provided in this embodiment, after static correction to obtain the statically corrected Image 2, the original external camera parameters can be calibrated. Then, based on the calibrated external camera parameter data, epipolar correction is performed. After completion, feature matching and depth calculation are carried out. During feature matching, corrections are also made to the search direction, search step size, baseline distance, etc. This method improves the accuracy of epipolar correction, thereby improving the accuracy of feature matching and depth information calculation and enhancing the image shooting effect. The above process is referred to as dynamic correction (or dynamic calibration).

[0190] Referring to Figure 19 , the application scenario of this embodiment and the process of triggering depth calculation can be the same as those in Embodiment 1. In this embodiment, static correction can be performed first, followed by dynamic correction, and depth calculation is carried out after dynamic correction.

[0191] That is to say, first, steps S101 to S110 in the above Embodiment 1 are executed. Specifically, reference can be made to Embodiment 1, which will not be elaborated here.

[0192] Exemplarily, Figure 20 is a schematic flowchart of another method for determining depth information provided in the embodiment of the present application. After step S110, the steps in Figure 20 can be executed, including:

[0193] S201. The static correction module sends the distorted-corrected Image 1, the statically corrected Image 2, and the original calibration data to the dynamic correction module.

[0194] Optionally, the static correction module may not send the original calibration data to the dynamic correction module, and instead, the dynamic correction module obtains the original calibration data from the calibration file. The embodiment of the present application does not make any limitation on this.

[0195] S202. The dynamic correction module extracts feature points from the distortion-corrected image 1 and the statically corrected image 2 respectively, obtaining the feature points of image 1 (also referred to as the feature points of the first image) and the feature points of image 2 (also referred to as the feature points of the third image).

[0196] S203. The dynamic correction module performs feature point matching on the feature points of image 1 and the feature points of image 2.

[0197] S204. The dynamic correction module calculates the essential matrix E based on the feature point matching result and the original camera extrinsic parameters.

[0198] It can be understood that the essential matrix E constrains the relationship of a three-dimensional point P in the world coordinate system under the camera coordinate systems of camera 1 and camera 2. The projection of point P on the imaging plane of camera 1 is denoted as p1, and the projection of point P on the imaging plane of camera 2 is denoted as p2. Then, based on the original translation parameter T and the original rotation matrix R, the essential matrix E can be calculated according to formulas (7) and (8):

[0199] p2 T *E*p1 = 0 (7)

[0200] E = T^R (8)

[0201] S205. The dynamic correction module performs singular value decomposition (SVD) on the essential matrix E to solve the corrected rotation matrix R' and the preliminary translation parameter T' (also referred to as the third translation parameter).

[0202] The essential matrix E inherits the properties of the matrix T^, so its singular values are in the form of [σ σ 0]. Therefore, the rotation matrix R and the translation parameter T can be obtained through SVD decomposition. In this embodiment, the rotation matrix obtained by SVD decomposition is called the corrected rotation matrix, denoted as R'. The translation parameter obtained by SVD decomposition is called the preliminary translation parameter, denoted as T'. T' = (T' x , T' y ). Among them, T' x is also called the third horizontal translation amount, and T' y is also called the third vertical translation amount.

[0203] If the SVD decomposition of the essential matrix E is E = U∑V T , then the results of R' and T' are as follows:

[0204] T'^ = UZU T or T'^ = UZ T U T ;

[0205] R′ = UWV T or R′ = UW T V T ;

[0206] Wherein, W is an orthogonal matrix, Z is a skew-symmetric matrix, If the vector z = (0, 0, 1) T , then z^ = Z.

[0207] It can be understood that the calibration translation parameter T′ is scale-free.

[0208] In the above steps, based on the distorted-corrected image 1 and the statically-corrected image 2, the essential matrix is calculated and SVD decomposition is performed, so as to determine the accurate rotation matrix and translation parameter, and realize the accurate calibration of the original rotation matrix and the original translation parameter.

[0209] S206. The dynamic correction module performs scale calibration on the preliminary translation parameter T′ based on the original translation parameter T to obtain the calibrated translation parameter T″ (also called the second translation parameter).

[0210] Optionally, the calibration coefficient s can be determined according to the original translation parameter T and the preliminary translation parameter T′. Optionally, the calibration coefficient s = T x / T′ x , or s = T y / T′ y .

[0211] After that, based on the calibration coefficient s, the preliminary translation parameter T′ is calibrated to obtain the calibrated translation parameter T″. T″ = sT′. Specifically, T″ = (T″ x , T″ y ), T″ x = sT′ x , T″ y = sT′ y . Wherein, T″ x is also called the second horizontal translation amount, and T″ y is also called the second vertical translation amount.

[0212] The process of the above steps S201 to S206 is also called calibrating the calibration data.

[0213] In this step, through scale calibration, the accuracy of the translation parameter can be improved, and further the accuracy of subsequent epipolar correction can be improved.

[0214] S207. The dynamic correction module calculates the dynamic correction matrix H′ (also called the second correction matrix) according to the original camera internal parameters of camera 1, the original camera internal parameters of camera 2, and the calibrated rotation matrix R′ (also called the second rotation matrix).

[0215] In the embodiments of the present application, the correction matrix determined based on the calibrated data after correction is called the dynamic correction matrix, denoted as H'.

[0216] Optionally, the dynamic correction matrix H' can be calculated by formula (9):

[0217] H' = K1 * R' * K2 -1 (9)

[0218] This step is the same as the principle of step S109 above, the difference is that the rotation matrix used is different, which will not be elaborated here.

[0219] S208. The dynamic correction module corrects the statically corrected image 2 according to the dynamic correction matrix H' and the corrected translation parameter T'', to obtain the dynamically corrected image 2 (also called the fourth image).

[0220] It should be noted that the dynamic correction module keeps the distorted-corrected image 1 unchanged and only corrects the statically corrected image 2.

[0221] This step is the same as the principle of step S110 above, the difference is that the correction matrix used is different, which will not be elaborated here.

[0222] Optionally, if dynamic correction is further performed after static correction, in step S110 above, translation correction may not be performed through the original translation parameter T, but translation correction is performed through the corrected translation parameter in this step. This can simplify the static correction process and improve the algorithm operation efficiency.

[0223] S209. The dynamic correction module determines the correction search direction k2, the correction search step size step2, and the correction baseline distance B2 according to the original camera internal parameters of camera 1, the original camera internal parameters of camera 2, the corrected rotation matrix R', and the corrected translation parameter T''.

[0224] Among them, the correction search direction k2 is a straight line with a slope of T x ' represents the translation amount in the x direction in the corrected translation parameter, and T y ' represents the translation amount in the y direction in the corrected translation parameter.

[0225] Optionally, the correction search step size step2 is f x1 *T x ''. Optionally, the correction search step size step2 can also be f y1 *T y ''. Optionally, the correction search step size step2 can also be

[0226] Correction baseline distance

[0227] This step is the same as the principle of step S111 above. The difference is that different rotation matrices, translation parameters, etc. are used, so the obtained results are different, which will not be elaborated here.

[0228] S210. The dynamic correction module sends the correction search direction k2, the correction search step size step2, and the correction baseline distance B2 to the depth estimation module.

[0229] S211. The depth estimation module performs feature matching on the distorted-corrected image 1 and the dynamically-corrected image 2 according to the correction search direction k2 and the correction search step size step2, and obtains multiple groups of feature pairs.

[0230] Among them, the features obtained by extracting features from the dynamically-corrected image 2 can be represented as feature F-3, also known as the fourth feature.

[0231] S212. The depth estimation module calculates the correlation of each group of feature pairs according to the multiple groups of feature pairs, the original camera intrinsics of camera 1, the original camera intrinsics of camera 2, the correction rotation matrix R′, and the correction translation parameter T″.

[0232] Among them, the correction rotation matrix R′ and the correction translation parameter T″ are also collectively referred to as the corrected camera extrinsics.

[0233] S213. The depth estimation module calculates the disparity d corresponding to each pixel point according to the correlation of each group of feature pairs.

[0234] S214. The depth estimation module calculates the depth information Z corresponding to each pixel point according to the disparity d corresponding to each pixel point and the correction baseline distance B2.

[0235] The above steps S210 to S214 are the same as the principle of steps S112 to S116 in Embodiment 1. The difference is that different parameters are used, which will not be elaborated here.

[0236] It can be understood that the principle of performing epipolar correction and depth calculation in Embodiment 2 is the same as that in Embodiment 1. Therefore, Embodiment 2 has all the beneficial effects of Embodiment 1, which will not be elaborated here. In addition, as described above, Embodiment 2 can correct the original camera extrinsics, improve the accuracy of epipolar correction, and further improve the accuracy of feature matching and depth information calculation, improve the effect of the finally obtained image, and improve the user experience.

[0237] The above text has introduced in detail an example of the method for determining depth information provided by the embodiments of the present application. It can be understood that in order for an electronic device to implement the above functions, it includes corresponding hardware and / or software modules for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to exceed the scope of the present application.

[0238] The embodiments of the present application can divide the functional modules of the electronic device according to the above method examples. For example, each functional module can be divided corresponding to each function, such as a detection unit, a processing unit, a display unit, etc., or two or more functions can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation.

[0239] It should be noted that all relevant contents of each step involved in the above method embodiments can be cited in the function description of the corresponding functional module, and will not be repeated here.

[0240] The electronic device provided in this embodiment is used to execute the above method for determining depth information, so it can achieve the same effect as the above implementation method.

[0241] In the case of adopting an integrated unit, the electronic device may further include a processing module, a storage module, and a communication module. Among them, the processing module can be used to control and manage the actions of the electronic device. The storage module can be used to support the electronic device to execute stored program codes and data, etc. The communication module can be used to support the communication between the electronic device and other devices.

[0242] Among them, the processing module can be a processor or a controller. It can implement or execute various exemplary logical blocks, modules, and circuits described in combination with the disclosure content of the present application. The processor can also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, and so on. The storage module can be a memory. The communication module can specifically be a device for interacting with other electronic devices, such as a radio frequency circuit, a Bluetooth chip, a Wi-Fi chip, etc.

[0243] In one embodiment, when the processing module is a processor and the storage module is a memory, the electronic device involved in this embodiment can be a device having Figure 10 the structure shown.

[0244] The embodiment of the present application also provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the processor is caused to execute the method for determining depth information in any of the above embodiments.

[0245] The embodiment of the present application also provides a computer program product. When the computer program product runs on a computer, the computer is caused to execute the above-related steps to implement the method for determining depth information in the above embodiment.

[0246] In addition, the embodiment of the present application also provides a device, which may specifically be a chip, a component or a module. The device may include a processor and a memory connected to each other. The memory is used to store computer execution instructions. When the device runs, the processor may execute the computer execution instructions stored in the memory, so that the chip executes the method for determining depth information in each of the above method embodiments.

[0247] Among them, the electronic device, computer-readable storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here.

[0248] Through the description of the above embodiments, those skilled in the art can understand that for the convenience and brevity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0249] In several embodiments provided by the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0250] The unit described as a separation component may or may not be physically separated. The component shown as a unit may be a single physical unit or multiple physical units, that is, it may be located in one place, or it may be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0251] In addition, each functional unit in various embodiments of the present application may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0252] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs and other various media that can store program codes.

[0253] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for determining depth information, the method being executed by an electronic device, the electronic device including a first camera and a second camera, characterized in that, The method includes: Obtaining a first image collected by the first camera and a second image collected by the second camera; Obtaining the camera internal parameters of the first camera and the second camera, as well as a first rotation matrix and a first translation parameter between the second cameras; the first translation parameter includes a first horizontal translation amount and a first vertical translation amount; Correcting the second image according to the camera internal parameters and the first rotation matrix to obtain a third image; wherein, the third image is consistent with the imaging of the second camera in the target binocular structure, and the included angle between the optical axes of the first camera and the second camera in the target binocular structure is determined according to the ratio of the first horizontal translation amount to the first vertical translation amount; Performing a feature matching operation based on the first image and the third image, and calculating the depth information of each pixel point in the first image; wherein, the search direction during the feature matching operation is determined according to the first horizontal translation amount and the first vertical translation amount.

2. The method according to claim 1, characterized in that, The included angle between the optical axes of the first camera and the second camera in the target binocular structure is equal to the arctangent value of the ratio of the first horizontal translation amount to the first vertical translation amount.

3. The method according to claim 1 or 2, characterized in that, The camera internal parameters include the first camera internal parameter of the first camera and the second camera internal parameter of the second camera; the correcting the second image according to the camera internal parameters and the first rotation matrix to obtain a third image includes: Determining a first correction matrix according to the product of the first camera internal parameter, the first rotation matrix, and the second camera internal parameter; Correcting the second image based on the first correction matrix and the first translation parameter to obtain the third image.

4. The method according to claim 3, characterized in that The performing a feature matching operation based on the first image and the third image includes: Performing feature extraction on the first image to obtain a plurality of first features; Performing feature extraction on the third image to obtain a plurality of third features; For a second feature, searching in the third image along a first search direction at a first search step to determine a fifth feature that matches the second feature among the plurality of third features; the second feature is any one of the plurality of first features, and the first search direction is a straight line with a slope equal to the ratio of the first horizontal translation amount to the first vertical translation amount; Determining the second feature and the fifth feature as a pair of features.

5. The method according to claim 4, characterized in that, The first search step size is f x1 *T x , or f y1 *T y , or Among them, f x1 represents the value of the focal length of the first camera along the horizontal direction in the first camera internal parameters, f y1 represents the value of the focal length of the first camera along the vertical direction in the first camera internal parameters, T x represents the first horizontal translation amount, T y represents the first vertical translation amount.

6. The method according to claim 5, characterized in that, The calculating the depth information of each pixel point in the first image includes: Calculating the correlation information of each pair of features according to the camera internal parameters, the first rotation matrix, and the first translation parameter; Determining the disparity corresponding to each pixel point in the first image according to the correlation information of each pair of features; Based on a first baseline distance, determining the depth information corresponding to each pixel point in the first image according to the disparity corresponding to each pixel point; the first baseline distance is determined according to the first camera internal parameter, the second camera internal parameter, the first horizontal translation amount, and the first vertical translation amount.

7. The method according to claim 6, characterized in that, The first baseline distance is determined according to the following formula: Among them, B1 represents the first baseline distance.

8. The method according to claim 1 or 2, characterized in that Based on the first image and the third image, performing a feature matching operation and calculating the depth information of each pixel point in the first image, including: According to the first image and the third image, respectively correcting the first rotation matrix and the first translation parameter to obtain a second rotation matrix and a second translation parameter; Based on the camera internal parameters, the second rotation matrix, and the second translation parameter, correcting the third image to obtain a fourth image; Based on the first image and the fourth image, performing the feature matching operation and calculating the depth information of each pixel point in the first image.

9. The method according to claim 8, wherein The step of respectively correcting the first rotation matrix and the first translation parameter according to the first image and the third image to obtain a second rotation matrix and a second translation parameter includes: Calculating an essential matrix according to the first image, the third image, the first rotation matrix, and the first translation parameter; Performing a singular value decomposition on the essential matrix to solve for the second rotation matrix and a third translation parameter; According to the first translation parameter, performing a scale correction on the third translation parameter to obtain the second translation parameter.

10. The method according to claim 9, wherein The step of calculating the essential matrix according to the first image, the third image, the first rotation matrix, and the first translation parameter includes: Performing feature point extraction on the first image to obtain the feature points of the first image; Performing feature point extraction on the third image to obtain the feature points of the third image; Performing feature point matching on the feature points of the first image and the feature points of the third image to obtain a feature point matching result; Calculating the essential matrix according to the feature point matching result, the first rotation matrix, and the first translation parameter.

11. The method according to claim 9 or 10, characterized in that The second translation parameter includes a second horizontal translation amount and a second vertical translation amount, and the third translation parameter includes a third horizontal translation amount and a third vertical translation amount; the step of performing a scale correction on the third translation parameter according to the first translation parameter to obtain the second translation parameter includes: Calculating the ratio of the first horizontal translation amount to the third horizontal translation amount, or calculating the ratio of the first vertical translation amount to the third vertical translation amount to obtain a correction coefficient; Calculating the product of the correction coefficient and the third horizontal translation amount to obtain the second horizontal translation amount; Calculating the product of the correction coefficient and the third vertical translation amount to obtain the second vertical translation amount.

12. The method according to any one of claims 8 to 11, characterized in that The camera internal parameters include the first camera internal parameters of the first camera and the second camera internal parameters of the second camera; the step of correcting the third image based on the camera internal parameters, the second rotation matrix, and the second translation parameter to obtain a fourth image includes: Determining a second correction matrix according to the product of the first camera internal parameters, the second rotation matrix, and the second camera internal parameters; Based on the second correction matrix and the second translation parameter, correcting the third image to obtain the fourth image.

13. The method according to any one of claims 8 to 12, characterized in that The second translation parameter includes a second horizontal translation amount and a second vertical translation amount. Performing the feature matching operation based on the first image and the fourth image includes: Performing feature extraction on the first image to obtain a plurality of first features; Performing feature extraction on the fourth image to obtain a plurality of fourth features; For a second feature, searching in the third image along a second search direction with a second search step to determine a sixth feature that matches the second feature among the plurality of third features; the second feature is any one of the plurality of first features, and the second search direction is a straight line with a slope equal to the ratio of the second horizontal translation amount to the second horizontal translation amount; Determining the second feature and the sixth feature as a pair of features.

14. The method according to claim 13, wherein The camera internal parameters include the first camera internal parameters of the first camera and the second camera internal parameters of the second camera; the second search step size is f x1 *T x ″, or f y1 *T y ″, or Among them, f x1 represents the value of the focal length of the first camera along the horizontal direction in the first camera internal parameters, f y1 represents the value of the focal length of the first camera along the vertical direction in the first camera internal parameters, T x ″ represents the second horizontal translation amount, T y ″ represents the second vertical translation amount.

15. The method according to claim 14, wherein Calculating the depth information of each pixel point in the first image includes: Calculating the correlation information of each pair of features according to the camera internal parameters, the second rotation matrix, and the second translation parameter; Determining the disparity corresponding to each pixel point in the first image according to the correlation information of each pair of features; Based on a second baseline distance, determining the depth information corresponding to each pixel point in the first image according to the disparity corresponding to each pixel point; the second baseline distance is determined according to the first camera internal parameter, the second camera internal parameter, the second horizontal translation amount, and the second vertical translation amount.

16. The method according to claim 15, wherein The second baseline distance is determined according to the following formula: where B2 represents the second baseline distance.

17. The method according to any one of claims 8 to 16, characterized in that The camera internal parameters include the first camera internal parameter of the first camera and the second camera internal parameter of the second camera; correcting the second image according to the camera internal parameters and the first rotation matrix to obtain the third image includes: Determining a first correction matrix according to the product of the first camera internal parameter, the first rotation matrix, and the second camera internal parameter; Based on the first correction matrix, correcting the second image to obtain the third image.

18. The method according to any one of claims 1 to 17, characterized in that, Obtaining the first image collected by the first camera and the second image collected by the second camera includes: Collecting a first original image through the first camera; Collecting a second original image through the second camera; Obtaining the first distortion parameter of the first camera and the second distortion parameter of the second camera; Performing distortion correction on the first original image according to the first distortion parameter to obtain the first image; Performing distortion correction on the second original image according to the second distortion parameter to obtain the second image.

19. An electronic device, characterized in that, Includes: A processor, a memory, and an interface; The processor, the memory, and the interface cooperate with each other to enable the electronic device to execute the method according to any one of claims 1 to 18.

20. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium. When the computer program is executed by the processor, the processor calls instructions to enable the electronic device to execute the method according to any one of claims 1 to 18.

Citation Information

Patent Citations

  • Binocular camera disparity map based road obstacle detection system and method

    CN103679707A

  • Binocular panoramic image acquisition method and device

    CN107666606A

  • Medical equipment machine vision image processing method and computer readable storage medium

    CN114331855A

  • Method and device for determining target position information and storage medium

    CN116723264A

  • Normalized-metadata generation apparatus, object occlusion detection apparatus, and methods thereof

    WO2018062647A1

Cited By

  • Generating depth information in automotive imaging systems

    US20260127753A1