Method of determining depth information and electronic device

By using epipolar correction methods for binocular cameras and adjusting the optical axis angle to reduce computation and power consumption, the problem of long depth information determination time and high power consumption caused by optical axis tilt settings is solved, achieving more efficient and accurate depth information calculation.

CN120279073BActive Publication Date: 2026-01-16HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311873094.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2026-01-16
Estimated Expiration
2043-12-29

AI Technical Summary

Technical Problem

The tilted optical axis setting of binocular cameras results in long depth information determination time and high device power consumption, which is difficult to solve effectively with existing technologies.

Method used

By acquiring the camera images and intrinsic parameters of the binocular camera, the second camera image is corrected using a rotation matrix and translation parameters. The first image remains unchanged, and only the second image is subjected to epipolar correction. The optical axis angle is adjusted to reduce computation and power consumption, thereby improving the accuracy of feature matching.

Benefits of technology

This reduces the computational load and time for feature matching, lowers device power consumption, and improves the accuracy and efficiency of depth information computation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279073B_ABST
    Figure CN120279073B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a method for determining depth information and an electronic device. The method is performed by the electronic device, and the electronic device comprises a first camera and a second camera. The method comprises: obtaining a first image collected by the first camera and a second image collected by the second camera; obtaining camera intrinsic parameters of the first camera and the second camera, and a first rotation matrix and a first translation parameter between the second camera; the first translation parameter comprises a first horizontal translation and a first vertical translation; rectifying the second image according to the camera intrinsic parameters and the first rotation matrix to obtain a third image; performing a feature matching operation based on the first image and the third image, and calculating depth information of each pixel point in the first image; wherein a search direction in the feature matching operation is determined according to the first horizontal translation and the first vertical translation. The method can reduce the calculation amount of depth calculation and reduce the power consumption of the device.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of electronics, in particular to a method for determining depth information and an electronic device. BACKGROUND

[0002] With the development of electronic technology, the shooting effect of digital cameras and cameras in terminal devices is getting better and better. For example, many cameras are binocular stereo vision depth cameras (referred to as binocular cameras or depth cameras). Binocular cameras include two cameras (also known as binoculars), similar to human eyes. Binocular cameras can not only obtain images of the photographed object, but also obtain depth information of each point of the photographed object. Based on the depth information, the image can be further processed, such as blurring processing, to improve the visual effect of the image.

[0003] Generally, the optical axes of the binoculars of a binocular camera are as parallel as possible when being set. However, binocular cameras with obliquely arranged optical axes have the advantages of robustness and are of great significance in practical applications.

[0004] The inventors have found that when the optical axes of the binoculars of a binocular camera are arranged obliquely, that is, the included angle between the optical axes of the binoculars is large (for example, 45°), the binocular camera takes a long time to determine the depth information, and the power consumption of the device is high. SUMMARY

[0005] The present application provides a method for determining depth information and an electronic device, which can reduce the distortion angle of the image in the image correction process, thereby reducing the calculation amount of depth calculation and reducing the power consumption of the device.

[0006] In a first aspect, the present application provides a method for determining depth information, which is executed by an electronic device including a first camera and a second camera. The method includes: obtaining a first image collected by the first camera and a second image collected by the second camera; obtaining camera intrinsic parameters of the first camera and the second camera, and a first rotation matrix and a first translation parameter of the second camera relative to the first camera; the first translation parameter includes a first horizontal translation and a first vertical translation; correcting the second image according to the camera intrinsic parameters and the first rotation matrix to obtain a third image; wherein the third image is consistent with the imaging of the second camera under a target binocular structure, and the included angle of the optical axes of the first camera and the second camera in the target binocular structure is determined according to the ratio of the first horizontal translation to the first vertical translation; performing a feature matching operation based on the first image and the third image, and calculating the depth information of each pixel point in the first image; wherein the search direction when performing the feature matching operation is determined according to the first horizontal translation and the first vertical translation.

[0007] The method for determining depth information provided in the first aspect of the present application keeps the first image unchanged and only performs epipolar rectification on the second image. In this way, the first image is not distorted, the resolution of the first image is not increased, and the invalid area of the first image is not increased. Therefore, the calculation amount and the calculation time of feature matching can be reduced, and the device power consumption can be saved. At the same time, the features of the pixel points in the first image are not damaged, the accuracy of feature matching is improved, and the accuracy of depth information calculation is further improved. Moreover, in the method, the included angle between the optical axis of the first camera and the optical axis of the second camera in the target binocular structure is determined according to the ratio of the first horizontal translation amount to the first vertical translation amount. Therefore, the included angle between the two optical axes can be kept basically unchanged. As a result, the adjustment angle of the second image is small, the distortion angle of the second image is small, the resolution of the second image is increased slightly, and the invalid area of the second image is increased slightly. The calculation amount and the calculation time of feature matching are further reduced, the device power consumption is saved, the distortion angle of the second image is small, the features of the pixel points in the second image are not easily damaged, the accuracy of feature matching is improved, and the accuracy of depth information calculation is further improved.

[0008] In a possible implementation manner, the included angle between the optical axis of the first camera and the optical axis of the second camera in the target binocular structure is equal to the inverse tangent value of the ratio of the first horizontal translation amount to the first vertical translation amount.

[0009] That is, the target of the epipolar rectification is that the included angle between the optical axis of the first camera and the optical axis of the second camera is equal to the inverse tangent value of the ratio of the first horizontal translation amount to the first vertical translation amount. For example, in the case where the ratio of the first horizontal translation amount to the first vertical translation amount is 1, the included angle between the optical axis of the first camera and the optical axis of the second camera in the target binocular structure is 45°. The second image is corrected according to the included angle. The distortion angle of the second image is the smallest, the calculation amount and the calculation time of feature matching are further reduced, the device power consumption is saved, and the accuracy of depth calculation is improved.

[0010] In a possible implementation manner, the camera intrinsic parameters include first camera intrinsic parameters of the first camera and second camera intrinsic parameters of the second camera; and the second image is rectified to obtain the third image according to the camera intrinsic parameters and the first rotation matrix, including: determining a first rectification matrix according to the product of the first camera intrinsic parameters, the first rotation matrix and the second camera intrinsic parameters; and rectifying the second image based on the first rectification matrix and the first translation parameter to obtain the third image.

[0011] Optionally, the first rectification matrix is obtained by multiplying the first camera intrinsic parameters on the left of the first rotation matrix and multiplying the second camera intrinsic parameters on the right.

[0012] In a possible implementation manner, the feature matching operation is performed based on the first image and the third image, including: performing feature extraction on the first image to obtain a plurality of first features; performing feature extraction on the third image to obtain a plurality of third features; for a second feature, searching in the third image along a first search direction according to a first search step to determine a fifth feature that matches the second feature from the plurality of third features; the second feature is any one of the plurality of first features, and the first search direction is along a direction of a straight line with a slope equal to a ratio of the first horizontal translation amount to the first vertical translation amount; and the second feature and the fifth feature are determined as a feature pair.

[0013] In this implementation manner, the first search direction is along a direction of a straight line with a slope equal to a ratio of the first horizontal translation amount to the first vertical translation amount, one-dimensional search is implemented, and the calculation amount during feature matching is not increased.

[0014] In a possible implementation manner, the first search step is , or , or ; wherein, denotes a value of a focal length of the first camera along a horizontal direction in the first camera intrinsic parameter, denotes a value of a focal length of the first camera along a vertical direction in the first camera intrinsic parameter, denotes the first horizontal translation amount, denotes the first vertical translation amount.

[0015] In a possible implementation manner, the depth information of each pixel point in the first image is calculated, including: calculating correlation information of each feature pair according to the camera intrinsic parameter, the first rotation matrix and the first translation parameter; determining a parallax corresponding to each pixel point in the first image according to the correlation information of each feature pair; and determining depth information corresponding to each pixel point in the first image based on the first baseline distance and the parallax corresponding to each pixel point. The first baseline distance is determined according to the first camera intrinsic parameter, the second camera intrinsic parameter, the first horizontal translation amount and the first vertical translation amount.

[0016] In a possible implementation manner, the first baseline distance is determined according to the following formula: B1= ; wherein, B1 denotes the first baseline distance.

[0017] The first baseline distance is a distance of a baseline in a target binocular structure. The first baseline distance is matched with the above-mentioned first search direction, and the depth information corresponding to each pixel point can be accurately calculated based on the first baseline distance.

[0018] In a possible implementation, the feature matching operation is performed and the depth information of each pixel point in the first image is calculated based on the first image and the third image, including: correcting the first rotation matrix and the first translation parameter respectively according to the first image and the third image to obtain a second rotation matrix and a second translation parameter; rectifying the third image based on the camera intrinsic parameter, the second rotation matrix and the second translation parameter to obtain a fourth image; performing the feature matching operation and calculating the depth information of each pixel point in the first image based on the first image and the fourth image.

[0019] It can be understood that in the process of using the electronic device, the position of one or both cameras in the binocular camera may change due to collision, falling, etc., resulting in a change in the relative position of the two cameras. In this implementation, the accuracy of the calibration data is improved by correcting the first rotation matrix and the first translation parameter. Then, based on the corrected second rotation matrix and the second translation parameter, the accuracy of image rectification is improved, and then the accuracy of feature matching and depth information calculation is improved, and the image shooting effect is improved.

[0020] In a possible implementation, the first rotation matrix and the first translation parameter are corrected according to the first image and the third image to obtain a second rotation matrix and a second translation parameter, including: calculating an essential matrix according to the first image, the third image, the first rotation matrix and the first translation parameter; singular value decomposition is performed on the essential matrix to obtain the second rotation matrix and a third translation parameter; and the third translation parameter is corrected in scale according to the first translation parameter to obtain the second translation parameter.

[0021] In this implementation, by calculating the essential matrix and performing singular value decomposition on the essential matrix, the accurate rotation matrix and translation parameter can be determined to accurately correct the first rotation matrix and the first translation parameter. Meanwhile, in this implementation, the accuracy of the translation parameter is further improved by correcting the first translation parameter in scale, and then the accuracy of subsequent epipolar rectification is improved.

[0022] In a possible implementation, the essential matrix is calculated according to the first image, the third image, the first rotation matrix and the first translation parameter, including: performing feature point extraction on the first image to obtain feature points of the first image; performing feature point extraction on the third image to obtain feature points of the third image; performing feature point matching on the feature points of the first image and the feature points of the third image to obtain a feature point matching result; and calculating the essential matrix according to the feature point matching result, the first rotation matrix and the first translation parameter.

[0023] In a possible implementation manner, the second translation parameter comprises a second horizontal translation amount and a second vertical translation amount, and the third translation parameter comprises a third horizontal translation amount and a third vertical translation amount; the third translation parameter is scale-corrected according to the first translation parameter to obtain the second translation parameter, including: calculating a ratio of the first horizontal translation amount to the third horizontal translation amount, or calculating a ratio of the first vertical translation amount to the third vertical translation amount to obtain a correction coefficient; calculating a product of the correction coefficient and the third horizontal translation amount to obtain the second horizontal translation amount; and calculating a product of the correction coefficient and the third vertical translation amount to obtain the second vertical translation amount.

[0024] In this implementation manner, the correction coefficient can be determined simply and quickly through the ratio of the translation amount in the horizontal direction or the vertical direction, and the accuracy of the scale correction is improved.

[0025] In a possible implementation manner, the camera intrinsic parameter comprises a first camera intrinsic parameter of the first camera and a second camera intrinsic parameter of the second camera; the third image is corrected based on the camera intrinsic parameter, the second rotation matrix and the second translation parameter to obtain the fourth image, including: determining a second correction matrix according to a product of the first camera intrinsic parameter, the second rotation matrix and the second camera intrinsic parameter; and correcting the third image based on the second correction matrix and the second translation parameter to obtain the fourth image.

[0026] In a possible implementation manner, the second translation parameter comprises a second horizontal translation amount and a second vertical translation amount; the feature matching operation is performed based on the first image and the fourth image, including: performing feature extraction on the first image to obtain a plurality of first features; performing feature extraction on the fourth image to obtain a plurality of fourth features; for a second feature, searching in the third image along a second search direction according to a second search step to determine a sixth feature that matches the second feature in the plurality of third features; the second feature is any one of the plurality of first features, and the second search direction is along a direction of a straight line with a slope equal to a ratio of the second horizontal translation amount to the third horizontal translation amount; and the second feature and the sixth feature are determined as a feature pair.

[0027] In a possible implementation manner, the camera intrinsic parameter comprises a first camera intrinsic parameter of the first camera and a second camera intrinsic parameter of the second camera; the second search step is , or , or ; wherein, f represents a value of a focal length of the first camera in the first camera intrinsic parameter along the horizontal direction, f represents a value of the focal length of the first camera in the first camera intrinsic parameter along the vertical direction, f represents the second horizontal translation amount, f represents the second vertical translation amount.

[0028] In other words, after correcting the first rotation matrix and the first translation parameter, the first search direction and search step size are further corrected to ensure that the parameters of epipolar correction are consistent with the parameters during feature matching search, thereby improving the accuracy of feature search and thus improving the accuracy of depth information calculation.

[0029] In one possible implementation, calculating the depth information of each pixel in the first image includes: calculating the correlation information of each feature pair based on camera intrinsics, a second rotation matrix, and a second translation parameter; determining the disparity corresponding to each pixel in the first image based on the correlation information of each feature pair; and determining the depth information corresponding to each pixel in the first image based on the disparity corresponding to each pixel, using a second baseline distance; wherein the second baseline distance is determined based on the first camera intrinsics, the second camera intrinsics, a second horizontal translation amount, and a second vertical translation amount.

[0030] In one possible implementation, the second baseline distance is determined according to the following formula: B2 = Where B2 represents the second baseline distance.

[0031] In this implementation, after the translation parameters are corrected, the second baseline distance is further adjusted synchronously so that the obtained second baseline distance matches the corrected translation parameters, thereby improving the accuracy of depth information calculation.

[0032] In one possible implementation, the camera intrinsics include first camera intrinsics of a first camera and second camera intrinsics of a second camera; the second image is corrected based on the camera intrinsics and a first rotation matrix to obtain a third image, including: determining a first correction matrix based on the product of the first camera intrinsics, the first rotation matrix, and the second camera intrinsics; and correcting the second image based on the first correction matrix to obtain the third image.

[0033] In other words, if the third image is further dynamically corrected, translation correction can be omitted during static correction, simplifying the correction process and improving the efficiency of the algorithm.

[0034] In one possible implementation, acquiring a first image captured by a first camera and a second image captured by a second camera includes: acquiring a first original image through the first camera; acquiring a second original image through the second camera; acquiring a first distortion parameter of the first camera and a second distortion parameter of the second camera; performing distortion correction on the first original image based on the first distortion parameter to obtain the first image; and performing distortion correction on the second original image based on the second distortion parameter to obtain the second image.

[0035] In this implementation, before epipolar correction, the image is first distorted to eliminate deviations introduced by the camera due to manufacturing precision or assembly process, thereby improving image accuracy and thus improving image display effect.

[0036] Secondly, this application provides an apparatus included in an electronic device, which has the function of implementing the behaviors of the electronic device in the first aspect and possible implementations thereof. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above-described functions. For example, a receiving module or unit, a processing module or unit, etc.

[0037] Thirdly, this application provides an electronic device, which includes a processor, a memory, and an interface; the processor, memory, and interface cooperate with each other to enable the electronic device to execute any one of the methods in the first aspect of the technical solution.

[0038] Optionally, the electronic device can be an imaging device such as a digital camera or a monitoring device, or a terminal device with a camera such as a mobile phone or a tablet computer. This application embodiment does not limit it in any way.

[0039] Fourthly, this application provides a chip including a processor. The processor is used to read and execute a computer program stored in a memory to perform the methods in the first aspect and any possible implementation thereof.

[0040] Optionally, the chip may also include memory, which is connected to the processor via circuitry or wires.

[0041] Alternatively, the chip may also include a communication interface.

[0042] Fifthly, this application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform any one of the methods in the first aspect of the technical solution.

[0043] Sixthly, this application provides a computer program product, which includes computer program code that, when executed on an electronic device, causes the electronic device to perform any one of the methods in the first aspect of the technical solution. Attached Figure Description

[0044] Figure 1 This is a schematic diagram of a camera imaging coordinate system provided in an embodiment of this application;

[0045] Figure 2 This is a schematic diagram of the geometric structure of binocular stereo vision before epipolar correction provided in an embodiment of this application;

[0046] Figure 3 is an example of image feature matching provided by an embodiment of the present application;

[0047] Figure 4 is an example of a planar binocular standard geometry provided by an embodiment of the present application;

[0048] Figure 5 is another example of image feature matching provided by an embodiment of the present application;

[0049] Figure 6 is an example of a principle of calculating depth information based on a planar binocular standard geometry provided by an embodiment of the present application;

[0050] Figure 7 is an example of camera position provided by an embodiment of the present application;

[0051] Figure 8 is another example of camera position provided by an embodiment of the present application;

[0052] Figure 9 is an example of image comparison before and after epipolar rectification provided by an embodiment of the present application;

[0053] Figure 10 is an example of structure of an electronic device provided by an embodiment of the present application;

[0054] Figure 11 is an example of software structure of an electronic device provided by an embodiment of the present application;

[0055] Figure 12 is an example of application scenario of determining depth information provided by an embodiment of the present application;

[0056] Figure 13 is an example of flow of a method of determining depth information provided by an embodiment of the present application;

[0057] Figure 14 is an example of principle of determining depth information provided by an embodiment of the present application;

[0058] Figure 15 is an example of principle of feature search provided by an embodiment of the present application;

[0059] Figure 16 is an example of principle of calculating depth information after epipolar rectification provided by an embodiment of the present application;

[0060] Figure 17 is another example of image comparison before and after epipolar rectification provided by an embodiment of the present application;

[0061] Figure 18is a search direction schematic diagram provided by an embodiment of the present application;

[0062] Figure 19 is another example of a principle diagram for determining depth information provided by an embodiment of the present application;

[0063] Figure 20 is a flowchart of another example of a method for determining depth information provided by an embodiment of the present application. DETAILED DESCRIPTION

[0064] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application. In the description of the embodiments of the present application, unless otherwise specified, " / " represents the meaning of or, for example, A / B can represent A or B; "and / or" in this document only describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.

[0065] Hereinafter, the terms "first", "second", "third" are only used for description purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first", "second", "third" can explicitly or implicitly include one or more of the features.

[0066] In the present application, the reference "one embodiment" or "some embodiments" means that in one or more embodiments of the present application, the specific features, structures or characteristics described in connection with the embodiment are included. Therefore, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in another some embodiments" and the like appearing in the present application specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "include", "contain", "have" and their variants mean "include but not limited to", unless otherwise specifically emphasized.

[0067] In order to better understand the embodiments of the present application, the following explains the terms or concepts that may be involved in the embodiments.

[0068] 1. Binocular stereo vision

[0069] Binocular stereo vision is an important form of machine vision, which is a method of obtaining three-dimensional geometric information of an object. The method uses two cameras to obtain two images of the measured object from different positions, and calculates the positional deviation between the corresponding points of the two images based on the principle of parallax to obtain the three-dimensional geometric information (including depth information) of the object.

[0070] 2. Camera intrinsic parameters, camera extrinsic parameters and distortion parameters

[0071] Reference Figure 1 Camera imaging mainly involves four coordinate systems: a world coordinate system, a camera coordinate system, an image coordinate system (also referred to as an image physical coordinate system), and a pixel coordinate system (also referred to as an image pixel coordinate system, a video coordinate system, etc.). Several coordinate system conversions are required in the imaging process. First, a point in space is converted from the world coordinate system to the camera coordinate system, which is also referred to as rigid body transformation. The point in the world coordinate system can be represented as (X, Y, Z). , , The point in the camera coordinate system can be represented as (x, y, z). , , Then, the point is converted from the camera coordinate system to the image coordinate system, which is also referred to as perspective projection. The point in the image coordinate system can be represented as (x, y).

[0072] During the conversion of the coordinate points between the various coordinate systems, relevant parameters are required.

[0073] The internal parameters of the camera, referred to as camera intrinsic parameters, refer to parameters describing the internal properties of the camera, which are used for converting the coordinate points from the camera coordinate system to the pixel coordinate system. Optionally, the camera intrinsic parameters can include the focal length of the camera, the principal point (optical center, referred to as optical center) coordinates, etc.

[0074] The external parameters of the camera, referred to as camera extrinsic parameters, refer to parameters describing the external properties of the camera, which are used for converting the coordinate points from the world coordinate system to the camera coordinate system. The camera extrinsic parameters can include a rotation matrix and a translation parameter.

[0075] Due to manufacturing precision and assembly process deviations, the lens of the camera will introduce distortion, resulting in distortion of the original image. In order to correct the distortion, distortion parameters are introduced. Optionally, the distortion parameters can include radial distortion parameters and tangential distortion parameters. It can be understood that the distortion parameters can also be classified as camera intrinsic parameters. In the embodiments of the present application, the camera intrinsic parameters do not include distortion parameters.

[0076] 3. Epipolar rectification

[0077] The process of determining depth information by the binocular camera generally includes: 1) obtaining two images by respectively acquiring images by the left and right cameras; 2) respectively matching the projection points of each point in space in the two images, that is, performing feature matching on the two images; 3) calculating the parallax between each pixel point based on the feature matching result according to the principle of triangulation; and 4) determining the depth information according to the parallax.

[0078] In step 2), when performing feature matching, if the projection points of the same point in the two images are not in the same pixel row, searching needs to be performed in the whole image, that is, two-dimensional matching search needs to be performed, which is large in calculation amount, time-consuming, and prone to false matching points and poor in accuracy. Therefore, epipolar rectification needs to be performed on the two images to simplify the two-dimensional matching search into one-dimensional matching search, so as to save the calculation amount, shorten the calculation time, eliminate the false matching points, and improve the accuracy.

[0079] The epipolar rectification is also called parallel rectification, or stereoscopic rectification, and the essence is to rectify the images obtained by the binocular camera with a non-standard binocular stereo vision geometry structure to a standard parallel binocular geometry structure through a series of transformations. In the standard parallel binocular geometry structure, the imaging planes of the left and right cameras are on the same plane and are aligned along the horizontal direction.

[0080] Exemplarily, Figure 2 An example of a binocular stereo vision geometry structure before epipolar rectification is provided for the embodiments of the present application. As shown in the figure, the camera coordinate system of the left camera is , the image coordinate system thereof is , and the pixel coordinate system thereof is . Among them, the coordinate origin of the camera coordinate system of the left camera is the optical center of the left camera, is the optical axis direction of the left camera, , are respectively parallel to the horizontal and vertical directions of the imaging plane (hereinafter referred to as the left imaging plane) of the left camera. The coordinate origin of the image coordinate system of the left camera is the intersection of the optical axis of the left camera and the left imaging plane. The camera coordinate system of the right camera is , the image coordinate system thereof is , and the pixel coordinate system thereof is . Among them, the coordinate origin of the camera coordinate system of the right camera is the optical center of the right camera, is the optical axis direction of the right camera, , These are perpendicular to the horizontal and vertical directions of the imaging plane of the right camera (hereinafter referred to as the right imaging plane), respectively. The origin of the image coordinate system of the right camera... This is the intersection of the optical axis and the right imaging plane. and The distance between them is the focal length of the left camera. , and The distance between them is the focal length of the left camera. The optical core of the left camera Optical core of the right camera The line segment between them is the baseline, and the distance between the baselines can be represented as B.

[0081] As can be seen, in a non-standard binocular stereo vision geometry, the optical axes of the left and right cameras are not parallel, and the left and right imaging planes are not on the same plane. Therefore, the images of the same point in space captured by the left camera (referred to as the left image) and the right camera (referred to as the right image) cannot be aligned horizontally. For details, see [link to documentation]. Figure 3 For example, points a, b, and c in the left and right images cannot be aligned horizontally.

[0082] Therefore, epipolar correction is needed to ensure that the images captured by the binocular camera are aligned with the epipolar coordinates. Figure 4 The standard geometry of a head-up binocular camera is shown. It can be seen that in this geometry, the optical axes of the left and right cameras are parallel, and the left and right imaging planes are on the same plane and aligned horizontally. This means that the same point in space is horizontally aligned in both the left and right images. For details, see [link to documentation]. Figure 5 In the corrected left and right images, all points, including points a, b, and c, can be aligned horizontally. Therefore, during feature matching, only the vertical direction needs to be searched, saving computation.

[0083] The specific implementation process of epipolar correction involves determining the rotation matrix and translation parameters based on the actual positions of the two cameras and the standard geometry of the binocular head-up view. Then, the images acquired by the two cameras are corrected according to the rotation matrix and translation parameters.

[0084] In summary, epipolar correction typically uses the standard binocular geometry as the correction standard, correcting both images acquired by the binoculars separately to ensure that the corrected images lie within the standard binocular geometry. This allows for subsequent feature matching through one-dimensional search based on the standard binocular geometry, and the calculation of depth information based on that geometry.

[0085] For example,Figure 6 An example of a principle of calculating depth information based on a head-up binocular standard geometry is provided for an embodiment of the present application. Referring to FIG. 1, Figure 6 , a head-up binocular standard geometry, the focal lengths of the left camera and the right camera are the same, denoted as f herein, and the baseline distance is B. According to the basic principle of similar triangles, for any point P in space, the horizontal coordinate of the perspective projection point p1 of P in the left imaging plane is , and the horizontal coordinate of the perspective projection point p2 of P in the right imaging plane is Therefore, the depth Z can be derived based on formula (1):

[0086] (1)

[0087] That is, the depth Z can be calculated by formula (2):

[0088] (2)

[0089] where d represents the disparity between points p1 and p2.

[0090] The technical problem of the present application is described below.

[0091] As introduced above, if the structure of a binocular camera is a head-up binocular standard geometry, it can save the calculation amount of feature matching. Therefore, when setting the cameras, the optical axes of the binoculars of a general binocular camera are preferably arranged in parallel. For example, as shown in FIG. 2(a), taking a binocular camera in a mobile phone as an example, the cameras 701 and 702 can be arranged along the x-axis, or as shown in FIG. 2(b), the cameras 701 and 702 are arranged along the y-axis, so that the optical axes of the two cameras are close to parallel. Figure 7 Figure 7 However, in actual applications, cameras with oblique optical axes of binoculars also have important significance and wide applications. The oblique arrangement of the optical axes of binoculars means that the optical axes of the binoculars have a certain angle. For example, as shown in FIG. 3, taking a mobile phone as an example, the optical axes of the cameras 701 and 702 can have an angle of 45°.

[0092] It can be understood that the larger the angle between the optical axes of the binoculars, the larger the rotation (also known as warp) angle of the image during epipolar rectification. Therefore, for binocular cameras with oblique optical axes, the warp angle of the image after epipolar rectification is self-evident. Referring to FIG. 4, in a specific embodiment, the original left image obtained by a binocular camera with an optical axis angle of 45° is shown in FIG. 4(a1), and the original right image obtained is shown in FIG. 4(a2). Figure 8

[0093] It can be understood that the larger the angle between the optical axes of the binoculars, the larger the rotation (also known as warp) angle of the image during epipolar rectification. Therefore, for binocular cameras with oblique optical axes, the warp angle of the image after epipolar rectification is self-evident. Referring to FIG. 4, in a specific embodiment, the original left image obtained by a binocular camera with an optical axis angle of 45° is shown in FIG. 4(a1), and the original right image obtained is shown in FIG. 4(a2). Figure 9 Figure 9 Figure 5 ​​​​As shown in Figure (a2). The corrected left image obtained after epipolar correction is as follows. Figure 9 As shown in Figure (b1), the corrected right image obtained after epipolar correction is as follows: Figure 9 As shown in Figure (b2), epipolar correction introduces a large distortion angle into the image. This leads to the following problems: 1) The larger the distortion angle, the higher the resolution of the corrected image, increasing the computational load for subsequent calculations. It may even require splitting the corrected image into two parts, matching them separately, and then stitching them together to calculate the depth, resulting in extremely high computation time and device power consumption. 2) The larger the distortion angle, the larger the invalid region, such as... Figure 9 The grid-filled areas in Figures (b1) and (b2) represent invalid regions. The larger the invalid region, the greater the amount of unnecessary computation, further increasing computation time and device power consumption. 3) The larger the distortion angle, the easier it is to destroy the features of pixels in the image, leading to inaccurate subsequent feature matching, and consequently, inaccurate determination of the final depth information.

[0094] To address this, this application provides a method for determining depth information. For two images acquired through binoculars, one image (e.g., the left image) is kept unchanged, while the other image (e.g., the right image) is corrected. The correction aims to satisfy the condition that the angle between the optical axes of the two cameras is equal to the translation parameter in the camera's extrinsic parameters. The determined angle, This is the x-axis translation (also known as the horizontal translation). This is the y-axis translation (also known as the vertical translation). Then along... The system performs feature search and matching along the direction and calculates the depth. In this way, based on the calibration data, only one image is fine-tuned, greatly reducing the image distortion angle, the resolution of the corrected image, and the size of the invalid region. This saves computation time, reduces device power consumption, and is less likely to damage the features of pixels in the image, thus improving the accuracy of depth calculation.

[0095] The method for determining depth information provided in this application can be applied to any electronic device with a stereo camera, including but not limited to digital cameras, mobile phones, tablets, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), etc. This application does not impose any restrictions on the specific type of electronic device.

[0096] For example, Figure 10Fig. 1 is a structural schematic diagram of an electronic device 100 provided by an embodiment of the present application. The electronic device 100 can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a loudspeaker 170A, a receiver 170B, a microphone 170C, a headset interface 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 can include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0097] It can be understood that the structure illustrated by the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can include more or fewer components than those illustrated, or combine certain components, or split certain components, or different arrangement of components. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0098] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices, or can be integrated in one or more processors.

[0099] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to instruction operation codes and timing signals, and complete the control of fetching instructions and executing instructions.

[0100] The processor 110 can also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can hold instructions or data that the processor 110 has just used or is using in a loop. If the processor 110 needs to use the instructions or data again, it can be called directly from the memory. This avoids repeated access and reduces the waiting time of the processor 110, thus improving the efficiency of the system.

[0101] The electronic device 100 can implement the photographing function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc.

[0102] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, and the optical signal is converted into an electrical signal. The camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also optimize the algorithm of the noise, brightness, and skin color of the image. The ISP can also optimize the exposure, color temperature, and other parameters of the shooting scene. In some embodiments, the ISP can be arranged in the camera 193.

[0103] The camera 193 is used to capture still images or videos. Objects generate optical images through lenses and project them onto photosensitive elements. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into a standard RGB, YUV, etc. format image signal. In some embodiments, the electronic device 100 can include at least two cameras 193.

[0104] The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservice architecture, or a cloud architecture. The embodiment of the present application takes the Android system with a layered architecture as an example to exemplarily illustrate the software structure of the electronic device 100.

[0105] In some embodiments, as shown in Figure 11 The system architecture of the electronic device 100 includes an application layer, an application framework layer, a hardware abstraction layer (HAL), a driver layer, and a hardware layer (Hardwork).

[0106] It can be understood that,Figure 11 As one example, the layers divided in the electronic device 100 are not limited to Figure 11 As shown, the layers can also include an Android runtime and libraries layer, etc. between the application framework layer and the HAL layer.

[0107] The application layer can include a series of application packages. As Figure 11 As shown, the application packages can include a camera and other applications including, but not limited to, camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc.

[0108] Generally, the applications are developed using Java language and are completed by calling application programming interfaces (APIs) and programming frameworks provided by the application framework layer.

[0109] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications of the application layer. The application framework layer includes some pre-defined functions.

[0110] For example, the application framework layer can include a camera access interface. The camera access interface can include a camera service and camera management. The camera service can be used to provide an interface for accessing a camera, and the camera management can be used to provide an interface for managing access to the camera.

[0111] In addition, the application framework layer can also include a content provider, a resource manager, a notification manager, a window manager, a view system, a phone manager, etc. Similarly, the camera application can also call the content provider, the resource manager, the notification manager, the window manager, the view system, etc. according to actual business needs, and the embodiments of the present application do not make any limitation thereto.

[0112] The hardware abstraction layer is used to abstract hardware, such as encapsulating the driver programs in the driver layer and providing calling interfaces to the application framework layer, and shielding the implementation details of the low-level hardware.

[0113] For example, the hardware abstraction layer can include a camera hardware abstraction layer (Camera HAL), and of course, other hardware device abstraction layers. The Camera HAL is a camera core software framework, and the Camera HAL includes interface modules, camera abstraction devices, image processing modules, and calibration files, etc. Among them, the interface modules, camera abstraction devices, and image processing modules are components in the image data and control instruction transmission pipeline in the Camera HAL, and of course, different components correspond to different functions. For example, the interface module can be a software interface facing the application framework layer, used for data interaction with the application framework layer, and of course, the interface module can also interact with other modules (such as camera abstraction devices, image processing modules) in the HAL. For another example, the camera abstraction device can be a software interface facing the driver layer, used for data interaction with the driver layer, such as calling the camera device driver in the driver layer. For another example, the image processing module can process the raw image data returned by the camera device.

[0114] The calibration file stores original calibration data of each camera of the electronic device, and the original calibration data is a parameter of the binocular camera pre-calibrated when the electronic device is shipped and stored in the calibration file. The original calibration data can include original camera intrinsic parameters, original camera extrinsic parameters, and original distortion parameters, etc. For example, the calibration file can be a binary (Bin) file.

[0115] The image processing module can include a software algorithm, such as an algorithm for determining depth information, which is used to calculate the depth information of each pixel point according to the images obtained by the two cameras according to the method provided in the embodiments of the present application. Optionally, the depth information can be output in the form of a depth map. In a specific embodiment, the image processing module can include a static rectification module, a dynamic rectification module, and a depth estimation module, etc. The static rectification module is used to rectify the images obtained by the two cameras according to the calibration data in the calibration file. The dynamic rectification module is used to correct the calibration data in the calibration file according to the images obtained by the two cameras, and then further rectify the images obtained by the static rectification based on the corrected calibration data. The depth estimation module is used to estimate the depth of each pixel point in the image based on the images obtained by the static rectification or the images obtained by the dynamic rectification, to obtain the depth information corresponding to each pixel point.

[0116] In addition, the image processing module can also include nodes or algorithms with other image processing capabilities, which can be referred to related technologies, and will not be described here.

[0117] The driver layer is used to provide drivers for different hardware devices. For example, the driver layer can include a camera device driver, and of course, other hardware device drivers.

[0118] In addition, the hardware layer includes hardware modules that can be driven, such as a camera device. For example, the camera device includes a plurality of cameras, such as camera 1, camera 2, …, and camera n, where n is an integer greater than 2. In addition, the camera device can also include a multispectral sensor and the like, which are not limited in the embodiments of the present application.

[0119] In the present application, by calling the hardware abstraction layer interface in the hardware abstraction layer, the connection between the application layer and the application framework layer above the hardware abstraction layer and the driver layer and the hardware layer below can be achieved, and camera data transmission and function control can be achieved.

[0120] The working process of the software and hardware of the electronic device 100 will be described below in conjunction with a capture and photographing scenario.

[0121] The camera application in the application layer can be displayed on the screen of the electronic device 100 in the form of an icon. When the icon of the camera application is clicked by the user to trigger, the electronic device 100 starts to run the camera application. When the camera application runs on the electronic device 100, the camera application calls the interface corresponding to the camera application in the application framework layer, and then starts the camera device driver by calling the hardware abstraction layer, turns on the plurality of cameras on the electronic device 100, and performs photographing through the plurality of cameras to obtain a photographed image. The photographed image obtained is stored in a gallery.

[0122] During the photographing through the plurality of cameras 193, the camera application can call the image processing module in the application framework layer to process the images obtained by the binoculars to determine the depth information corresponding to each pixel point in the images. The determined depth information can be used for subsequent blurring processing or other scenarios, which are not limited in the embodiments of the present application.

[0123] For ease of understanding, the embodiments of the present application will take an electronic device with the structures shown in Figure 10 and Figure 11 as an example, and the method for determining depth information provided by the embodiments of the present application will be described in detail in conjunction with the drawings and application scenarios. For ease of understanding, the embodiments of the present application will take the first camera (also referred to as the first camera) and the second camera (also referred to as the second camera) included in the plurality of cameras configured by the electronic device as an example, and determine the depth information based on the images photographed by the first camera and the second camera. In other words, the first camera and the second camera are the binoculars of a binocular camera. The first camera and the second camera are arranged obliquely, and the translation parameters between the first camera and the second camera are T=(T x, T y). , ). and are not equal to 0.

[0124] Optionally, the camera 1 can be a main camera, and the camera 2 can be an auxiliary camera. The main camera refers to a camera that sends a display to the display screen. The auxiliary camera refers to a camera that is used to assist the main camera in calculating the depth and does not send a display to the display screen. It should be understood that in other embodiments, the camera 1 can also be an auxiliary camera, and the camera 2 can be a main camera, which is not limited in the embodiments of the present application.

[0125] First, the application scenario is described.

[0126] It can be understood that in some scenarios or shooting modes, the camera application needs to perform depth calculation to determine depth information. In this case, the camera application will call the binocular camera to shoot an image and perform depth calculation. For example, when the user selects a portrait mode, the camera application can call the binocular camera to shoot an image and calculate the depth information, and then perform background blurring on the background to improve the display effect of the person in the image. For another example, when the user selects an aperture mode, the camera application can also call the binocular camera to shoot an image, and then calculate the depth information and perform blurring processing on the image based on the depth information to improve the image effect.

[0127] For example, Figure 12 An example of an application scenario for determining depth information is provided in the embodiments of the present application. As shown in Figure 12 (a) of FIG. 1, taking a mobile phone as an example, the mobile phone desktop includes a camera icon 1201. The user clicks the camera icon to open the camera application, as shown in Figure 12 (b) of FIG. 1. Taking the default entry of the portrait mode after the camera application is started as an example, the user can click the shooting control 1202 in the interface shown in Figure 12 (b) of FIG. 1. The camera application responds to the user's operation and sends a calling instruction to the camera 1 and the camera 2 to simultaneously call the camera 1 and the camera 2 for shooting. Then, the depth information is calculated according to the method provided in the embodiments of the present application, and based on the depth information, the background blurring processing is performed to obtain the image 1203 shown in Figure 12 (c) of FIG. 1.

[0128] It should be noted that the above is only an example of the application scenario of the method provided in the embodiments of the present application, and is not limited. The method can be applied to any shooting scenario that needs to determine the depth information, which is not listed here.

[0129] The application scenario shown in Figure 12 is taken as an example to describe the method for determining the depth information provided in the present application.

[0130] Embodiment one:

[0131] The method for determining depth information provided in this embodiment includes: correcting an image acquired by the camera 2 according to original calibration data calibrated when the electronic device is manufactured (referred to as static correction), and performing depth calculation.

[0132] Exemplarily, Figure 13 is a flowchart of a method for determining depth information provided in an embodiment of the present application, Figure 14 is a schematic diagram of a principle of determining depth information provided in an embodiment of the present application, please refer to Figure 13 and Figure 14 , the method comprises:

[0133] S101, the camera application responds to a user inputting a shooting instruction in a portrait mode, and sends a calling instruction to the camera 1 and the camera 2.

[0134] As described above, the method provided in this embodiment can also be applied to other application scenarios, and thus, in this step, the operation of triggering the camera application to send a calling instruction to the camera 1 and the camera 2 can also be other operations. Optionally, the instruction input by the user for triggering the method to execute can be uniformly referred to as a binocular shooting instruction (also referred to as a first instruction). The binocular shooting instruction is used to instruct to shoot an image by the camera 1 and the camera 2 in the binocular camera and determine depth information of each pixel point in the image. The user can input the binocular shooting instruction by clicking a shooting control when the camera application is in a portrait shooting mode.

[0135] Optionally, in different scenarios, the camera application can call the same camera or different cameras. For example, the electronic device includes three cameras a, b and c. In scenario 1, the camera application can call the camera a and the camera b, and in scenario 2, the camera application can call the camera b and the camera c. The method provided in this embodiment can be triggered when it is necessary to call two cameras that are tilted. Whether the two cameras are tilted can be determined according to whether the camera extrinsic parameters of the two cameras are not equal to 0. and . and are not equal to 0, it is considered that the two cameras are tilted. As a possible implementation manner, it can be further limited that when the difference between the camera extrinsic parameters of the two cameras is greater than a preset threshold, the method provided in this application is triggered to be executed.

[0136] S102, the camera 1 acquires an image 1 (also referred to as a first original image) in response to the calling instruction of the camera application.

[0137] The camera 1 starts a shooting function according to the calling instruction of the camera application, and acquires the image 1 by the shooting function.

[0138] ​The shooting function can be either a photo capture function or a video recording function. That is, image 1 can be a single frame image captured by camera 1 using the photo capture function, or it can be an image frame from a video frame recorded by camera 1 using the video recording function.

[0139] S103, Camera 1 sends image 1 to the static correction module in the image processing module.

[0140] S104, Camera 2 responds to the call command from the camera application and acquires Image 2 (also known as the second raw image).

[0141] Camera 2 initiates the shooting function according to the call command of the camera application, and captures image 2 through the shooting function.

[0142] The shooting function can be either a photo capture function or a video recording function. That is, image 2 can be a single frame image captured by camera 2 using the photo capture function, or it can be an image frame from a video frame recorded by camera 2 using the video recording function.

[0143] S105, Camera 2 sends Image 2 to the static correction module in the image processing module.

[0144] S106. The static correction module obtains the original calibration data of camera 1 and camera 2 from the calibration file.

[0145] In this embodiment of the application, the original calibration data may include the original camera intrinsic parameters of camera 1 (also known as the first camera intrinsic parameters), the original camera intrinsic parameters of camera 2 (also known as the second camera intrinsic parameters), the original camera extrinsic parameters, the original distortion parameters of camera 1 (also known as the first distortion parameters), and the original distortion parameters of camera 2 (also known as the second distortion parameters), etc.

[0146] In this embodiment, the original camera intrinsic parameters of camera 1 are represented as follows: . This can include the focal length of camera 1 (denoted as...). ) and optical center coordinates The value in its pixel coordinate system. Focal length The focal length value, converted to pixel coordinates in the pixel coordinate system of camera 1, can be denoted as: and . Indicates focal length The value along the x-direction in the pixel coordinate system of camera 1 (i.e., the value of the focal length of the first camera along the horizontal direction). Indicates focal length The value along the y-direction in the pixel coordinate system of camera 1 (i.e., the value of the focal length of the first camera along the vertical direction). Optical center coordinates. Transform to pixel coordinates, and denote the value along the x-direction as: The value along the y-direction is denoted as .

[0147] Optional, raw camera internals It can be represented in matrix form, as shown in the following formula (3):

[0148] (3)

[0149] The raw camera intrinsics of camera 2 are represented as follows . and Similarly, it can be expressed as the following formula (4):

[0150] (4)

[0151] Indicates the focal length of camera 2 The value along the x-direction in its pixel coordinate system. Indicates focal length The value along the y-direction in the pixel coordinate system of camera 2 is 6. Indicates the optical center coordinates of camera 2 Transform to the x-axis value in its pixel coordinate system. Indicates the optical center coordinates of camera 2 Transform to the value along the y-direction in its pixel coordinate system.

[0152] The original camera extrinsic parameters may include the original rotation matrix (also known as the first rotation matrix, denoted as R) and the original translation parameters (also known as the first translation parameters, denoted as T). Optionally, the original rotation matrix R represents adjusting camera 2 from its actual position to an angle of arctan(θ) with the optical axis of camera 1. The required translation amount. The original translation parameter T represents the amount of translation required to adjust camera 2 from its actual position to the position of camera 1. Original translation parameter T = ( , ),in, The x-direction translation (also known as the first horizontal translation) is equal to the coordinate difference between camera 1 and camera 2 in the x-axis direction. The vertical translation (also known as the first vertical translation) is equal to the difference in coordinates between camera 1 and camera 2 along the y-axis. For details, please refer to [link to relevant documentation]. Figure 8 .

[0153] S107. The static correction module uses the original distortion parameters of camera 1 to correct the distortion of image 1, and obtains the distortion-corrected image 1 (also known as the first image).

[0154] S108, the static correction module corrects the image 2 by using the original distortion parameters of the camera 2 to obtain a distortion-corrected image 2 (also referred to as a second image).

[0155] The distortion correction is performed on the image 1 and the image 2 respectively, the deviation introduced by the manufacturing precision or the assembly process of the camera is eliminated, the accuracy of the image is improved, and the image display effect is improved.

[0156] S109, the static correction module calculates a static correction matrix H (also referred to as a first correction matrix) according to the original calibration data of the camera 1 and the camera 2.

[0157] The correction matrix is also referred to as a distortion matrix. In the embodiment of the application, the correction matrix determined based on the original calibration data in the calibration file is referred to as a static correction matrix, and is denoted as H.

[0158] Optionally, the static correction matrix H can be calculated by formula (5):

[0159] (5)

[0160] wherein, represents the original camera intrinsic parameter of the camera 1. R represents an original rotation matrix. represents the original camera intrinsic parameter of the camera 2.

[0161] S110, the static correction module corrects the distortion-corrected image 2 according to the static correction matrix H and the original translation parameter T to obtain a static-corrected image 2 (also referred to as a third image).

[0162] It should be noted that the static correction module keeps the distortion-corrected image 1 unchanged, and only corrects the distortion-corrected image 2.

[0163] The static correction matrix H contains the original camera intrinsic parameter of the camera 1 and the original camera intrinsic parameter of the camera 2 , and thus the static correction matrix H can correct the error caused by the internal structure of the camera 1 and the camera 2.

[0164] As described above, the original rotation matrix R represents the rotation amount required for adjusting the camera 2 from the actual position to the position in which the angle between the optical axis of the camera 1 and the optical axis of the camera 2 is . The original translation parameter T represents the translation amount required for adjusting the camera 2 from the actual position to the position of the camera 1. Then, the static-corrected image 2 obtained by the correction of the static correction matrix H containing R and the original translation parameter T while keeping the distortion-corrected image 1 unchanged is equivalent to correcting the angle between the optical axes of the two cameras to , and two imaging planes are aligned in x direction and y direction, which is also called target binocular structure. That is, in the binocular stereo vision geometry structure shown in Figure 2 , the right imaging plane is translated and slightly adjusted in angle. In this way, the image rotation angle is adjusted to be very small, and the distortion degree of image 2 is very small.

[0165] S111, the static correction module determines a search direction k1 (also referred to as a first search direction), a search step step1 (also referred to as a first search step), and a baseline distance B1 (also referred to as a first baseline distance) according to the original calibration data of camera 1 and camera 2.

[0166] The search direction k1 is a direction along a straight line with a slope of .The x direction translation amount is the coordinate difference between camera 1 and camera 2 in the x direction, , and the y direction translation amount is the coordinate difference between camera 1 and camera 2 in the y direction.

[0167] Optionally, the search step step1 can be . Optionally, the search step step1 can also be . Optionally, the search step step1 can also be .

[0168] The baseline distance B1 is .

[0169] S112, the static correction module sends the search direction k1, the search step step1, and the baseline distance B1 to the depth estimation module.

[0170] S113, the depth estimation module performs feature matching on the distortion-corrected image 1 and the static-corrected image 2 according to the search direction k1 and the search step step1, and obtains a plurality of groups of feature pairs.

[0171] Please continue to see Figure 14 , optionally, the depth estimation module can include a deep learning feature extraction model, a feature matching model, a correlation calculation model, a disparity calculation model, etc.

[0172] The depth estimation module can input the distortion-corrected image 1 and the static-corrected image 2 into the deep learning feature extraction model for feature extraction, respectively, to obtain features. The features extracted from the distortion-corrected image 1 are referred to as feature F-1 (also referred to as first feature), and the features extracted from the static-corrected image 2 are referred to as feature F-2 (also referred to as third feature).

[0173] Then, the depth estimation module inputs the features F-1, the features F-2, the search direction k1 and the search step step1 into a feature matching model to match the features F-2 corresponding to the features F-1. Specifically, for each feature F-1, the feature matching model searches for the corresponding feature F-2 in the static rectified image 2 along the search direction k1 according to the search step step1. Each feature F-1 and the corresponding feature F-2 are called a feature pair, and the feature matching model outputs all the feature pairs.

[0174] Specifically, please refer to Figure 15 . For example, a feature can be represented by a feature block. As shown in (a) of FIG. 1, Figure 15 , for any feature block 1501 in the rectified image 1, the feature matching model can search for the corresponding feature block in the rectified image 2 along the search direction k1, the search direction k1 is along the direction of a straight line with a slope of , and the search step step1 is , as shown in (b) of FIG. 1. Figure 15

[0175] S114, the depth estimation module calculates the correlation of each feature pair (also referred to as correlation information) according to the multiple feature pairs and the original calibration data.

[0176] Please continue to refer to Figure 14 , the depth estimation module inputs the features, the original camera intrinsic parameter of the camera 1, the original camera extrinsic parameter of the camera 2 and the original camera extrinsic parameter into a correlation calculation model. The correlation calculation model calculates the correlation between the feature F-1 and the feature F-2 in each feature pair based on the original camera intrinsic parameter of the camera 1, the original camera extrinsic parameter of the camera 2 and the original camera extrinsic parameter.

[0177] S115, the depth estimation module calculates the parallax d corresponding to each pixel point according to the correlation of each feature pair.

[0178] Please continue to refer to Figure 14 , the depth estimation module inputs the correlation of each feature pair into a parallax calculation model. The parallax calculation model calculates the parallax d corresponding to each pixel point. The parallax d corresponding to all the pixel points can be represented in the form of a parallax map, that is, the output of the parallax calculation model can be a parallax map.

[0179] S116, the depth estimation module calculates the depth information Z corresponding to each pixel point according to the parallax d corresponding to each pixel point and the baseline distance B1.

[0180] For example, Figure 16 ​A schematic diagram of a principle of calculating depth information after polar rectification is provided in an embodiment of the present application. It can be understood that after polar rectification, the focal lengths of the two cameras are the same, and are uniformly denoted as f.

[0181] Specifically, for any point p1 in the rectified image 1 (i.e., any point P of the photographed object), the depth information thereof can be calculated according to the above formula (6):

[0182] (6)

[0183] After the depth information of each pixel point is calculated, the camera application can perform background blurring processing on the rectified image 1 based on the depth information, to obtain a final image, and send the final image to a display for display, so as to display an interface as shown in (c) of FIG. 1. Figure 12 The background blurring processing and the like will not be elaborated here.

[0184] In the embodiment, the rectified image 1 remains unchanged, and only the rectified image 2 is rectified, without distorting the image 1, increasing the resolution of the image 1, or increasing the invalid area, reducing the calculation amount and calculation time in subsequent feature matching, damaging the features of the pixel points in the image 1, improving the accuracy of subsequent feature matching, and further improving the accuracy of depth information calculation. Moreover, in the embodiment, the standard of polar rectification, or the target of polar rectification, is to make the angle between the camera 1 and the camera 2 be the angle determined in the original calibration data, and to basically maintain the original angle, so that the adjustment angle of the entire process to the image 2 is very small, the distortion angle is very small, the increase in the resolution of the image 2 is also very small, the increase in the invalid area is very small, the calculation amount and calculation time in subsequent feature matching are further reduced, the features of the pixel points in the image 2 are not easily damaged, the accuracy of subsequent feature matching is improved, and the accuracy of depth information calculation is further improved. An exemplary

[0185] A schematic diagram of a comparison between images before and after polar rectification is provided in an embodiment of the present application. The original image 1 obtained by the camera 1 is as shown in (a1) of FIG. 1, and the original image 2 obtained by the camera 2 is as shown in (a2) of FIG. 1. The rectified image 1 obtained after rectification by the method provided in the embodiment is as shown in (b1) of FIG. 1, and the rectified image 2 obtained after rectification by the method provided in the embodiment is as shown in (b2) of FIG. 1. By comparison, it can be obviously seen that the method of the present application does not distort the image 1, and the distortion angle of the image 2 is also very small. Figure 17 Figure 17 Figure 17 Figure 17 Figure 17 Figure 9 Figure 17 ​​​​​​The black filled part in the (b2) figure in the (b1) figure is very small.

[0186] In addition, referring to Figure 18 The method provided by the embodiment of the present application searches in the direction, which is one-dimensional search and does not increase the calculation amount.

[0187] Embodiment two:

[0188] It can be understood that in the process of using the electronic device, the position of one or both cameras in the binocular camera may change due to collision, impact, etc., resulting in a change in the relative position of the two cameras. In this case, if the correction is continued based on the camera extrinsic parameters in the original calibration data, the corrected image obtained is inaccurate, and the depth information calculation is inaccurate, and the image effect obtained by shooting is poor.

[0189] In view of this, in the method for determining depth information provided by the embodiment, after obtaining the static corrected image 2, the original camera extrinsic parameters can be corrected, and then the epipolar rectification is performed based on the corrected camera extrinsic parameter data. After completion, feature matching and depth calculation are performed. During feature matching, the search direction, search step, baseline distance, etc. are also corrected. This method improves the accuracy of epipolar rectification, and further improves the accuracy of feature matching and depth information calculation, and improves the image shooting effect. The above process is called dynamic correction (or dynamic correction).

[0190] In addition, referring to Figure 19 The application scenarios and the process of triggering depth calculation of the embodiment can be the same as those of embodiment one, and in the embodiment, static correction can be performed first, and then dynamic correction is performed, and depth calculation is performed after dynamic correction.

[0191] That is, steps S101 to S110 in the above embodiment one are first performed. For details, see embodiment one, which will not be described here.

[0192] Exemplarily, Figure 20 is a flowchart of another example of the method for determining depth information provided by the embodiment of the present application, and after step S110, the steps in Figure 20 may be performed, including:

[0193] S201, the static correction module sends the distortion corrected image 1, the static corrected image 2 and the original calibration data to the dynamic correction module.

[0194] Optionally, the static correction module can also not send the original calibration data to the dynamic correction module, but obtain the original calibration data from the calibration file by the dynamic correction module, and the embodiment of the present application does not make any limitation thereto.

[0195] S202, the dynamic correction module performs feature point extraction on the image 1 after distortion correction and the image 2 after static correction respectively, to obtain feature points of the image 1 (also referred to as feature points of the first image) and feature points of the image 2 (also referred to as feature points of the third image).

[0196] S203, the dynamic correction module performs feature point matching on the feature points of the image 1 and the feature points of the image 2.

[0197] S204, the dynamic correction module calculates an essential matrix E according to the feature point matching result and the original camera extrinsic parameter.

[0198] It can be understood that the essential matrix E constrains the relationship of a three-dimensional point P in the world coordinate system in the camera coordinate system of the camera 1 and the camera coordinate system of the camera 2. The projection of the P point in the imaging plane of the camera 1 is represented as p1, and the projection of the P point in the imaging plane of the camera 2 is represented as p2. Then, according to the original translation parameter T and the original rotation matrix R, the essential matrix E can be calculated based on formula (7) and formula (8):

[0199] (7)

[0200] (8)

[0201] S205, the dynamic correction module performs singular value decomposition (SVD) on the essential matrix E to solve a correction rotation matrix and a preliminary translation parameter (also referred to as a third translation parameter).

[0202] The essential matrix E inherits the properties of the matrix T^, so its singular values are , so the rotation matrix R and the translation parameter T can be obtained by SVD decomposition. In this embodiment, the rotation matrix obtained by SVD decomposition is referred to as a correction rotation matrix, denoted as . The translation parameter obtained by SVD decomposition is referred to as a preliminary translation parameter, denoted as . . Wherein, is also referred to as a third horizontal translation, is also referred to as a third vertical translation.

[0203] If the SVD decomposition of the essential matrix E is , then and The results are as follows:

[0204] or ;

[0205] or ;

[0206] wherein, is an orthogonal matrix, . is an anti-symmetric matrix, . If the vector , then .

[0207] It can be understood that the corrected translation parameter is scale-free.

[0208] In the above steps, based on the distortion-corrected image 1 and the static-corrected image 2, the essential matrix is calculated, and SVD decomposition is performed, so that the accurate rotation matrix and translation parameter can be determined, and the original rotation matrix and original translation parameter can be accurately corrected.

[0209] S206, the dynamic correction module performs scale correction on the preliminary translation parameter based on the original translation parameter T, to obtain the corrected translation parameter (also referred to as the second translation parameter).

[0210] Optionally, the correction coefficient s can be determined according to the original translation parameter T and the preliminary translation parameter . Optionally, the correction coefficient , or, .

[0211] Then, the preliminary translation parameter is corrected based on the correction coefficient s, to obtain the corrected translation parameter . =s . Specifically, , . Wherein, is also referred to as the second horizontal translation amount, is also referred to as the second vertical translation amount.

[0212] The processes of the above steps S201 to S206 are also referred to as correcting the calibration data.

[0213] In this step, through scale correction, the accuracy of the translation parameter can be improved, and then the accuracy of subsequent epipolar line correction can be improved.

[0214] S207, the dynamic correction module calculates a dynamic correction matrix (also referred to as the second correction matrix) according to the original camera intrinsic parameter of the camera 1, the original camera intrinsic parameter of the camera 2, and the corrected rotation matrix (also referred to as the second rotation matrix).

[0215] In the embodiments of the present application, the correction matrix determined based on the calibrated data after correction is referred to as a dynamic correction matrix, denoted as .

[0216] Optionally, the dynamic correction matrix may be calculated by formula (9):

[0217] (9)

[0218] This step has the same principle as the above step S109, and the difference is that the rotation matrix used is different, which will not be described here.

[0219] S208, the dynamic correction module corrects the static corrected image 2 according to the dynamic correction matrix and the correction translation parameter to obtain a dynamically corrected image 2 (also referred to as a fourth image).

[0220] It should be noted that the dynamic correction module keeps the distortion-corrected image 1 unchanged and only corrects the static corrected image 2.

[0221] This step has the same principle as the above step S110, and the difference is that the correction matrix used is different, which will not be described here.

[0222] Optionally, if dynamic correction is further performed after static correction, in the above step S110, the translation correction can not be performed by the original translation parameter T, but in this step, the translation correction is performed by the correction translation parameter. In this way, the static correction process can be simplified and the algorithm running efficiency can be improved.

[0223] S209, the dynamic correction module determines the correction search direction k2, the correction search step step2 and the correction baseline distance B2 according to the original camera intrinsic parameter of the camera 1, the original camera intrinsic parameter of the camera 2, the correction rotation matrix and the correction translation parameter .

[0224] The correction search direction k2 is the direction along the straight line with a slope of . The x direction translation amount in the correction translation parameter is denoted as , and the y direction translation amount in the correction translation parameter is denoted as .

[0225] Optionally, the correction search step step2 is . Optionally, the correction search step step2 can also be . Optionally, the correction search step step2 can also be .

[0226] Corrected baseline distance B2 .

[0227] The step is the same as the principle of step S111 described above, the difference is that the rotation matrix, translation parameters and other parameters used are different, so the results obtained are different, which will not be repeated here.

[0228] S210, the dynamic correction module sends the corrected search direction k2, the corrected search step step2 and the corrected baseline distance B2 to the depth estimation module.

[0229] S211, the depth estimation module performs feature matching on the distortion-corrected image 1 and the dynamic-corrected image 2 according to the corrected search direction k2 and the corrected search step step2, and obtains a plurality of feature pairs.

[0230] Among them, the features extracted from the dynamic-corrected image 2 can be represented as feature F-3, also known as the fourth feature.

[0231] S212, the depth estimation module calculates the correlation of each group of feature pairs according to the plurality of feature pairs, the original camera intrinsic parameters of camera 1, the original camera intrinsic parameters of camera 2, the corrected rotation matrix and the corrected translation parameters .

[0232] Among them, the corrected rotation matrix and the corrected translation parameters are also collectively referred to as the corrected camera extrinsic parameters.

[0233] S213, the depth estimation module calculates the disparity d corresponding to each pixel point according to the correlation of each group of feature pairs.

[0234] S214, the depth estimation module calculates the depth information Z corresponding to each pixel point according to the disparity d corresponding to each pixel point and the corrected baseline distance B2.

[0235] The above steps S210 to S214 are the same as the principles of steps S112 to S116 in Embodiment One, the difference is that different parameters are used, which will not be repeated here.

[0236] It can be understood that the principles of polar rectification and depth calculation of Embodiment Two and Embodiment One are the same, so Embodiment Two has all the beneficial effects of Embodiment One, which will not be repeated. In addition, as described above, Embodiment Two can correct the original camera extrinsic parameters, improve the accuracy of polar rectification, and thus improve the accuracy of feature matching and depth information calculation, improve the effect of the final obtained image, and improve the user experience.

[0237] The above describes in detail the method for determining the depth information provided by the embodiments of the present application. It can be understood that the electronic device includes hardware and / or software modules corresponding to the functions to implement the above functions. Those skilled in the art should easily realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented in hardware or a combination of hardware and computer software. Whether a certain function is implemented in hardware or computer software driven hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered beyond the scope of the present application.

[0238] The embodiments of the present application can divide the functional modules of the electronic device according to the above method examples. For example, each functional module such as a detection unit, a processing unit, and a display unit can be divided according to each function, or two or more functions can be integrated in one module. The integrated module can be implemented in the form of hardware or software functional module. It should be noted that the division of the modules in the embodiments of the present application is illustrative and is only a logical function division. Actual implementation can have another division method.

[0239] It should be noted that all related contents of each step involved in the above method embodiments can be cited to the function description of the corresponding functional module, which will not be repeated here.

[0240] The electronic device provided by the embodiments of the present application is used to execute the above method for determining the depth information, and thus can achieve the same effect as the above implementation method.

[0241] In the case of using integrated units, the electronic device can further include a processing module, a storage module, and a communication module. The processing module can be used to control and manage the actions of the electronic device. The storage module can be used to support the electronic device to execute program codes and data, etc. The communication module can be used to support the communication between the electronic device and other devices.

[0242] The processing module can be a processor or a controller. It can implement or execute various exemplary logical blocks, modules, and circuits described in combination with the disclosure of the present application. The processor can also be a combination of computing functions, such as one or more microprocessor combinations, combinations of digital signal processing (DSP) and microprocessors, etc. The storage module can be a memory. The communication module can be a device for interacting with other electronic devices, such as a radio frequency circuit, a Bluetooth chip, a Wi-Fi chip, etc.

[0243] In one embodiment, when the processing module is a processor and the storage module is a memory, the electronic device related to the embodiment can be a device with the structure as shown in the figure. Figure 10

[0244] The embodiment of the present application further provides a computer readable storage medium, and the computer readable storage medium stores a computer program. When the computer program is executed by a processor, the processor executes the method for determining depth information in any of the above embodiments.

[0245] The embodiment of the present application further provides a computer program product. When the computer program product is run on a computer, the computer executes the related steps to implement the method for determining depth information in the above embodiments.

[0246] In addition, the embodiment of the present application further provides an apparatus, which can be a chip, a component or a module, and the apparatus can include a processor and a memory connected to each other. The memory is used to store computer execution instructions. When the apparatus is running, the processor can execute the computer execution instructions stored in the memory, so that the chip executes the method for determining depth information in the above method embodiments.

[0247] The electronic device, the computer readable storage medium, the computer program product or the chip provided by the embodiment can achieve the beneficial effects of the corresponding method provided above, and thus the beneficial effects of the electronic device, the computer readable storage medium, the computer program product or the chip are not described here.

[0248] Through the description of the above embodiments, those skilled in the art can understand that, for the convenience and brevity of description, only the division of the above functional modules is taken as an example for illustration. In actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the apparatus is divided into different functional modules to complete all or part of the functions described above.

[0249] In the several embodiments provided by the present application, it should be understood that the disclosed apparatus and method can be implemented in other ways. For example, the apparatus embodiments described above are only schematic. The division of the modules or units is only a logical function division, and there can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another apparatus, or some features can be ignored or not executed. In addition, the displayed or discussed mutual coupling or direct coupling or communication connection between the units can be indirect coupling or communication connection through some interfaces, apparatuses or units, and can be electrical, mechanical or other forms.

[0250] ​The units described as separate components may or may not be physically separate, and the components displayed as units may be a physical unit or multiple physical units, that is, may be located in one place, or also can be distributed to multiple different places. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0251] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present alone, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0252] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a readable storage medium. Based on such understanding, the technical scheme of the embodiments of the present application essentially or the part that contributes to the prior art or the whole or part of the technical scheme can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions to make a device (which can be a single-chip microcomputer, a chip, etc.) or a processor execute all or part of the steps of the method of each embodiment of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.

[0253] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method of determining depth information, the method being performed by an electronic device including a first camera and a second camera, the method comprising: The method comprises: acquiring a first image collected by the first camera and a second image collected by the second camera; acquiring camera intrinsic parameters of the first camera and the second camera, and a first rotation matrix and a first translation parameter of the second camera relative to the first camera; the first translation parameter comprises a first horizontal translation and a first vertical translation; correcting the second image according to the camera intrinsic parameters and the first rotation matrix to obtain a third image; the third image is consistent with imaging of the second camera under a target binocular structure, and an included angle between optical axes of the first camera and the second camera in the target binocular structure is determined according to a ratio of the first horizontal translation to the first vertical translation; performing a feature matching operation based on the first image and the third image, and calculating depth information of each pixel point in the first image; a search direction when the feature matching operation is performed is determined according to the first horizontal translation and the first vertical translation.

2. The method of claim 1, wherein, The included angle between the optical axes of the first camera and the second camera in the target binocular structure is equal to an inverse tangent value of the ratio of the first horizontal translation to the first vertical translation.

3. The method according to claim 1 or 2, characterized in that, The camera intrinsic parameters comprise first camera intrinsic parameters of the first camera and second camera intrinsic parameters of the second camera; and the correcting the second image according to the camera intrinsic parameters and the first rotation matrix to obtain a third image comprises: determining a first correction matrix according to a product of the first camera intrinsic parameters, the first rotation matrix and the second camera intrinsic parameters; correcting the second image based on the first correction matrix and the first translation parameter to obtain the third image.

4. The method of claim 3, wherein, The performing a feature matching operation based on the first image and the third image comprises: performing feature extraction on the first image to obtain a plurality of first features; performing feature extraction on the third image to obtain a plurality of third features; for a second feature, searching in a first search direction in the third image according to a first search step to determine a fifth feature matching the second feature in the plurality of third features; the second feature is any one of the plurality of first features, and the first search direction is along a straight line with a slope equal to the ratio of the first horizontal translation to the first vertical translation; determining the second feature and the fifth feature as a feature pair.

5. The method of claim 4, wherein, The first search step is , or , or ; wherein, denotes a value of the focal length of the first camera along the horizontal direction in the first camera intrinsic parameter, denotes a value of the focal length of the first camera along the vertical direction in the first camera intrinsic parameter, denotes the first horizontal translation amount, denotes the first vertical translation amount.

6. The method of claim 5, wherein, The calculating depth information of each pixel point in the first image comprises: calculating correlation information of each feature pair according to the camera intrinsic parameters, the first rotation matrix and the first translation parameter; determining a parallax corresponding to each pixel point in the first image according to the correlation information of each feature pair; determining depth information corresponding to each pixel point in the first image according to the parallax corresponding to each pixel point based on a first baseline distance; the first baseline distance is determined according to the first camera intrinsic parameters, the second camera intrinsic parameters, the first horizontal translation and the first vertical translation.

7. The method of claim 6, wherein, The first baseline distance is determined according to the following formula: B1= ; The first baseline distance is represented by B1.

8. The method of claim 1 or 2, wherein, The feature matching operation is performed based on the first image and the third image, and depth information of each pixel point in the first image is calculated, including: The first rotation matrix and the first translation parameter are respectively corrected according to the first image and the third image, to obtain a second rotation matrix and a second translation parameter; The third image is corrected based on the camera intrinsic parameter, the second rotation matrix and the second translation parameter, to obtain a fourth image; The feature matching operation is performed based on the first image and the fourth image, and depth information of each pixel point in the first image is calculated.

9. The method of claim 8, wherein, The first rotation matrix and the first translation parameter are respectively corrected according to the first image and the third image, to obtain a second rotation matrix and a second translation parameter, including: An essential matrix is calculated according to the first image, the third image, the first rotation matrix and the first translation parameter; The second rotation matrix and a third translation parameter are solved by singular value decomposition of the essential matrix; The third translation parameter is scaled and corrected according to the first translation parameter, to obtain the second translation parameter.

10. The method of claim 9, wherein, The essential matrix is calculated according to the first image, the third image, the first rotation matrix and the first translation parameter, including: Feature points of the first image are extracted to obtain feature points of the first image; Feature points of the third image are extracted to obtain feature points of the third image; Feature point matching is performed on the feature points of the first image and the feature points of the third image to obtain a feature point matching result; The essential matrix is calculated according to the feature point matching result, the first rotation matrix and the first translation parameter.

11. The method according to claim 9 or 10, characterized in that, The second translation parameter includes a second horizontal translation and a second vertical translation, and the third translation parameter includes a third horizontal translation and a third vertical translation; the third translation parameter is scaled and corrected according to the first translation parameter, to obtain the second translation parameter, including: A correction coefficient is calculated by calculating a ratio of the first horizontal translation to the third horizontal translation, or calculating a ratio of the first vertical translation to the third vertical translation; The second horizontal translation is calculated by multiplying the correction coefficient by the third horizontal translation; The second vertical translation is calculated by multiplying the correction coefficient by the third vertical translation.

12. The method of claim 8, wherein, The camera intrinsic parameter includes a first camera intrinsic parameter of the first camera and a second camera intrinsic parameter of the second camera; the third image is corrected based on the camera intrinsic parameter, the second rotation matrix and the second translation parameter, to obtain a fourth image, including: A second correction matrix is determined according to a product of the first camera intrinsic parameter, the second rotation matrix and the second camera intrinsic parameter; The third image is corrected based on the second correction matrix and the second translation parameter, to obtain the fourth image.

13. The method of claim 8, wherein, The second translation parameter includes a second horizontal translation amount and a second vertical translation amount, and the feature matching operation is performed based on the first image and the fourth image, including: performing feature extraction on the first image to obtain a plurality of first features; performing feature extraction on the fourth image to obtain a plurality of fourth features; for a second feature, searching in the fourth image along a second search direction according to a second search step to determine a sixth feature matched with the second feature from the plurality of fourth features; the second feature is any one of the plurality of first features, and the second search direction is along a straight line with a slope equal to a ratio of the second horizontal translation amount to the second horizontal translation amount; determining the second feature and the sixth feature as a feature pair.

14. The method of claim 13, wherein, The camera intrinsic parameters include first camera intrinsic parameters of the first camera and second camera intrinsic parameters of the second camera; the second search step size is , or , or ; wherein, denotes a value of the focal length of the first camera along the horizontal direction in the first camera intrinsic parameter, denotes a value of the focal length of the first camera along the vertical direction in the first camera intrinsic parameter, denotes the second horizontal translation amount, denotes the second vertical translation amount.

15. The method of claim 14, wherein, The depth information of each pixel point in the first image is calculated, including: calculating the correlation information of each feature pair according to the camera intrinsic parameter, the second rotation matrix and the second translation parameter; determining the parallax corresponding to each pixel point in the first image according to the correlation information of each feature pair; determining the depth information corresponding to each pixel point in the first image based on the second baseline distance and the parallax corresponding to each pixel point; the second baseline distance is determined according to the first camera intrinsic parameter, the second camera intrinsic parameter, the second horizontal translation amount and the second vertical translation amount.

16. The method of claim 15, wherein, The second baseline distance is determined according to the following formula: B2= ; wherein B2 represents the second baseline distance.

17. The method of claim 8, wherein, The camera intrinsic parameter includes a first camera intrinsic parameter of the first camera and a second camera intrinsic parameter of the second camera; and the second image is rectified to obtain a third image according to the camera intrinsic parameter and the first rotation matrix, including: determining a first rectification matrix according to a product of the first camera intrinsic parameter, the first rotation matrix and the second camera intrinsic parameter; rectifying the second image based on the first rectification matrix to obtain the third image.

18. The method of claim 1, wherein, The first image captured by the first camera and the second image captured by the second camera are obtained, including: capturing a first original image by the first camera; capturing a second original image by the second camera; obtaining a first distortion parameter of the first camera and a second distortion parameter of the second camera; rectifying the first original image according to the first distortion parameter to obtain the first image; rectifying the second original image according to the second distortion parameter to obtain the second image.

19. An electronic device, comprising: including: a processor, a memory and an interface; the processor, the memory and the interface cooperate with each other, so that the electronic device executes the method in any one of claims 1 to 18.

20. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, when the computer program is executed by the processor, the processor calls instructions, so that the electronic device executes the method in any one of claims 1 to 18. The computer readable storage medium stores a computer program, when the computer program is executed by the processor, the processor calls instructions, so that the electronic device executes the method in any one of claims 1 to 18.

Citation Information

Patent Citations

  • Medical equipment machine vision image processing method and computer readable storage medium

    CN114331855A

  • Method and device for determining target position information and storage medium

    CN116723264A