Methods for image processing, electronic devices, and computer-readable storage media

By performing specific correction and rotation of the binocular image to align it in the second tilt direction, the problems of low computing efficiency and high power consumption caused by the invalid area after the binocular image are aligned are solved, and more efficient parallax information calculation and lower device power consumption are achieved.

CN118446903BActive Publication Date: 2025-05-16HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311869841.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-05-16
Estimated Expiration
2043-12-29

AI Technical Summary

Technical Problem

A large number of invalid areas are introduced after row alignment or column alignment of binocular images, resulting in low computational efficiency of parallax information and high device power consumption.

Method used

By performing specific corrections and rotations on the basis of arranging the first image and the second image in the first tilt direction, the third image and the fourth image are aligned in the second tilt direction, thereby reducing the invalid area area.

Benefits of technology

It reduces the complexity of parallax information calculation, improves computing efficiency, reduces device power consumption, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118446903B_ABST
    Figure CN118446903B_ABST
Patent Text Reader

Abstract

The present invention provides a method for image processing, an electronic device and a computer-readable storage medium. The method is used in an electronic device, and the electronic device includes a first camera and a second camera arranged in a first oblique direction. The method includes: acquiring a first image and a second image based on the first camera and the second camera; correcting the first image and the second image to obtain a third image and a fourth image; aligning the third image and the fourth image in the second oblique direction; the difference between the extension angles of the second oblique direction and the first oblique direction is a+kπ / 2, and |a| is less than the oblique angle of the first oblique direction; matching the fourth image with the third image along the second oblique direction, determining the disparity information between the feature points of the third image and the corresponding points of the fourth image, and determining the depth information of the first image based on the disparity information. The method for image processing can reduce the invalid area after binocular image correction and improve the calculation efficiency of disparity information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of image processing, and in particular, relates to a method, electronic device and computer-readable storage medium for image processing. Background Art

[0002] In certain usage scenarios, electronic devices such as mobile phones involve the application of depth information. For example, when using the "large aperture mode" or "portrait mode" of a mobile phone camera to take pictures, it is necessary to use a blur algorithm to blur the image to be blurred to obtain a blurred image, so that the blurred image is close to the natural blur effect of a SLR (i.e., the subject is clear and the background is blurred in layers). The blur effect depends on the depth information of the image to be blurred. In order to obtain the depth information of the image to be blurred, usually the main and auxiliary images (i.e., binocular images) collected by dual cameras are first used, and then the binocular stereo matching technology is used to calculate the disparity information between the main and auxiliary images collected by the dual cameras, and then the depth information of the main image is obtained through the disparity information.

[0003] Binocular stereo matching technology involves correcting the main image and the secondary image to be row-aligned or column-aligned to reduce the search range of binocular stereo matching. However, after the main image and the secondary image are aligned in rows or columns, a large number of invalid areas may be introduced. It should be noted that when calculating disparity information, the computational complexity is proportional to the image size. Based on this, when a large number of invalid areas are introduced after the correction of the main image and the secondary image, the complexity of calculating the disparity information is increased, resulting in low efficiency in calculating the disparity information and increased device power consumption, greatly reducing the user experience. Summary of the invention

[0004] The present application provides a method for image processing, an electronic device and a computer-readable storage medium, which can solve the technical problems of low calculation efficiency of disparity information and high power consumption of equipment caused by the introduction of a large number of invalid areas after the current binocular image row alignment or column alignment correction.

[0005] In a first aspect, an embodiment of the present application provides a method for image processing, which is applied to an electronic device. The electronic device includes at least a first camera and a second camera, and the first camera and the second camera are arranged in a first oblique direction. The method includes: acquiring a first image and a second image; wherein the first image and the second image are each obtained by shooting the same shooting target by the first camera and the second camera, and the first camera and the second camera are arranged in the first oblique direction; correcting the first image and the second image to obtain a third image corresponding to the first image and a fourth image corresponding to the second image; wherein the third image and the fourth image are aligned in the second oblique direction; the difference between the extension angle of the second oblique direction and the extension angle of the first oblique direction is a (a can be positive or negative, depending on the increase or decrease of the extension angle) + kπ / 2, the absolute value of a is less than the inclination angle of the first oblique direction, and k is an integer; matching the fourth image with the third image along the second oblique direction, determining the disparity information between the feature points of the third image and the corresponding points of the fourth image, and determining the depth information of the first image according to the disparity information.

[0006] In the method for image processing, according to the epipolar constraint principle, when the third image and the fourth image are aligned in the second oblique direction, the corresponding point of the fourth image appears in the second oblique direction of the fourth image. Therefore, when the fourth image and the third image are matched in the second oblique direction of the fourth image, the feature point of the third image can be matched to the corresponding point. The disparity information is the position difference between the feature point and the corresponding point. Therefore, when the feature point of the third image is matched to the corresponding point, the disparity information between the feature point of the third image and the corresponding point of the fourth image can be obtained, so that the depth information of the first image can be further obtained based on the disparity information.

[0007] In the method for image processing, the difference between the extension angle of the second oblique direction and the extension angle of the first oblique direction is a+kπ / 2, and |a| is smaller than the oblique angle of the first oblique direction, so that the invalid area of ​​the third image and the fourth image can be smaller than the invalid area when the rows are aligned or the columns are aligned, thereby obtaining the benefit of the invalid area. The following is a specific analysis and explanation.

[0008] For ease of explanation, the third tilt direction is introduced as the second tilt direction when k=0, that is, the difference between the extension angle of the third tilt direction and the extension angle of the first tilt direction is a. In this case, the extension angle of the second tilt direction and the extension angle of the third tilt direction are in a relationship of kπ / 2.

[0009] |a| can be understood as the angle between the third tilt direction and the first tilt direction. Taking the tilt angle of the first tilt direction as β as an example, |a| < β means that the angle between the third tilt direction and the first tilt direction is less than β. In this case, -β < a < β, and the third tilt direction is within the range where the first tilt direction deflects β in each of the two opposite directions.

[0010] When |a| is equal to β, a = β or a = -β. This means that the third tilt direction can be one of the two directions obtained by deflecting the first tilt direction β in each of the two opposite directions. Since the tilt angle β of the first tilt direction is the smaller one of the angles formed by the first tilt direction with the row direction (i.e., one of the length direction and the width direction of the image) and the column direction (i.e., the other of the length direction and the width direction of the image) respectively, there is one direction among the two directions obtained by deflecting the first tilt direction β in each of the two opposite directions that is the row direction or the column direction. And for the other direction, since it deflects the same angle β, no matter a = β or a = -β, the invalid area of the third image and the fourth image aligned in the third tilt direction is the same, and is equal to the invalid area when row alignment or column alignment occurs.

[0011] It can be seen that when |a| is equal to β, the invalid area of the third image and the fourth image aligned in the third tilt direction is equal to the invalid area when row alignment or column alignment occurs. Within 45° (the tilt angle ≤ 45°, so |a| is also less than 45°), the smaller the absolute value of the rotation angle, the smaller the invalid area. Therefore, compared with the case where |a| is equal to β (i.e., row alignment or column alignment), when |a| < β, the invalid area of the third image and the fourth image aligned in the third tilt direction is smaller.

[0012] It should be understood that when rotating an integer multiple of 90°, the invalid area does not change. Therefore, the third image and the fourth image aligned in the second tilt direction, which has a relationship of kπ / 2 with the extension angle of the third tilt direction, have the same invalid area. Therefore, in the embodiments of the present application, the difference between the extension angle of the second tilt direction and the extension angle of the first tilt direction is a + kπ / 2, and |a| is less than the tilt angle of the first tilt direction, which can make the invalid area of the third image and the fourth image less than the invalid area when row alignment or column alignment occurs. In this way, when calculating the parallax information, the calculation complexity is reduced, thereby improving the calculation efficiency of the parallax information and reducing the power consumption of the electronic device, and further greatly reducing the user experience.

[0013] In some embodiments of the present application, correcting the first image and the second image to obtain a third image corresponding to the first image and a fourth image corresponding to the second image includes: rotating the first image and the second image by a+kπ / 2 to obtain the third image and the fourth image.

[0014] Ideally, the first camera and the second camera are arranged in the first tilted direction, and the first image and the second image are also aligned in the first tilted direction. It can be seen that the extension angle of the alignment direction of the first image and the second image before and after correction differs by a+kπ / 2. Therefore, by rotating the first image and the second image by the difference angle, the third image and the fourth image can be obtained.

[0015] The embodiment of the present application defines that in a plane rectangular coordinate system, a ray is rotated with the positive semi-axis as the starting point. When rotating counterclockwise, the angle decreases, which is expressed as a negative angle; when the angle rotates clockwise, the angle increases, which is expressed as a positive angle. Therefore, the difference a+kπ / 2 between the extension angle of the second tilt direction and the extension angle of the first tilt direction can be positive or negative. If a+kπ / 2 is positive, it means that the extension angle gradually increases from the first tilt direction to the second tilt direction, and the rotation from the first tilt direction to the second tilt direction requires a clockwise rotation of |a+kπ / 2|; if a+kπ / 2 is negative, it means that the extension angle gradually decreases from the first tilt direction to the second tilt direction, and the rotation from the first tilt direction to the second tilt direction requires a counterclockwise rotation of |a+kπ / 2|. In other words, the direction of rotation of the first image and the second image is determined based on the positive or negative value of a+kπ / 2.

[0016] In some embodiments of the present application, a first image and a second image are corrected to obtain a third image corresponding to the first image and a fourth image corresponding to the second image, including: performing Bouguet stereo correction on the first image and the second image so that the first image after Bouguet stereo correction and the second image after Bouguet stereo correction are aligned in a first direction, and the first direction is row alignment or column alignment (that is, the purpose of Bouguet stereo correction is to correct the first image and the second image to be row alignment or column alignment); wherein the difference between the extension angle of the first direction and the extension angle of the first tilt direction is γ (γ may be positive or negative, depending on the direction of rotation to achieve row alignment or column alignment); rotating the first image after Bouguet stereo correction and the second image after Bouguet stereo correction by a-γ+kπ / 2 to obtain the third image and the fourth image.

[0017] Ideally, the first image and the second image are aligned in the first oblique direction. Therefore, by rotating the first image and the second image by a+kπ / 2, the third image and the fourth image in the second oblique direction can be obtained. However, in the specific implementation process, due to various factors, the first image and the second image are distorted, so that they cannot be aligned in the first oblique direction, and even cannot be coplanar. Simply rotating the first image and the second image cannot obtain the third image and the fourth image aligned in the second oblique direction, and more complex processing is required.

[0018] Currently, there is a Bouguet stereo correction technology that can align the first image and the second image in the row direction or column direction, that is, row alignment or column alignment. This embodiment uses the Bouguet stereo correction technology to first perform row alignment or column alignment correction on the first image and the second image. In this case, the first image and the second image after Bouguet stereo correction are aligned in the first direction. This embodiment further rotates the first image and the second image after Bouguet stereo correction by a-γ+kπ / 2 to obtain the third image and the fourth image aligned in the second oblique direction.

[0019] Bouguet stereo correction corrects the first image and the second image to be aligned in the first direction, and the difference between the extension angle in the first direction and the extension angle in the first oblique direction is γ, so the extension angle in the second oblique direction and the extension angle in the first direction are a-γ+kπ / 2. Therefore, by rotating the first image and the second image after Bouguet stereo correction by a-γ+kπ / 2, the third image and the fourth image can be obtained.

[0020] In some embodiments of the present application, when the inclination angle of the first inclination direction is greater than 22.5°, the extension angle of the second inclination direction is 45°+kπ / 2.

[0021] In this embodiment, the extension angle of the second inclined direction is 45°+kπ / 2, such as 45°, 135°, 225°(-135°), 315°(-45°). The pixel points of each phase in the second inclined direction with an extension angle of 45°+kπ / 2 are diagonally aligned. When matching the third pixel and the fourth image in the second inclined direction of the fourth image, each time a pixel is matched, the next pixel is jumped to match, and the next pixel is unique. It should be understood that when the next pixel is not unique, more computing power is required to match multiple pixels, resulting in low computing efficiency. Based on this, this embodiment can improve computing efficiency by correcting the first image and the second image into the third image and the fourth image aligned in the direction of 45°+kπ / 2.

[0022] It should be noted that, when the inclination angle of the first inclination direction is 22.5°, the extension angle of the first inclination direction is one of 22.5°+kπ / 2 and 62.5°+kπ / 2. Taking 22.5°+kπ / 2 as an example, the difference between 45°+kπ / 2 and 22.5°+kπ / 2, that is, a+kπ / 2=22.5°+kπ / 2. It can be seen that a is just equal to the inclination angle of the first inclination direction (22.5), which does not meet the conditions for obtaining the invalid area area gain. Based on this, in this embodiment, for the case where the inclination angle of the first inclination direction is greater than 22.5°, the first image and the second image can be corrected to the third image and the fourth image aligned in the direction of the extension angle of 45°+kπ / 2 to obtain the invalid area area gain.

[0023] In some embodiments of the present application, the fourth image is matched with the third image along the second inclined direction to determine the disparity information between the feature points of the third image and the corresponding points of the fourth image, including: inputting the third image and the fourth image into a disparity model to obtain the disparity information output by the disparity model, wherein the disparity model is used to match the fourth image with the third image along the second inclined direction and determine the disparity information.

[0024] In this embodiment, a disparity model capable of matching in the second oblique direction of the fourth image is used to determine the disparity information between the feature points of the third image and the corresponding points of the fourth image. According to the foregoing content, the corresponding points of the fourth image are in the second oblique direction of the fourth image. Therefore, the disparity model can match and search for corresponding points in the second oblique direction of the fourth image, thereby effectively calculating the disparity information of the feature points of the third image and the corresponding points of the fourth image in the second oblique direction.

[0025] Specifically, the fourth image is matched with the third image along the second oblique direction and the disparity information is determined, including: using a convolution kernel to process the fourth image so that the fourth image is shifted in sequence by 1 step to N steps along the second oblique direction to obtain N feature maps respectively; N is a positive integer greater than 1; determining the correlation between each feature map and each pixel pair of the third image respectively to obtain N correlation maps; based on the N correlation maps, determining the disparity information; wherein the disparity information is n*d, d is the step size of the fourth image shifted by one step, n is the sequence number of the nth correlation map, the nth correlation map is the correlation map with the highest pixel value among the pixel points with the same coordinates as the feature points in the N correlation maps (referred to as the same coordinate points for short), and n is a positive integer less than or equal to N.

[0026] This embodiment provides a solution for using a disparity model to match the fourth image with the third image along the second oblique direction and determine the disparity information. In this embodiment, the disparity model uses a convolution kernel as a sliding window to slide on the n-1th feature map for convolution. When the n-1th feature map is slid, the fourth image can be offset by one step in the second oblique direction to obtain the nth feature map, 0≤n≤N, and n is a positive integer (n is 0, that is, the 0th feature map is the fourth image). The correlation between the nth feature map and the third image is calculated to obtain the pixel value of the same coordinate point in the nth feature map. The pixel value is the degree of match between a certain pixel point in the second oblique direction through the same coordinate point in the fourth image and the feature point. This process is equivalent to the above-mentioned process of jumping to the next pixel point for matching.

[0027] It should be understood that when the nth correlation graph is the correlation graph with the highest pixel value of the same coordinate point in the N correlation graphs, it means that the same coordinate point in the nth correlation graph is the most matched with the feature point of the third image. From the above content, it can be known that the same coordinate point in the nth correlation graph is a pixel point in the fourth image that is offset by n*d in the second oblique direction passing through the same coordinate point, and the pixel point is the corresponding point. Therefore, the disparity information between the corresponding point and the feature point can be obtained as n*d.

[0028] For example, the convolution kernel is a matrix T (2m+1)×(2m+1) ; T (2m+1)×(2m+1) The element t ij = 1, T (2m+1)×(2m+1) Except t ij The rest of the elements are zero; among them, t ij Point to t mm The direction of the second tilt direction is the same as that of the second tilt direction, t ij T (2m+1)×(2m+1) The element in the i-th row and j-th column in t mm T (2m+1)×(2m+1) The element in the mth row and mth column in T (2m+1)×(2m+1) The central element, m is a positive integer, 1≤i≤2m+1, 1≤j≤2m+1, i, j, m are all positive integers.

[0029] In this embodiment, a convolution kernel is provided that can shift the fourth image in the second tilt direction. When the convolution kernel is used as a sliding window to slide on the fourth image, since only t ij =1, and the rest are zero. Therefore, each time we slide to an area on the fourth image, we will ij The corresponding pixel value is filled to t mm The corresponding pixel position, due to t ij Point to t mm The direction of the second tilt is the same as that of the second tilt, so t ijThe corresponding pixel value is filled to t mm The position of the corresponding pixel point is equivalent to t ij The value of the corresponding pixel point is shifted by one step in the second oblique direction. By sliding the convolution kernel on the fourth image, the value of each pixel point in the fourth image is shifted by one step in the second oblique direction, which is equivalent to shifting the fourth image by one step in the second oblique direction.

[0030] For example, when the extension angle of the second tilt direction is 45°+2kπ, the convolution kernel is In this embodiment, t 11 Point to t 22 The direction is 45°.

[0031] In some embodiments of the present application, before inputting the third image and the fourth image into the disparity model to obtain the disparity information output by the disparity model, the method further includes: obtaining a training sample set, the training sample set including multiple groups of training samples, a single group of training samples including a first training image and a second training image, the first training image and the second training image being aligned in a second inclined direction; training an initial model based on the training sample set to obtain a disparity model; wherein the initial model is used to match the second training image with the first training image along the second inclined direction, and determine the disparity information between feature points of the first training image and corresponding points of the second training image.

[0032] In the process of training the disparity model, the present embodiment uses the first training image and the second training image aligned in the second oblique direction as inputs of the initial model, and constrains the initial model to perform a matching search in the second oblique direction of the second training image. On the one hand, the initial model has the ability to match in the second oblique direction. On the other hand, through the training of the initial model, the parameters of the initial model can be continuously adjusted, so that the accuracy of the initial model in matching search in the second oblique direction and determining the disparity information is improved.

[0033] In some embodiments of the present application, after determining the depth information of the first image, the method further includes: blurring the first image according to the depth information to obtain a blurred image.

[0034] This embodiment provides an application scenario of depth information, namely, an image blurring scenario. In the image blurring scenario, blurring the first image depends on the depth information of the first image. Based on this, when the amount of calculation and the complexity of calculation of the parallax information are reduced and the calculation efficiency is improved, the efficiency of obtaining the depth information of the first image is also improved accordingly. Naturally, the efficiency of blurring the first image to obtain the blurred image is improved.

[0035] In a second aspect, an embodiment of the present application provides an electronic device. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the method described in various possible implementations in the first aspect is implemented.

[0036] In a third aspect, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in various possible implementations of the first aspect is implemented.

[0037] In a fourth aspect, an embodiment of the present application provides a chip system. The chip system includes a processor and a memory; wherein the processor is coupled to the memory, and the memory is used to store programs or instructions, and when the program or instructions are executed by the processor, the chip system implements the method described in various possible implementations in the first aspect.

[0038] It can be understood that the beneficial effects that can be achieved by the technical solutions of the second to fourth aspects provided above can refer to the beneficial effects of the method for determining depth information in the first aspect and any possible design thereof, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1A A schematic diagram of the spatial distance provided in an embodiment of the present application;

[0040] Figure 1B A schematic diagram of depth information provided by an embodiment of the present application;

[0041] Figure 1C A schematic diagram of the tilt direction, tilt angle and extension angle provided in an embodiment of the present application;

[0042] Figure 2 A comparison chart of the effects before and after image blurring provided in the embodiment of the present application;

[0043] Figure 3 It is a schematic diagram of the principle of binocular stereo matching technology;

[0044] Figure 4 A schematic diagram of the relationship between the horizontally arranged binocular cameras and their imaging provided in an embodiment of the present application;

[0045] Figure 5 A schematic diagram of the relationship between the obliquely arranged binocular cameras and their imaging provided in an embodiment of the present application;

[0046] Figure 6 A schematic diagram of the Bouguet stereo correction process provided in an embodiment of the present application;

[0047] Fig. 7A A schematic diagram of the back structure of an electronic device provided in an embodiment of the present application;

[0048] Figure 7B A schematic diagram of an application architecture of an electronic device provided in an embodiment of the present application;

[0049] Figure 7C A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application;

[0050] Figure 8 A flowchart of a method for image processing provided in an embodiment of the present application is shown in FIG1 ;

[0051] Fig. 9 A schematic diagram of a shooting interface with a portrait mode provided in an embodiment of the present application;

[0052] Fig.10 A schematic diagram of the position distribution of the first tilt direction and the second tilt direction provided in an embodiment of the present application;

[0053] Fig.11 A comparison diagram of the invalid area after the image provided by the embodiment of the present application is rotated clockwise and counterclockwise by the same angle;

[0054] Fig.12 A sine curve diagram provided for an embodiment of the present application;

[0055] Fig.13 A schematic diagram of a third image and a fourth image aligned at 45° provided in an embodiment of the present application;

[0056] Fig.14 A schematic diagram of a method for image processing provided in an embodiment of the present application Figure 2 ;

[0057] Fig.15 A schematic diagram of a process of correcting a first image and a second image to obtain a third image and a fourth image according to an embodiment of the present application;

[0058] Fig.16 A schematic diagram of the principle of matching a third image and a fourth image in a second tilt direction of the fourth image provided in an embodiment of the present application;

[0059] Fig.17 An example flow chart of determining disparity information using a disparity model provided in an embodiment of the present application;

[0060] Fig.18 A before-and-after comparison diagram of the fourth image provided in an embodiment of the present application shifted by one step in the 45° direction;

[0061] Fig.19A schematic diagram of a fourth image provided by an embodiment of the present application being shifted by n steps in a second oblique direction;

[0062] Fig. 20 An example flow chart of a disparity model provided in an embodiment of the present application outputting disparity information based on a third image and a fourth image;

[0063] Fig.21 A flowchart of a method for training a disparity model provided in an embodiment of the present application;

[0064] Fig. 22 A schematic diagram of a process for obtaining a training sample set provided in an embodiment of the present application;

[0065] Fig.23 A schematic diagram of the processing process of the original training samples provided in the embodiment of the present application;

[0066] Fig.24 A schematic diagram of a method for image processing provided in an embodiment of the present application Figure 3 . DETAILED DESCRIPTION

[0067] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0068] The terms "first" and "second" in the specification and claims herein are used to distinguish different objects rather than to describe a specific order of objects. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "multiple" refers to two or more than two, for example, multiple processing units refer to two or more processing units, etc.; multiple elements refer to two or more elements, etc.

[0069] In the embodiments of the present application, the words "exemplarily" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.

[0070] In the embodiment of the present application, the angle between the two objects is the minimum positive angle. In the embodiment of the present application, the angle of counterclockwise rotation is "-", and the angle of clockwise rotation is "+". In the embodiment of the present application, a plane rectangular coordinate system with the following characteristics is used for explanation: a ray is rotated with the positive half axis of the x-axis as the starting point. When rotating counterclockwise, the angle decreases, which is expressed as a negative angle; when the angle rotates clockwise, the angle increases, which is expressed as a positive angle. Of course, in other embodiments, in the plane rectangular coordinate system, a ray can be rotated with the positive half axis of the x-axis as the starting point. When rotating counterclockwise, the angle increases, which is expressed as a positive angle; when the angle rotates clockwise, the angle decreases, which is expressed as a negative angle.

[0071] To facilitate understanding, the terms involved in this application are first described.

[0072] 1. Binocular image: two images obtained by two cameras at different positions shooting the same target.

[0073] for example, Fig. 7A In the mobile phone 100 shown, two images obtained by shooting the same shooting target with the main camera W and the ultra-wide-angle camera UW are binocular images.

[0074] 2. Parallax: In the field of image processing, parallax refers to the position difference of the same point in the physical world in the binocular images captured by cameras at two different positions. It is inversely proportional to the spatial distance of the point (the distance between the point and the midpoint between the two cameras) and the distance between the two cameras. Specifically, z = b*f / d, z is the spatial distance, b is the baseline distance, d is the parallax, and f is the focal length of the two lenses.

[0075] for example, Figure 1A , point K is a point in the physical world; the midpoint O2 between camera B1 and camera B2 is the midpoint of the line connecting camera B1 and camera B2, and the midpoint O2 is equidistant from camera B1 and camera B2 respectively; the distance from point K to the midpoint O2 is the spatial distance of point K.

[0076] When calculating the disparity, one image is selected from the binocular images as the main image (also called the reference image), for example, the image taken by a camera with higher imaging quality is used as the main image; the other image is used as the secondary image (also called the auxiliary image or the image to be matched). The feature point in the main image (a pixel point in the main image, for the sake of easy distinction, it is called a feature point) is used as the reference point, and then the surrounding area of ​​the pixel point with the same coordinates as the feature point in the secondary image is searched to find the corresponding point that best matches the feature point, and finally the disparity between the feature point and the corresponding point is calculated.

[0077] 3. Depth information: The depth information of a pixel in an image refers to the distance between the point of the pixel in the physical world obtained from the image and the camera in the direction of the optical axis of the camera, which can be understood as the above z. It can be seen that the depth information can be inferred from the parallax.

[0078] for example, Figure 1B In the figure, pixel K' is a pixel in the image, and point K is a point of pixel K' in the physical world. Dotted line S1 is the optical axis of camera B1, and the direction in which dotted line S1 extends is the direction of the optical axis of camera B1. The depth information of pixel K' is shown in figure L, where L is the distance between point K and camera B1 in the direction of the optical axis of camera B1.

[0079] 4. Tilt direction, tilt angle of the tilt direction, and extension angle of the tilt direction

[0080] The tilt direction refers to the direction in which the directions of the two perpendicular axes of the plane rectangular coordinate system are deflected, that is, there is an angle between the tilt direction and the directions of the two perpendicular axes. In the embodiment of the present application, the angle is positive and negative, and the angle is the minimum positive angle. There is an angle between the tilt direction and the directions of the two perpendicular axes, which means that there is an angle between the reference line extending along the tilt direction and the two perpendicular axes. When the angles in the two directions are involved later, it can be understood by reference.

[0081] The tilt angle of the tilt direction is the smaller of the angles at which the tilt direction is deflected relative to the directions of two perpendicular axes, that is, the smaller of the angles between the tilt direction and the directions of the two perpendicular axes. From this definition, it can be seen that the tilt angle is ≤45° and is a positive angle.

[0082] The extension angle of the tilt direction is the angle required to rotate in the plane rectangular coordinate system with the 0° direction as the starting side and the tilt direction as the terminal side. Taking the 0° direction as the starting side means taking the reference line extending along the 0° direction as the starting side; taking the tilt direction as the terminal side means taking the reference line extending along the tilt direction as the terminal side. Similar descriptions can be referred to for understanding in the following.

[0083] It should be understood that the 0° direction as the starting edge can be rotated counterclockwise by a certain angle to the terminal edge where the tilt direction is located, or it can be rotated clockwise by a certain angle to the terminal edge where the tilt direction is located. Moreover, after the 0° direction as the starting edge is rotated by a certain angle to the terminal edge where the tilt direction is located, it can also reach the terminal edge where the tilt direction is located by continuing to rotate 2kπ. Therefore, the same tilt direction can have multiple extension angles, and the multiple extension angles are positive and negative. Except for separate descriptions, the extension angles within ±2π (i.e., ±360°) are used as an example for description in the embodiments of the present application.

[0084] For example, please refer to Figure 1C , Figure 1C In the plane rectangular coordinate system O1-X1Y1 shown, the direction S2 is deflected compared to the positive direction of the X1 axis and the positive direction of the Y1 axis, so the direction S2 is an inclined direction. The angle a1 between the direction S2 and the positive direction of the X1 axis is 60°; the angle a2 between the direction S2 and the positive direction of the Y1 axis is 30°, and the angle a1 and the angle a2 are complementary. The inclination angle of the direction S2 is the smaller of the angle a1 and the angle a2. Figure 1C The inclination angle of the middle direction S2 is angle a2 = 30°.

[0085] Figure 1C In the equation, the positive direction of the X1 axis is the 0° direction, with the 0° direction as the starting edge and the direction S2 as the ending edge, the angle required to rotate from the starting edge to the ending edge is a3+2kπ, that is, the extension angle of the direction S2 is a3+2kπ. Wherein, a3 is the minimum angle required to rotate counterclockwise or clockwise from the 0° direction (i.e. the starting edge) to the direction S2 (i.e. the ending edge). If rotating counterclockwise from the 0° direction to the direction S2, a3=-315°; if rotating clockwise from the 0° direction to the direction S2, a3=60°.

[0086] In the embodiment of the present application, for a more intuitive understanding, the tilt direction can be referred to by the extension angle a3. For example, Figure 1C The S2 direction can be referred to as the 60° direction or the -315° direction.

[0087] 5. Align the two images in a certain direction

[0088] The premise for aligning two images in a certain direction is that the two images are in the same plane, that is, coplanar. Take the alignment of two images, a first image and a second image, in the oblique direction Z as an example. When the coplanar first image and second image overlap, if the two corresponding pixel points in the first image and the second image are aligned in the oblique direction Z, then the first image and the second image are aligned in the oblique direction Z. The embodiment of the present application is explained with the oblique direction Z pointing from the pixel point in the second image to the pixel point in the first image. Of course, in other embodiments, the oblique direction Z can also be pointed from the pixel point in the first image to the pixel point in the second image, and the schemes related to the definition of this direction need to be adjusted synchronously.

[0089] The two corresponding pixels in the first image and the second image refer to the pixels that image the same point in the physical world in the first image and the second image. For example, Figure 5 In (b), a pixel point M on the head of a person in the left image and a pixel point M' on the head of a person in the right image are two corresponding pixel points. Figure 5 In (b), the left and right images are aligned in the tilt direction Z1, Figure 5As can be seen from (c) in the figure, the tilt direction Z1 points from the pixel point M' in the right figure to the pixel point M in the left figure.

[0090] The technical solutions involved in the embodiments of the present application are described below in conjunction with the accompanying drawings.

[0091] In some usage scenarios, mobile phones and other electronic devices involve the application of depth information. For example, when using the "large aperture" mode or "portrait" mode of a mobile phone camera to take pictures, it is necessary to use a blur algorithm to blur the image to obtain a blurred image, so that the blurred image is close to the natural blur effect of a SLR (i.e., the subject is clear and the background is blurred in layers).

[0092] For example, please refer to Figure 2 , Figure 2 This is a comparison chart of the effects before and after the image blurring provided in the embodiment of the present application. Figure 2 (a) is the image before blurring (i.e., the image to be blurred), in which the human body and the background are clearly displayed; Figure 2 (b) in the figure is the blurred image (i.e., the blurred image), in which the human body is clearly displayed while the background is blurred.

[0093] It should be noted that the blur effect depends on the depth information of the image to be blurred. In order to obtain the depth information of the image to be blurred, usually a main image and a secondary image (i.e., a binocular image) are first captured from different perspectives by dual cameras, and then a disparity map of the main and secondary images is obtained by using binocular stereo matching technology. Finally, the depth information of the image to be blurred, such as the main image, is obtained by converting the disparity map.

[0094] Specifically, whether binocular stereo matching technology adopts traditional algorithms or deep learning algorithms, its general principles are as follows:

[0095] For example, please refer to Figure 3 , Figure 3 Schematic diagram of the principle of binocular stereo matching technology. Figure 3 One grid represents one pixel. Figure 3 The left image is the main image, and the right image is the secondary image.

[0096] For a pixel point P on the left image (for the sake of distinction, this embodiment of the application refers to it as a feature point), search for the pixel point with the highest matching degree with the feature point P among the pixels (such as P1, P2, P3, etc.) within a certain search range Dmax (set as needed) on the right image, that is, search for the pixel point that best matches the feature point P (for the sake of distinction, this embodiment of the application refers to the pixel point that best matches the feature point P as a corresponding point). This process is referred to as a matching search in this embodiment of the application. Finally, the disparity between each pair of feature points and corresponding points in the left and right images is calculated.

[0097] because Figure 3 The left and right images are both two-dimensional images, and matching and searching for corresponding points in two-dimensional space is very time-consuming. In order to reduce the matching search range, it is necessary to perform epipolar (Bouguet) stereo correction to align the rows or columns of the left and right images, so as to reduce the matching search in two dimensions to one dimension, thereby achieving the purpose of reducing the matching search range. It should be noted that Bouguet stereo correction is a stereo correction technology that corrects the imaging plane of the binocular image to the same plane (i.e., coplanar) through steps such as rotation and translation, and aligns the pixels of each row (or each column) of the binocular image. The following is combined with Figure 4 and Figure 5 Let's take row alignment as an example.

[0098] Please refer to Figure 4 , Figure 4 A schematic diagram of the relationship between the horizontally arranged binocular cameras and their imaging provided in an embodiment of the present application.

[0099] Figure 4 The binocular cameras (camera B1 and camera B2) shown in (a) are arranged in the horizontal direction. It should be noted that Figure 4 (a) in the figure shows the back of the phone from the perspective of the phone being taken horizontally. Figure 4 In (a), an O-XY coordinate system is established. The X-axis direction is the length direction of the mobile phone, and the Y-axis direction is the width direction of the mobile phone. Cameras B1 and B2 are arranged in the X-axis direction (the positive direction of the X-axis in the figure). Figure 4 From the horizontal shooting perspective shown in (a), the above-mentioned camera B1 and camera B2 are arranged horizontally. Of course, in the vertical shooting perspective, the camera B1 and camera B2 arranged in the X-axis direction (such as the positive direction of the X-axis in the figure) can also be regarded as being arranged vertically.

[0100] Figure 4 The left and right images of (b) are Figure 4 (a) is taken by camera B1 and camera B2. Figure 4 The left image shown in (b) is the main image. Figure 4 The right image shown in (b) is the secondary image. According to the projection principle, if the camera is on the right, the person features in the image are on the left; if the camera is on the left, the person features in the image are on the right. Therefore, the left image is taken by camera B1, and the right image is taken by camera B2.

[0101] Figure 4 The left image and the right image shown in (b) are both in the pixel coordinate system o-xy and aligned in the row direction (in this embodiment of the application, alignment in the row direction is referred to as row alignment). The so-called row direction refers to the x-axis direction of the pixel coordinate system o-xy; the y-axis direction of the pixel coordinate system o-xy is the column direction. Figure 4 The row alignment of the left and right images shown in (b) means that the corresponding two pixels in the left and right images are aligned in the x-axis direction. For example, Figure 4 A pixel point M on the head of a person in the left image and a pixel point M' on the head of a person in the right image are two corresponding pixel points, and the pixel point M and the pixel point M' are aligned in the positive direction of the x-axis, that is, the left image and the right image are aligned in rows. Of course, in other embodiments, the pixel point M and the pixel point M' can also be aligned in the negative direction of the x-axis to achieve the row alignment of the left image and the right image. It should be understood that column alignment refers to the alignment of the left image and the right image in the column direction (the embodiment of the present application will refer to the alignment in the row direction as row alignment), that is, the two corresponding pixel points in the left image and the right image are aligned in the y-axis direction (such as the positive or negative direction of the y-axis).

[0102] By overlapping the left and right images, we can get Figure 4 (c) in Figure 4 (c) in the figure can more intuitively show the positional relationship between the two corresponding pixels in the left and right images. Figure 4 As can be seen from (c) in the figure, when the left and right images are aligned in rows, there is a difference Δx between pixel point M and pixel point M' in the x-axis direction, but no difference in the y-axis direction, that is, the x-coordinates are different but the y-coordinates are the same. It should be understood that column alignment means that the y-coordinates are different but the x-coordinates are the same.

[0103] It should be understood that when the left image and the right image are aligned, the two corresponding pixel points in the left image and the right image are aligned in the x-axis direction. Therefore, when the pixel point M in the left image is used as a feature point for matching search of corresponding points, it is only necessary to perform matching search in the x-axis direction of the right image, and the corresponding point M' that best matches the feature point M in the left image can be searched in the right image, thereby achieving the purpose of reducing the matching search range. It should be understood that the difference Δx between the x-coordinates of the feature point M and the corresponding point M' is the disparity between the feature point M and the corresponding point M'.

[0104] It can be seen that since the cameras B1 and B2 are arranged in the X-axis direction, the left and right images captured are aligned in rows without correction. Figure 4 (b) is an ideal image. In actual implementation, even if cameras B1 and B2 are arranged horizontally, the left and right images taken by them may not be aligned in rows. However, when the two cameras are arranged diagonally, the two images taken by them are not aligned in rows. Figure 4 Provide explanation.

[0105] For example, please refer to Figure 5 , Figure 5 A schematic diagram of the relationship between the obliquely arranged binocular cameras and their imaging provided in the embodiment of the present application. Figure 5The definition of the coordinate system in (a) can be found in Figure 3 Definition of the coordinate system in (a).

[0106] Figure 5 The binocular cameras shown in (a) (i.e., camera B1 and camera B2) are arranged obliquely. The oblique arrangement of camera B1 and camera B2 means that the arrangement direction Z1 between camera B1 and camera B2 is inclined relative to both the X-axis direction and the Y-axis direction, and the arrangement direction Z1 refers to the direction in which camera B1 points to camera B2. When camera B1 and camera B2 are arranged obliquely, there are differences in the coordinates of camera B1 and camera B2 in both the X-axis direction and the Y-axis direction, that is, there are differences in the X-coordinate and Y-coordinate of camera B1 and camera B2. Taking the coordinates (X1, Y1) of camera B1 and the coordinates (X2, Y2) of camera B2 as examples, the difference ΔX between X1 and X2 is not equal to zero, and the difference ΔY between Y1 and Y2 is not equal to zero.

[0107] It should be understood that the arrangement direction Z1 is inclined relative to both the X-axis direction and the Y-axis direction, which means that the arrangement direction Z1 in the figure has an angle b1 (b1=arctan(ΔX / ΔY)=30°) with the X-axis and an angle b2 (b2=arctan(ΔY / ΔX)=30°) with the Y-axis direction (the Y-axis in the figure). In the embodiment of the present application, the smaller of the angles b1 and b2 is regarded as the arrangement angle b between the camera B1 and the camera B2. Figure 5 In (a), the arrangement angle b between camera B1 and camera B2 is angle b1. It can be seen that the arrangement direction Z1 is an inclined direction. Therefore, the so-called oblique arrangement means that camera B1 and camera B2 are arranged in an inclined direction, and the arrangement angle of the arrangement direction Z1 is the inclination angle of the inclined direction.

[0108] Figure 5 The left image in (b) is shown by Figure 5 The image (a) is taken by camera B1, which is the main image, and the right image is taken by Figure 5 The image (a) is taken by camera B2 and is the secondary image. Figure 5 The left and right images shown in (b) are both in the pixel coordinate system o-xy. The x-coordinate and y-coordinate of pixel point M and pixel point M' are different. By overlapping the left and right images, we can get Figure 5 (c) in . Figure 5 As can be seen from (c) in , the difference Δx between the x-coordinates of pixel point M and pixel point M' is not equal to zero, and the difference in y-coordinates is also not equal to zero. Pixel point M and pixel point M' are not aligned in the x-axis direction and the y-axis direction, that is, they are neither aligned in row nor in column. Figure 5The left and right images shown in (b) of the figure are subjected to Bouguet stereo correction, so that the left and right images are aligned in rows or columns after correction. Figure 6 , taking row alignment as an example to illustrate.

[0109] For example, please refer to Figure 6 , Figure 6 Shown Figure 5 Schematic diagram of the process of row alignment after Bouguet stereo correction on the left and right images in the middle.

[0110] Depend on Figure 6 It can be seen that by Figure 6 The left and right images shown in (a) are rotated to obtain Figure 6 The left and right images shown in (b) of Figure 6 The left and right images shown in (b) overlap to obtain Figure 6 (c) in . Figure 6 As can be seen from (c), Figure 6 The left and right images shown in (b) achieve row alignment.

[0111] It should be noted that since image processing requires the image to be square, Figure 6 After the left and right pictures shown in (a) in the figure, you need to Figure 4 Fill some pixels around the image. Invalid pixels refer to the original image (such as Figure 6 The pixels not included in the left and right images (shown in (a) in the figure) are invalid pixels, for example, invalid pixels can be filled noise pixels, etc. This causes the corrected left and right images to introduce a large number of invalid areas (invalid areas refer to the areas composed of invalid pixels filled around the original image after the original image is rotated, such as Figure 6 ), which increases the image size. It should be noted that when calculating disparity information (for example, for deep learning, such as convolution), the computational complexity and amount of computation are proportional to the image size. Based on this, when a large number of invalid areas are introduced after the left and right images are aligned, the complexity of calculating disparity information is increased, resulting in low computational efficiency of disparity information, which greatly reduces the user experience. The computing power of mobile devices such as mobile phones needs to focus on the energy consumption ratio of the algorithm.

[0112] In order to solve the above problems, an embodiment of the present application provides a method for image processing, which is applied to an electronic device with two cameras arranged obliquely. In the method for image processing, the first image and the second image taken by the two cameras arranged obliquely on the same shooting target are corrected so that the corrected first image (the embodiment of the present application refers to it as the third image) and the corrected second image (the embodiment of the present application refers to it as the fourth image) are aligned in a certain oblique direction (the embodiment of the present application refers to it as the second oblique direction). When the third image and the fourth image are aligned in the second oblique direction, the corresponding points of the fourth image will appear in the second oblique direction of the fourth image. Therefore, the method for image processing matches the fourth image with the third image along the second oblique direction, determines the parallax information of the feature points of the third image and the corresponding points of the fourth image, and then determines the information of the first image based on the parallax information.

[0113] For the sake of convenience of explanation and distinction, the embodiment of the present application refers to the arrangement direction of the two obliquely arranged cameras as the first oblique direction. Taking the two obliquely arranged cameras as an example, the first camera and the second camera are arranged in the first oblique direction, and the inclination angle of the first oblique direction is the arrangement angle of the first camera and the second camera.

[0114] In the method for image processing, the difference between the extension angle of the second oblique direction and the extension angle of the first oblique direction is a+kπ / 2, the absolute value of a is smaller than the oblique angle of the first oblique direction, and k is an integer. In this way, the invalid area of ​​the third image and the fourth image can be made smaller than the invalid area when the rows are aligned or the columns are aligned, thereby obtaining the benefit of the invalid area. In this way, when calculating the disparity information, the computational complexity is reduced, thereby improving the computational efficiency of the disparity information, and thus greatly reducing the user experience. The reasons will be analyzed in detail in conjunction with the accompanying drawings later, and will not be elaborated here.

[0115] Exemplarily, the electronic device in the embodiments of the present application may be a mobile phone, a tablet computer, a desktop, a laptop, a handheld computer, a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) device, a virtual reality (VR) device, or the like. The embodiments of the present application do not impose any special restrictions on the specific form of the device.

[0116] Take the electronic device 100 as a mobile phone as an example, please refer to Fig. 7A , Fig. 7A A schematic diagram of the back structure of an electronic device 100 provided in an embodiment of the present application.

[0117] The mobile phone 100 includes a multi-camera module 101, which refers to a camera module with multiple cameras. Fig. 7A The multi-camera module 101 shown has three cameras, namely a main camera W, an ultra-wide-angle camera UW, and a telephoto camera T. It should be understood that in other embodiments, the multi-camera module 101 may include more or fewer cameras, such as two, four, etc., which is not limited in the present embodiment. Fig. 7A An O-XY coordinate system is established in which the X-axis direction is the length direction of the mobile phone 100 and the Y-axis direction is the width direction of the mobile phone 100.

[0118] For aesthetic reasons, in the mobile phone 100, there is a situation where the cameras are arranged diagonally between each other. The definition of diagonal arrangement can refer to Figure 5 The definition of the oblique arrangement of cameras B1 and B2 in . For example, Fig. 7A In the figure, except for the vertical arrangement of the main camera W and the telephoto camera T, the ultra-wide-angle camera UW and the telephoto camera T, as well as the main camera W and the ultra-wide-angle camera UW are arranged diagonally. Among them, the main camera W and the ultra-wide-angle camera UW are arranged in the inclined direction Z1, that is, the direction in which the main camera W points to the ultra-wide-angle camera UW is the inclined direction Z1, and the arrangement angle (that is, the inclination angle of the inclined direction Z1) is 30°; the ultra-wide-angle camera UW and the telephoto camera T are arranged in the inclined direction Z2, that is, the direction in which the ultra-wide-angle camera UW points to the telephoto camera T is the inclined direction Z2, and the arrangement angle (that is, the inclination angle of the inclined direction Z2) is also 30°.

[0119] It should be understood that the first camera and the second camera may refer to Fig. 7A The ultra-wide-angle camera UW and the telephoto camera T arranged diagonally in the middle may also refer to the main camera W and the ultra-wide-angle camera UW arranged diagonally. In the subsequent embodiments, the method for image processing provided in the embodiments of the present application is described by taking the first camera as the main camera W and the second camera as the ultra-wide-angle camera UW as an example.

[0120] For example, Figure 7B is a schematic diagram of the application architecture (including software system and some hardware) of the electronic device provided in the embodiment of the present application. The application architecture can be Fig. 7A The application architecture of the electronic device shown.

[0121] like Figure 7B As shown, the application architecture is divided into several layers, each with clear roles and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the application architecture can be divided into five layers, from top to bottom, namely, the application layer, the application framework layer, the hardware abstraction layer HAL, the driver layer, and the hardware layer.

[0122] like Figure 7B As shown, the application layer includes the camera and the gallery. It can be understood that Figure 7B The application layer may include other applications, which are not limited in this application. For example, the application layer may include information, alarm clock, weather, stopwatch, compass, timer, flashlight, calendar, Alipay and other applications.

[0123] Generally speaking, applications are developed using the Java language and are completed by calling the application programming interface (API) and programming framework provided by the application framework layer.

[0124] The application framework layer provides an application programming interface (API) and a programming framework for the application programs in the application layer. The application framework layer includes some predefined functions.

[0125] For example, the application framework layer may include a camera access interface. The camera access interface may include a camera service and a camera management. The camera service may be used to provide an interface for accessing the camera, and the camera management may be used to provide an access interface for managing the camera.

[0126] In addition, the application framework layer may also include content providers, resource managers, notification managers, window managers, view systems, phone managers, etc. Similarly, the camera application may also call content providers, resource managers, notification managers, window managers, view systems, etc. according to actual business needs. The embodiments of the present application do not impose any restrictions on this.

[0127] The hardware abstraction layer is used to abstract the hardware. For example, it can encapsulate the driver in the driver layer and provide a calling interface to the application framework layer, shielding the implementation details of the low-level hardware.

[0128] For example, the hardware abstraction layer may include a camera hardware abstraction layer and an image processing module. The camera hardware abstraction layer includes multiple camera abstract devices. The camera abstract device may be a software interface facing the driver layer, used to interact with the driver layer for data, such as calling the camera device driver in the driver layer. For another example, the image processing module may process the image transmitted back by the camera device.

[0129] In an embodiment of the present application, the image processing module can be used to execute the method provided in the embodiment of the present application, and determine the depth information of the first image according to the first image and the second image acquired by the first camera and the second camera. Optionally, the depth information can be output in the form of a depth map. In a specific embodiment, the image processing module may include a correction module, a disparity determination module, and a depth determination module. The correction module is used to correct the first image and the second image acquired by the first camera and the second camera so that the corrected first image (i.e., the third image) and the corrected second image (i.e., the fourth image) are aligned in the second oblique direction; the disparity determination module matches the fourth image with the third image along the second oblique direction and determines the disparity information between the two; thereafter, the depth determination module can calculate the depth information of the first image according to the disparity information determined by the disparity determination module. In addition, the image processing module may also include a blur processing module, and the blur processing module may further blur the first image according to the depth information of the first image determined by the depth determination module.

[0130] Specifically, the image processing module sends the image to the graphics signal processor driver in the driver layer so that the graphics signal processor driver calls the graphics processor in the hardware layer to perform digital signal processing. The graphics processor can return the processed image to the image processing module through the graphics processor driver. In addition, the image output by the graphics signal processor can be sent to the camera device driver. The camera device driver can send the image output by the image signal processor to the camera hardware abstraction layer. The camera hardware abstraction layer can send the image to the image processing module for further processing.

[0131] It should be understood that the image processing module can also be placed in other layers. As a possible implementation, the algorithm processing module can be placed in the application layer or the application framework layer.

[0132] The driver layer is used to drive hardware resources. The driver layer can include multiple driver modules. Figure 7B As shown, the driver layer includes camera device driver and the like.

[0133] In addition, the hardware layer includes hardware modules that can be driven, such as a camera device. For example, the camera device includes multiple cameras, such as a first camera, a second camera, ...

[0134] For example, the camera application in the application layer can be displayed on the screen of the electronic device 100 in the form of an icon. When the icon of the camera application is clicked by the user to trigger it, the electronic device 100 starts running the camera application. When the photo control in the "portrait" mode or "large aperture" mode of the camera application is clicked by the user to trigger a shooting instruction, the shooting instruction can be sent to the camera hardware abstraction layer through the camera access interface. The camera hardware abstraction layer calls the camera abstract device 1 and the camera abstract device 2 to start the camera device driver, and turns on the first camera and the second camera to shoot to obtain the first image and the second image.

[0135] In the process of shooting with the first camera and the second camera, the camera application can call the image processing module in the application framework layer to process the first image and the second image acquired by the first camera and the second camera to determine the depth information of the first image. The determined depth information can be used in subsequent blur processing or other scenarios, which is not limited in the embodiments of the present application.

[0136] The above describes in detail the software system used in the embodiment of the present application. Figure 7C The hardware system of the electronic device 100 is described.

[0137] Figure 7C A schematic diagram of the hardware structure of an electronic device 100 suitable for the present application is shown.

[0138] The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include multiple sensors.

[0139] It should be noted that Figure 7C The structure shown does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include Figure 7C More or fewer components than those shown, or the electronic device 100 may include Figure 7C Combinations of some of the components shown, or the electronic device 100 may include Figure 7C Subcomponents of some of the components shown. For example, Figure 7C The illustrated proximity light sensor 180G may be optional. Figure 7C The components shown may be implemented in hardware, software, or a combination of software and hardware.

[0140] The processor 110 may include one or more processing units. For example, the processor 110 may include at least one of the following processing units: an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and a neural-network processing unit (NPU). Different processing units may be independent devices or integrated devices.

[0141] The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of instruction fetching and execution.

[0142] The processor 110 may also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory may store instructions or data that the processor 110 has just used or cyclically used. If the processor 110 needs to use the instruction or data again, it may be directly called from the memory. This avoids repeated access, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0143] In some embodiments, the processor 110 may include one or more interfaces. For example, the processor 110 may include at least one of the following interfaces: an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a SIM interface, and a USB interface.

[0144] Figure 7C The connection relationship between the modules shown is only a schematic illustration and does not constitute a limitation on the connection relationship between the modules of the electronic device 100. Optionally, the modules of the electronic device 100 may also adopt a combination of multiple connection modes in the above embodiments.

[0145] The wireless communication function of the electronic device 100 can be implemented through components such as the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, and the baseband processor.

[0146] The electronic device 100 can realize the display function through the GPU, the display screen 194 and the application processor. The GPU is a microprocessor for image processing, which connects the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.

[0147] The display screen 194 can be used to display images or videos. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a mini light-emitting diode (Mini LED), a micro light-emitting diode (Micro LED), a micro OLED (Micro OLED) or a quantum dot light emitting diode (QLED). In some embodiments, the electronic device 100 may include 1 or F display screens 194, where F is a positive integer greater than 1.

[0148] The electronic device 100 can realize the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194 and the application processor.

[0149] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and the light is transmitted to the camera photosensitive element through the lens. The light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can perform algorithm optimization on the noise, brightness and color of the image. The ISP can also optimize the exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0150] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard red green blue (RGB), YUV or other format. In an embodiment of the present application, the electronic device 100 includes a plurality of cameras 193. The plurality of cameras 193 include at least a first camera and a second camera, and the first camera and the second camera are arranged obliquely.

[0151] The digital signal processor is used to process digital signals, and can process not only digital image signals but also other digital signals. For example, when the electronic device 100 is selecting a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.

[0152] Video codecs are used to compress or decompress digital videos. The electronic device 100 may support one or more video codecs. Thus, the electronic device 100 may play or record videos in a variety of coding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.

[0153] NPU is a processor that draws on the structure of biological neural networks, such as the transmission mode between neurons in the human brain to quickly process input information, and can also continuously self-learn. Through NPU, the electronic device 100 can realize intelligent cognition and other functions, such as image recognition, face recognition, voice recognition and text understanding.

[0154] For example, please refer to Figure 8 , Figure 8 A flowchart diagram of a method for image processing provided in an embodiment of the present application is shown first. Figure 8 The method for image processing shown can be applied to 7A to 7C In the electronic device shown. Figure 8 The method for image processing shown includes the following steps S801 to S804:

[0155] S801, when photographing the same shooting target, the electronic device obtains a first image based on a first camera and obtains a second image based on a second camera.

[0156] That is to say, the first image and the second image are a pair of binocular images. The first camera and the second camera are arranged in the first oblique direction, that is, the direction in which the first camera points to the second camera is the first oblique direction. It can be seen that the first camera and the second camera are arranged obliquely. Ideally, the first image and the second image are also aligned in the first oblique direction.

[0157] Combination Fig. 7A , when the first camera is the main camera W and the second camera is the ultra-wide-angle camera UW, the first tilt direction is the tilt direction Z1 shown in the figure, the tilt angle of the tilt direction Z1 is 30°, and the extension angle of the tilt direction Z1 is 30° (or -330°). In this case, the image captured by the main camera W is the first image, and the image captured by the ultra-wide-angle camera UW is the second image. Fig. 7A The arrangement and Figure 5 The arrangement of cameras B1 and B2 is the same. Figure 5 Explanation is based on the display.

[0158] Please continue to refer to Figure 5 , ideally, Figure 5 The left picture shown in (b) is the first image taken by camera B1. Figure 5 The right image shown in (b) is the second image captured by camera B2. Figure 5 As can be seen from (c) in the figure, the left image and the right image are aligned in the tilt direction Z1, that is, the direction from the pixel point M' to the pixel point M in the figure is the tilt direction Z1.

[0159] In some embodiments, the method of triggering the electronic device to execute S801 may be, but is not limited to: the electronic device receives a shooting instruction, the shooting instruction is used to instruct the electronic device to shoot a binocular image. In this case, the electronic device responds to the shooting instruction and starts the first camera and the second camera to shoot the same shooting target, thereby achieving the acquisition of the first image and the second image in S801.

[0160] For example, combining Fig. 7A and Fig. 9 As shown, Fig. 7A The mobile phone 100 shown in FIG. 7 displays a shooting interface 901 in the “portrait” mode, and the shooting interface 901 includes a shooting control 901a. In response to a user clicking the shooting control 901a, the electronic device receives a shooting instruction. In response to the shooting instruction, the electronic device starts the main camera W and the ultra-wide-angle camera UW shown in FIG. 7 to shoot the same shooting target (such as a person), thereby obtaining Figure 5 The left and right images are shown in (b) of Figure 1. Since the “portrait” mode requires blurring of the image, Fig. 9The shooting instruction instructs the mobile phone to shoot a binocular image. Of course, in other embodiments, Fig. 7A The mobile phone 100 shown may also have other shooting modes that require image blur processing, such as a "large aperture" mode. In this case, the user can trigger the above-mentioned shooting instruction in these shooting modes that require image blur processing.

[0161] S802, correcting the first image and the second image to obtain a third image corresponding to the first image and a fourth image corresponding to the second image.

[0162] The third image is the corrected first image, and the fourth image is the corrected second image. The third image and the fourth image are aligned in the second tilt direction, and the difference between the extension angle of the second tilt direction and the extension angle of the first tilt direction (the extension angle of the second tilt direction minus the extension angle of the first tilt direction) is a+kπ / 2, and the absolute value of a is less than the tilt angle of the first tilt direction. For example, the tilt angle of the first tilt direction is β, that is, |a|<β.

[0163] For ease of explanation, the third tilt direction is introduced here as the second tilt direction when k=0, that is, the difference between the extension angle of the third tilt direction and the extension angle of the first tilt direction is a. In this case, the extension angle of the second tilt direction and the extension angle of the third tilt direction are in a relationship of kπ / 2. It should be understood that when the second tilt direction is the third tilt direction, the third image and the fourth image aligned in the second tilt direction are also referred to as the third image and the fourth image aligned in the third tilt direction, for example, Fig.10 The third tilt direction is the tilt direction Z21, and when the second tilt direction is the third tilt direction, the third image and the fourth image are aligned in the tilt direction Z21.

[0164] According to |a|<β, -β<a<β. Taking the extension angle of the third tilt direction as a2 and the extension angle of the first tilt direction as a1 as an example, -β<a2-a1<β. According to -β<a2-a1<β, a1-β<a2<a1+β. It can be seen that the third tilt direction can be located within the range of β deflected in two opposite directions of the first tilt direction.

[0165] Continue to use Figure 5 The tilt direction Z1 is taken as the first tilt direction, please refer to Fig.10 , Fig.10 The schematic diagram of the position distribution of the first tilt direction and the second tilt direction provided in the embodiment of the present application. In the figure, the extension angle of each tilt direction is marked in the bracket after the number of each tilt direction (only the values ​​within ±360° are marked).

[0166] The tilt direction Z1 in the figure is the first tilt direction, the tilt angle β of the tilt direction Z1 (the angle between the tilt direction Z1 and the positive direction of the x-axis, i.e., the 0° direction) = 30°, and the extension angle of the tilt direction Z1 is 30° (or -330°).

[0167] Tilt direction Z1 rotates clockwise by β to tilt direction Z3, and the extension angle of tilt direction Z3 is the extension angle of tilt direction Z1 + β. Conversely, the difference between the extension angle of tilt direction Z3 and the extension angle of tilt direction Z1 is β. Tilt direction Z1 rotates counterclockwise by β to the positive direction of the x-axis, i.e., the 0° direction. Therefore, the extension angle of the 0° direction is the extension angle of the tilt direction Z1 - β. Conversely, the difference between the extension angle of the 0° direction and the extension angle of the tilt direction Z1 is - β. Fig.10 It can be seen that the extension angle of the tilt direction Z3 is 60° (or -300°), and the extension angle of the 0° direction is 0°, satisfying the aforementioned relationship.

[0168] The tilt direction Z21 in the figure, which is located within the range of β deflected in two opposite directions of the first tilt direction, that is, between the tilt direction Z3 and the 0° direction, is the third tilt direction, and the difference between its extension angle and the extension of the tilt direction Z1 is a=15° (less than 30°). It should be noted that the tilt direction Z21 is only an example. In other embodiments, the tilt direction Z21 can be located in any direction between the tilt direction Z3 and the 0° direction. For example, the tilt direction Z21 can be the direction obtained by rotating the tilt direction Z1 counterclockwise at an angle less than β. In this case, a is negative. For example, the tilt direction Z21 can be the 15° direction, and a=-15°.

[0169] Combine the following Fig.11 The ineffective region areas of the third image and the fourth image when the tilt direction Z21 is located in the tilt direction Z3 and the 0° direction will be described.

[0170] Please refer to Fig.11 , Fig.11 The following is a comparison diagram of the invalid area after the image provided in the embodiment of the present application is rotated clockwise and counterclockwise by the same angle. By comparison, it can be seen that after the same image is rotated clockwise and counterclockwise by the same angle θ, the invalid area is the same, and the invalid area is sin2θ(L1 2 +L2 2 ) / 2, L1 and L2 are the length and width of the image. By sin2θ(L1 2 +L2 2 ) / 2, it can be seen that when the image size remains unchanged, the area of ​​the invalid region is related to the rotation angle. When the rotation angles are the same, the area of ​​the invalid region is the same.

[0171] Based on this, Fig.10In the figure, since the tilt direction Z3 and the 0° direction are the directions after the tilt direction Z1 is rotated clockwise and counterclockwise by the same angle β, respectively, the invalid area area of ​​the third image and the fourth image aligned in the tilt direction Z3 (i.e., the case where the tilt direction Z21 is the tilt direction Z3) is the same as the invalid area area of ​​the third image and the fourth image aligned in the 0° direction (i.e., the case where the tilt direction Z21 is the 0° direction). The third image and the fourth image are aligned in the 0° direction, so the invalid area area of ​​the third image and the fourth image aligned in the tilt direction Z3 is also the invalid area area when the row is aligned.

[0172] Combine the following Fig.12 The ineffective region areas of the third image and the fourth image when the tilt direction Z21 is between the tilt direction Z3 and the 0° direction will be described.

[0173] Please refer to Fig.12 , Fig.12 A sine curve diagram provided for the embodiment of the present application. Fig.12 From the sin2θ sinusoidal curve shown, it can be seen that when 2θ = π / 2 (i.e. θ = π / 4), sin2θ reaches the maximum value, that is, when θ = 45°, the invalid area is the largest. In the range of 0° to 45°, the larger θ is, the larger the invalid area is; in the range of 45° to 90°, the larger θ is, the smaller the invalid area is.

[0174] Based on this, Fig.10 In the figure, the tilt direction Z21 is between the tilt direction Z3 and the 0° direction, that is, the tilt direction Z21 is within the range of the tilt direction Z1 being deflected by β in two opposite directions. In other words, the tilt direction Z21 is obtained by deflecting the tilt direction Z1 by an angle less than β in two opposite directions. Since β≤45°, the tilt direction Z21 is obtained by deflecting the tilt direction Z1 by an angle less than β in two opposite directions within 45°. Obviously, combined with Fig.12 , within 45°, due to the smaller deflection angle, therefore, when the tilt direction Z21 is between the tilt direction Z3 and the 0° direction, the invalid area of ​​the third image and the fourth image is smaller than when the tilt direction Z21 is between the tilt direction Z3 and the 0° direction, that is, the invalid area of ​​the third image and the fourth image when the tilt direction Z21 is between the tilt direction Z3 and the 0° direction, that is, the invalid area when the rows are aligned.

[0175] Therefore, when |a|<β, the third image and the fourth image aligned in the third oblique direction can obtain the benefit of the invalid region area.

[0176] It should be understood that a square image remains a square image after being rotated (counterclockwise or clockwise) an integer multiple of 90°, so that invalid pixels will not be filled, and therefore the area of ​​the invalid region will not change. Based on this, the third image and the fourth image aligned in other directions where the extension angle and the extension angle of the third tilt direction are in a relationship of kπ / 2 have the same invalid region area. Therefore, in an embodiment of the present application, the difference between the extension angle of the second tilt direction and the extension angle of the first tilt direction is a+kπ / 2, and |a| is less than the tilt angle of the first tilt direction, which can make the invalid region area of ​​the third image and the fourth image smaller than the invalid region area when the rows are aligned or the columns are aligned. In this way, when calculating the disparity information, the computational complexity is reduced, thereby improving the computational efficiency of the disparity information, thereby greatly reducing the user experience.

[0177] For example, please refer to Fig.10 , except for the tilt direction Z21, Fig.10 Three tilt directions are also illustrated, namely: tilt direction Z22, tilt direction Z23 and tilt direction Z24, and their extension angles are in a relationship of kπ / 2 with the extension angle of tilt direction Z21. Among them, the extension angle is taken as a positive value for explanation, the tilt direction Z22 with an extension angle of 135° is the second tilt direction when k=1; the tilt direction Z23 with an extension angle of 225° is the second tilt direction when k=2; the tilt direction Z24 with an extension angle of 315° is the second tilt direction when k=3. The extension angle is taken as a negative value for explanation, the tilt direction Z24 with an extension angle of -45° is the second tilt direction when k=-1; the tilt direction Z23 with an extension angle of -135° is the second tilt direction when k=-2; the tilt direction Z22 with an extension angle of -225° is the second tilt direction when k=-3. Based on this, the second tilt direction can be Fig.10 Any one of the tilt directions Z21, Z22, Z23 and Z24 shown.

[0178] It should be noted that k here is obtained based on the extension angle of each tilt direction within ±360°. It should be understood that when the extension angle of each tilt direction is taken as a period of 2kπ, k will have more values.

[0179] The second tilt direction is Fig.10 Taking the tilt direction Z21 as an example, Figure 5 Based on the example of Fig.13 Continue with the explanation.

[0180] Please refer to Fig.13 , Fig.13 A schematic diagram of a third image and a fourth image aligned at 45° provided in an embodiment of the present application.

[0181] in, Fig.13 The left image shown in (a) is the third image, and the right image is the fourth image. Fig.13 The left and right images shown in (a) are overlapped to obtain Fig.13 (b) in . Fig.13 As can be seen from (b) in FIG. 1 , the third image and the fourth image are aligned in the tilt direction Z21, i.e., the direction in which the pixel point M' in the fourth image points to the pixel point M in the third image is the tilt direction Z21. By comparison Fig.13 (b) and Figure 6 From (b) in the figure, it can be found that the area of ​​the invalid region is reduced.

[0182] The above content illustrates an example of the second tilt direction in S802. The following describes an implementation of correcting the first image and the second image to obtain the third image and the fourth image.

[0183] In some embodiments of the present application, the above S802 specifically includes: rotating the first image and the second image by a+kπ / 2 to obtain a third image and a fourth image.

[0184] Ideally, the first camera and the second camera are arranged in the first tilted direction, and the first image and the second image are also aligned in the first tilted direction. The extension angle of the alignment direction of the first image and the second image before and after correction differs by a+kπ / 2. Therefore, the third image and the fourth image can be obtained by rotating the first image and the second image by the difference angle.

[0185] It should be noted that the difference a+kπ / 2 between the extension angle of the second oblique direction and the extension angle of the first oblique direction can be positive or negative. If a+kπ / 2 is positive, it means that the extension angle gradually increases from the first oblique direction to the second oblique direction, and the rotation from the first oblique direction to the second oblique direction requires a clockwise rotation of |a+kπ / 2|; if a+kπ / 2 is negative, it means that the extension angle gradually decreases from the first oblique direction to the second oblique direction, and the rotation from the first oblique direction to the second oblique direction requires a counterclockwise rotation of |a+kπ / 2|. In other words, the direction of rotation of the first image and the second image is determined based on the positive or negative value of a+kπ / 2.

[0186] Please continue to refer to Figure 5, ideally, the first image and the second image are also aligned in the tilt direction Z1, that is, the direction in which the pixel point M' in the second image points to the pixel point M in the first image is the tilt direction Z1. In other words, the alignment direction of the first image and the second image is the tilt direction Z1, and the alignment direction of the third image and the fourth image is the tilt direction Z21. It can be seen that the difference a+kπ / 2 between the extension angles of the alignment directions of the first image and the second image before and after correction exists in the following cases:

[0187] When the extension angle of the tilt direction Z21 is 45° and the extension angle of the tilt direction Z1 is 30°, the difference in the extension angles a+kπ / 2=15° (i.e., k=0); in this case, by Figure 5 The first image and the second image shown in (b) are rotated 15° clockwise to obtain Fig.13 The third image and the fourth image shown in (b) .

[0188] When the extension angle of the tilt direction Z21 is -315° and the extension angle of the tilt direction Z1 is -330°, the difference in the extension angles is a+kπ / 2=15° (i.e., k=0); in this case, by Figure 5 The first image and the second image shown in (b) are rotated 15° clockwise to obtain Fig.13 The third image and the fourth image shown in (b) .

[0189] When the extension angle of the tilt direction Z21 is 45° and the extension angle of the tilt direction Z1 is -330°, the difference in the extension angles is a+kπ / 2=375° (i.e., k=4); in this case, by Figure 5 The first image and the second image shown in (b) are rotated 375° clockwise to obtain Fig.13 The third image and the fourth image shown in (b) .

[0190] When the extension angle of the tilt direction Z21 is -315° and the extension angle of the tilt direction Z1 is 30°, the difference in the extension angles a+kπ / 2=-345° (i.e., k=-4); in this case, by Figure 5 The first image and the second image shown in (b) are rotated 345° counterclockwise to obtain Fig.13 The third image and the fourth image shown in (b) .

[0191] It should be noted that the above implementation is based on the situation that the first image and the second image are aligned in the first oblique direction under ideal conditions. However, in the specific implementation process, due to various factors, the first image and the second image are distorted, so that they cannot be aligned in the first oblique direction, or even cannot be coplanar. Simply rotating the first image and the second image cannot obtain the third image and the fourth image aligned in the second oblique direction, and more complex processing is required.

[0192] Based on this, in other embodiments of the present application, please refer to Fig.14 , Fig.14 A schematic diagram of a method for image processing provided in an embodiment of the present application Figure 2 . Fig.14 right Figure 8 The S802 in the embodiment is refined. The S802 specifically includes the following S802a and S802b.

[0193] S802a, performing Bouguet stereo correction on the first image and the second image, so that the first image and the second image after the Bouguet stereo correction are aligned in a first direction.

[0194] The first direction is a row direction or a column direction, and a difference between an extension angle of the first direction and an extension angle of the first inclined direction is γ.

[0195] Please continue to refer to Figure 6 , Figure 6 The first image and the second image are subjected to Bouguet stereo correction, so that the first image and the second image after Bouguet stereo correction are aligned in the positive direction of the x-axis, that is, the extension angle of the first direction is 0°, and the extension angle of the tilt direction Z1 is 30° as an example, then γ=-30°, and the Bouguet stereo correction is rotated in the counterclockwise direction. Of course, in other embodiments, Bouguet stereo correction can also be used to achieve column alignment, that is, the first image and the second image after Bouguet stereo correction are aligned in the y-axis direction. Taking the alignment of the first image and the second image after Bouguet stereo correction in the positive direction of the y-axis as an example, the extension angle of the first direction is 90°, and the extension angle of the tilt direction Z1 is 30° as an example, then γ=60°, and the Bouguet stereo correction is rotated in the clockwise direction. It can be seen that the above γ can be positive or negative, depending on the direction of rotation to achieve row alignment or column alignment.

[0196] S802b, rotating the first image and the second image after Bouguet stereo correction by a-γ+kπ / 2 to obtain a third image and a fourth image.

[0197] Bouguet stereo correction corrects the first image and the second image to be aligned in the first direction. Taking the extension angle of the first direction as a3, the extension angle of the first oblique direction as a1, and the extension angle of the second oblique direction as a2 as an example, a2-a1=a+kπ / 2, the difference between the extension angle of the first direction and the extension angle of the first oblique direction is γ, that is, a3-a1=γ, then the extension angle of the second oblique direction and the extension angle of the first direction are a2-a3=a-γ+kπ / 2. Therefore, by rotating the first image and the second image after Bouguet stereo correction by a-γ+kπ / 2, the above-mentioned third image and fourth image can be obtained. It should be noted that the rotation direction of the first image and the second image after Bouguet stereo correction is determined based on the positive and negative of a-γ+kπ / 2.

[0198] for Figure 6 For the first and second images after Bouguet stereo correction shown, when k=0, the difference between the extension angle of the tilt direction Z21 and the extension angle of the 0° direction is a-γ+kπ / 2=45°, where a=15°, γ=-30°, and k=0.

[0199] It should be noted that the currently available Bouguet stereo correction technology is used to achieve row alignment or column alignment. This embodiment is based on the currently available Bouguet stereo correction technology. Bouguet stereo correction is first performed on the first image and the second image so that the first image and the second image after Bouguet stereo correction are aligned in row or column. Then, the first image and the second image after Bouguet stereo correction are further rotated by a-γ+kπ / 2 to obtain the third image and the fourth image aligned in the second oblique direction. In other words, this embodiment only needs to rotate a-γ+kπ / 2 on the basis of Bouguet stereo correction to obtain the third image and the fourth image aligned in the second oblique direction, without the need to additionally design a new Bouguet stereo correction technology to achieve alignment in the oblique direction.

[0200] For ease of understanding, the following Fig.15 The complete calibration process is described below. Fig.15 , Fig.15 This is a schematic diagram of a process of correcting a first image and a second image to obtain a third image and a fourth image in an embodiment of the present application.

[0201] Fig.15 The left and right images shown in (a) are the first image taken by camera B1 and the second image taken by camera B2, that is, Figure 5 Next, by Fig.15The left and right images shown in (a) are subjected to Bouguet stereo correction, so that the first image and the second image after Bouguet stereo correction are aligned in the positive direction of the x-axis, showing Fig.15 The states of the left and right figures shown in (b) are shown in FIG. Fig.15 The left and right images shown in (b) are rotated 45° clockwise to obtain Fig.15 The left and right images shown in (c) above are: the left image is the third image, and the right image is the fourth image. Fig.15 The left and right images shown in (c) are Fig.13 The left and right figures shown in (a) are Fig.13 As can be seen in (b), the third image and the fourth image are aligned at 45°.

[0202] S803: Match the fourth image with the third image along the second oblique direction to determine disparity information between feature points of the third image and corresponding points of the fourth image.

[0203] It should be noted that the feature point generally refers to a pixel point in the third image, and the corresponding point is the pixel point in the fourth image that best matches the feature point of the third image. The best matching pixel point refers to the pixel point with the highest matching degree among all the pixels within a certain search range. The matching degree can be determined by the grayscale value and position information of the two matching pixels. The relevant image processing technology contains technical content on how to determine the matching degree of pixels and what is the pixel point with the highest matching degree, which will not be described in detail here.

[0204] Matching the fourth image with the third image along the second inclined direction refers to the process of determining the degree of matching between each pixel point of the fourth image within a search range (set as needed) in the second inclined direction (the second inclined direction passing through the pixel points with the same coordinates as the feature points) and the feature points of the third image, and determining the best matching pixel point, i.e., the corresponding point, based on each matching degree.

[0205] For example, please refer to Fig.16 , Fig.16 A schematic diagram of the principle of matching a third image and a fourth image in a second inclined direction of the fourth image provided in an embodiment of the present application.

[0206] Fig.16 In the example, the left image is the third image, the right image is the fourth image, and the second tilt direction is the 45° direction. For a feature point P in the left image, the 45° direction in the right image (the 45° direction passing through the pixel point with the same coordinates as the feature point) is within the search range. The matching degree between each pixel point in the right image and the feature point P is determined to determine the corresponding point. Specifically, the search range is 45° in the right image. The pixel points in the image are pixel Q1, pixel Q2, pixel Q3, etc. (the search range in the figure is The enclosed pixels include pixel points Q1, pixel point Q2, pixel point Q3, and the pixels indicated by ellipsis in the search range. After matching pixel point Q1, jump to the next pixel point in the 45° direction (i.e., pixel point Q2) for matching to determine the matching degree between feature point P and pixel point Q2; after matching pixel point Q2, jump to the next pixel point in the 45° direction (i.e., pixel point Q3) for matching to determine the matching degree between feature point P and pixel point Q3; and so on until the search range is completed. The pixel point with the highest matching degree is obtained by matching all the pixel points in the image, and the fourth image is matched with the third image along the second inclined direction.

[0207] In some embodiments of the present application, when the inclination angle of the first inclination direction is greater than 22.5°, the extension angle of the second inclination direction is 45°+kπ / 2.

[0208] In this embodiment, the extension angle of the second oblique direction is 45°+kπ / 2, such as 45°, 135°(-225°), 225°(-135°), 315°(-45°). When the extension angle of the second oblique direction is 45°+kπ / 2, the angles between the second oblique direction and the two axes of the pixel coordinate system are both 45°. Fig.16 , the pixels of each phase in this direction are aligned diagonally, and when matching the third pixel and the fourth image in the second inclined direction of the fourth image, each time a pixel is matched, the next pixel is jumped to match, and the next pixel is unique. It should be understood that when the next pixel is not unique, more computing power is required to match multiple pixels, resulting in low computing efficiency. Based on this, this embodiment can improve computing efficiency by correcting the first image and the second image into the third image and the fourth image aligned in the direction of 45°+kπ / 2.

[0209] It should be noted that, when the inclination angle of the first inclination direction is 22.5°, the extension angle of the first inclination direction is one of 22.5°+kπ / 2 and 62.5°+kπ / 2. Taking 22.5°+kπ / 2 as an example, the difference between 45°+kπ / 2 and 22.5°+kπ / 2, that is, a+kπ / 2=22.5°+kπ / 2. It can be seen that a is just equal to the inclination angle of the first inclination direction (22.5), which does not meet the conditions for obtaining the invalid area area gain. Based on this, in this embodiment, for the case where the inclination angle of the first inclination direction is greater than 22.5°, the first image and the second image can be corrected to the third image and the fourth image aligned in the direction of the extension angle of 45°+kπ / 2 to obtain the invalid area area gain.

[0210] The above-mentioned disparity information between the feature point and the corresponding point (also referred to as the disparity information corresponding to the feature point) refers to the position difference between the feature point and the corresponding point. Since the third image and the fourth image are aligned in the second oblique direction in the embodiment of the present application, the position difference between the feature point and the corresponding point is the position difference between the feature point and the corresponding point in the second oblique direction. In other words, the disparity information between the feature point and the corresponding point is the position difference between the feature point and the corresponding point in the second oblique direction. Fig.13 As shown in (b), the feature point M and the corresponding point M' are the best match, and the disparity information between the feature point M and the corresponding point M' is Δx.

[0211] The above content is explained by taking a pair of matching feature points and corresponding points as an example. It should be understood that there are multiple pixel points in the third image, that is, there are multiple feature points in the third image. Different feature points have different corresponding points, and there is a one-to-one correspondence between feature points and corresponding points. It can be seen that there are multiple pairs of matching feature points and corresponding points in the third image and the fourth image. A pair of matching feature points and corresponding points is called a pair of matching point pairs, and there are multiple pairs of matching point pairs in the third image and the fourth image.

[0212] Based on this, during a specific implementation, the electronic device may determine the disparity information of multiple pairs of matching points between the third image and the fourth image during the process of executing S803.

[0213] Optionally, the electronic device may store the disparity information of each pair of matching point pairs determined by the electronic device in the form of a disparity map. Specifically, the size and resolution of the disparity map are the same as those of the third image, and the pixel value of a disparity point (a pixel point, which is referred to as a disparity point for ease of distinction) in the disparity map is the disparity information corresponding to the feature point in the third image with the same coordinates as the disparity point.

[0214] In S803, since the third image and the fourth image are aligned in the second oblique direction, according to the epipolar constraint principle, the corresponding point of the fourth image will appear in the second oblique direction of the pixel point (referred to as the same coordinate point) passing through the feature point. Based on this, in this embodiment, the electronic device matches the fourth image with the third image in the second oblique direction of the fourth image, and can match and search for corresponding points in the second oblique direction of the fourth image, so that parallax information can be obtained based on the matched feature points and corresponding points. It should be understood that the matching search in the second oblique direction of the fourth image belongs to a one-dimensional search, which can reduce the matching search range.

[0215] S804: Determine depth information of the first image according to the disparity information.

[0216] The depth information of the first image includes the depth information of each feature point in the first image. In the specific implementation process, the electronic device may store the depth information of each feature point determined by the electronic device in the form of a depth map. Specifically, the size and resolution of the depth map are the same as those of the first image, and the pixel value of a depth point (a pixel point, which is referred to as a depth point for ease of distinction) in the depth map is the depth information corresponding to the pixel point with the same coordinates as the depth point in the first image.

[0217] In a specific implementation process, S804 may include: first, according to the disparity information of each matching point pair between the third image and the fourth image, the depth information of each feature point in the third image is obtained. For example, the electronic device may determine the depth information of each feature point in the third image based on the aforementioned relationship z=b*f / d. It should be noted that the depth information of each feature point in the third image may also be presented in the form of a depth map. The depth map of the first image is obtained by rotating the depth map corresponding to the third image to the orientation of the first image.

[0218] It should be understood that when the calculation amount and complexity of the disparity information are reduced and the calculation efficiency is improved, the efficiency of obtaining the depth information of the first image is also improved accordingly. This allows the electronic device to reduce time and energy consumption when determining the depth information of the first image, thereby improving the user experience.

[0219] As a possible implementation, the above S803 may include: Fig.14 Steps S803a and S803b in the embodiment are as follows:

[0220] S803a: Input the third image and the fourth image into a disparity model, so that the disparity model outputs disparity information.

[0221] The disparity model is used to match the fourth image with the third image along the second oblique direction and determine disparity information.

[0222] S803b, obtaining disparity information output by the disparity model.

[0223] According to the above content, the corresponding point of the fourth image will appear in the second oblique direction passing through the same coordinate point on the fourth image. When the disparity model is used to determine the disparity information, the disparity model should have the ability to match the third image and the fourth image in the second oblique direction passing through the same coordinate point in the fourth image. In this way, the disparity model matches the fourth image with the third image along the second oblique direction based on this ability, and can match and search for corresponding points, so that the disparity information of the feature points of the third image and the corresponding points of the fourth image in the second oblique direction can be effectively calculated.

[0224] In some possible implementations, the disparity model matches the fourth image with the third image along the second oblique direction and determines disparity information, including Fig.17 Steps S1701 to S1703 in .

[0225] For example, please refer to Fig.17 , Fig.17 The example flowchart of determining disparity information by using a disparity model provided in an embodiment of the present application includes the following steps S1701 to S1703.

[0226] S1701, process the fourth image using a convolution kernel so that the fourth image is shifted along the second tilt direction by 1 step to N steps in sequence, and obtain N feature maps respectively.

[0227] Among them, N is a positive integer greater than 1, and N is a preset search range, which can be set as needed to find the best matching corresponding point for each feature point in the third image in the second inclined direction of the fourth image.

[0228] The step length of the fourth image shifting one step in sequence along the second tilt direction can be set as needed. The larger the step length, the lower the accuracy of the matching search; the smaller the step length, the greater the calculation amount of the matching search.

[0229] Specifically, the disparity model uses a convolution kernel as a sliding window to slide on the fourth image for convolution. When the fourth image is slid, a convolution of the fourth image is completed, so that the fourth image is offset by 1 step along the second oblique direction to obtain the first feature map; the disparity model continues to use the convolution kernel as a sliding window to slide on the first feature map for convolution. When the first feature map is slid, a convolution of the first feature map is completed (that is, two convolutions of the fourth image are completed), so that the first feature map is offset by 1 step along the second oblique direction (the fourth image is offset by 2 steps along the second oblique direction), and the second feature map is obtained; ..., and so on; the disparity model uses the convolution kernel as a sliding window on the n-1th feature map The convolution is performed by sliding. When the n-1th feature map is slid, a convolution of the n-1th feature map is completed (that is, n convolutions of the fourth image are completed), so that the n-1th feature map is offset by 1 step along the second oblique direction (the fourth image is offset by n steps along the second oblique direction), and the nth feature map is obtained; ..., and so on; the disparity model uses the convolution kernel as a sliding window to slide on the N-1th feature map for convolution. When the N-1th feature map is slid, a convolution of the N-1th feature map is completed (that is, N convolutions of the fourth image are completed), so that the N-1th feature map is offset by 1 step along the second oblique direction (the fourth image is offset by N steps along the second oblique direction), and the Nth feature map is obtained.

[0230] The result of the fourth image shifting one step along the second inclined direction is that the row pixels of some rows and the column pixels of some columns of the fourth image located at the front end in the second inclined direction are removed, and the row pixels of some rows and the column pixels of some rows located at the end in the second inclined direction are filled with invalid pixels, for example, the pixel value of the invalid pixel is zero.

[0231] Please refer to Fig.18 , Fig.18 It shows a before and after comparison diagram of the fourth image shifted by one step in the 45° direction. Fig.18 (a) is the fourth image, i.e., no offset occurs; Fig.18 (b) in the figure is the image with the fourth image shifted by one step, i.e., the first feature map. Fig.18 The image shown in (b) is framed by the bold image frame in the figure. Fig.18 The fourth image is offset by one step Pixel distance, pixel distance refers to the size of a pixel. By comparison, it can be found that after the fourth image is shifted by one step in the 45° direction, 1 row of pixels and 1 column of pixels at the end of the 45° direction are filled with invalid pixels, that is, 1 row of pixels close to the x-axis and 1 column of pixels close to the y-axis are filled with invalid pixels, as indicated by darker shadows in the figure; and 1 row of pixels and 1 column of pixels at the front end of the 45° direction are removed, that is, 1 row of pixels away from the x-axis and 1 column of pixels away from the y-axis are removed, that is, the pixels outside the bold image frame in the figure, these pixels do not belong to the first feature map, and these pixels do not exist in the actual implementation process. The image obtained by shifting the fourth image by one step in the 45° direction is the part within the bold image frame, that is Fig.18 The image shown in (c) is Fig.18 The shaded area in (c) is the area filled with invalid pixels.

[0232] In some embodiments of the present application, the convolution kernel is a matrix T (2m+1)×(2m+1) ; T (2m+1)×(2m+1) The element t ij = 1, T (2m+1)×(2m+1) Except t ij The rest of the elements are zero; among them, t ij Point to t mm The direction of the second tilt direction is the same as that of the second tilt direction, t ij T (2m+1)×(2m+1) The element in the i-th row and j-th column in t mm T (2m+1)×(2m+1) The element in the mth row and mth column in T (2m+1)×(2m+1) The central element, m is a positive integer, 1≤i≤2m+1, 1≤j≤2m+1, i, j, m are all positive integers.

[0233] In this embodiment, a convolution kernel is provided that can shift the fourth image in the second oblique direction. When the convolution kernel is used as a sliding window to slide on the image, since only t ij =1, and the rest are zero. Therefore, each time we slide to an area on the fourth image, we will ij The corresponding pixel value is filled to t mm The position of the corresponding pixel point. ij Point to t mm The direction of the second tilt is the same as that of the second tilt, so t ij The corresponding pixel value is filled to t mm The position of the corresponding pixel point is equivalent to t ijThe value of the corresponding pixel point is shifted by one step in the second oblique direction. By sliding the convolution kernel on the fourth image, the value of each pixel point in the fourth image is shifted by one step in the second oblique direction, which is equivalent to shifting the fourth image by one step in the second oblique direction.

[0234] It should be noted that the convolution kernel is used to convolve the fourth image with a step length of one step. The result of the fourth image shifting one step along the second oblique direction is that |mi| rows of pixels and |mj| columns of pixels at the front end of the fourth image in the second oblique direction are removed, and |mi| rows of pixels and |mj| columns of pixels at the end of the fourth image in the second oblique direction are filled with invalid pixels.

[0235] For example, when the extension angle of the second tilt direction is 45°+2kπ, the convolution kernel is In this embodiment, t 11 Point to t 22 The direction is 45°. The fourth image is convolved by the convolution kernel. Each time the convolution is performed, the fourth image can be shifted by one step in the 45° direction. The step length of the fourth image shifting by one step in the 45° direction is The result of the fourth image being shifted by one step in the 45° direction is that one row of pixels and one column of pixels at the front end in the 45° direction of the fourth image are removed, and one row of pixels and one column of pixels at the end in the second oblique direction are filled with invalid pixels.

[0236] For example, please refer to Fig.19 , Fig.19 A schematic diagram of a fourth image provided in an embodiment of the present application being shifted by n steps in a second tilt direction. In the figure, the second tilt direction is a 45° direction. Fig.19 The convolution kernel is used as Convolve the fourth image n times, so that the fourth image is offset n steps in the 45° direction, and the offset distance is n*d. pixel distance, and finally get the nth feature map.

[0237] For another example, when the extension angle of the second tilt direction is 135°+2kπ, the convolution kernel is In this embodiment, t 13 Point to t 22 The direction is 135°.

[0238] S1702, determining the correlation between each feature map and each pixel pair of the third image, and obtaining N correlation maps.

[0239] The pixel pair refers to two pixels with the same coordinates in the feature map and the third image. For the nth feature map, the nth correlation map is obtained by determining the correlation of each pixel pair between the nth feature map and the third image.

[0240] During the specific implementation process, the correlation of each pixel pair between the nth feature map and the third image is calculated to determine the correlation of each pixel pair, that is, the aforementioned matching degree. At present, there are a large number of technologies on how to calculate the correlation between pixel pairs, which will not be described in detail here. It should be noted that the nth correlation map is an image with the same size and resolution as the third image, and the pixel value of a matching point in the nth correlation map (a pixel point in the nth correlation map, for the sake of convenience, it is called a matching point) is the correlation of the pixel pair with the same coordinates as the matching point. For example, the pixel value of the pixel point (x1, y1) of the nth correlation map is the correlation between the pixel point (x1, y1) of the third image and the pixel point (x1, y1) of the nth feature map.

[0241] It should be noted that in S1701, the fourth image is shifted in the second oblique direction by the convolution kernel, which is equivalent to moving each pixel on the fourth image in the second oblique direction. Fig. 20 In the figure, the pixel point M' moves from the upper left position in the figure along the 45° direction to the lower right position in the figure. By moving each pixel point on the fourth image in the second oblique direction, each pixel point in the search range in the second oblique direction of the fourth image is sequentially moved to the position with the same coordinates as the feature point, and the matching degree between each pixel point in the search range in the second oblique direction of the fourth image and the feature point of the third image is determined by S1702, so as to achieve the matching of the fourth image with the third image. S1701 and S1702 are equivalent to the above Fig.16 The process of matching one pixel and jumping to the next pixel for matching.

[0242] S1703: Determine disparity information based on the N correlation graphs.

[0243] Among them, the disparity information is n*d, d is the step size of the fourth image offset by one step, n is the serial number of the nth correlation map, the nth correlation map is the correlation map with the highest pixel value among the pixel points with the same coordinates as the feature point in the N correlation maps (referred to as the same coordinate points for short), and n is a positive integer less than or equal to N.

[0244] This embodiment provides a solution for matching the fourth image with the third image along the second oblique direction using a disparity model and determining disparity information. It should be understood that when the nth correlation graph is the correlation graph with the highest pixel value of the same coordinate point in the N correlation graphs, it means that the same coordinate point in the nth correlation graph is the most matched with the feature point of the third image. From the above content, it can be seen that the same coordinate point in the nth correlation graph is a pixel point in the fourth image in the second oblique direction passing through the same coordinate point offset by n*d, and this pixel point is the corresponding point. Therefore, the disparity information to be optimized between the corresponding point and the feature point can be obtained as n*d.

[0245] It should be noted that in the specific implementation process, Fig.17 The process of determining disparity information by the disparity model in may involve more steps. Fig. 20 Provide explanation.

[0246] For example, please refer to Fig. 20 , Fig. 20 An example flow chart of outputting disparity information based on the third image and the fourth image by the disparity model provided in the embodiment of the present application includes the following steps S2001 to S2004. The specific steps are as follows:

[0247] S2001, performing feature extraction on the third image and the fourth image to obtain a fifth image corresponding to the third image and a sixth image corresponding to the fourth image.

[0248] The third image and the fourth image are aligned in the second oblique direction. For example, the second oblique direction is a 45° direction, that is, the extension angle of the second oblique direction is 45°.

[0249] It should be understood that the fifth image is the third image after feature extraction, and the sixth image is the fourth image after feature extraction. It should be noted that there are related technologies in the prior art that describe the feature dimensions (such as color, texture, grayscale value, etc.) to be extracted, and no further description is given here.

[0250] S2002: Match the sixth image with the fifth image along a second oblique direction to determine the parallax information to be optimized between the feature points of the fifth image and the corresponding points of the sixth image.

[0251] The specific implementation of S2002 can be achieved through Fig.17 It should be understood that in this case, Fig.17 The third image and the fourth image involved in each step are adjusted to the fifth image and the sixth image accordingly. Fig.19 The fourth image in is substantially a sixth image obtained by feature extraction from the fourth image, such as a grayscale image.

[0252] S2003 , performing disparity optimization on the disparity information to be optimized, and obtaining disparity information between feature points of the fifth image and corresponding points of the sixth image.

[0253] Parallax optimization mainly involves improving accuracy, eliminating errors, optimizing weak texture areas and filling holes. It is a step in binocular stereo matching technology. You can refer to the current relevant description of binocular stereo matching technology, and no further description is given here.

[0254] S2004, outputting disparity information between the feature points of the fifth image and the corresponding points of the sixth image.

[0255] It should be understood that the disparity information between the feature points of the fifth image and the corresponding points of the sixth image is the disparity information between the feature points of the third image and the corresponding points of the fourth image, so that the output of the disparity information is achieved through the disparity model.

[0256] It should be understood that the above Fig.14 Before S803a in the above, the initial model needs to be trained to obtain the above disparity model. For example, please refer to Fig.21 , Fig.21 A flow chart of a method for training a disparity model provided in an embodiment of the present application. The method for training a disparity model includes the following steps S2101 to S2102:

[0257] S2101, obtain a training sample set.

[0258] The training sample set includes multiple groups of training samples, a single group of training samples includes a first training image and a second training image, and the first training image and the second training image are aligned in the second tilt direction.

[0259] In some embodiments, for the process of training using supervised learning, the single set of training samples includes, in addition to the first training image and the second training image, a disparity label, that is, disparity information of each matching pair between the first training image and the second training image. For example, a disparity map. Of course, for the process of unsupervised learning training, the single set of training samples may not include a disparity label.

[0260] S2102, training the initial model based on the training sample set to obtain a disparity model.

[0261] The initial model is used to match the second training image with the first training image along the second tilt direction, and determine the disparity information between the feature points of the first training image and the corresponding points of the second training image.

[0262] The initial model is used to match the second training image with the first training image along the second tilt direction. The specific content of determining the disparity information between the feature points of the first training image and the corresponding points of the second training image can be adaptively referred to the relevant content of S803, which will not be repeated here.

[0263] Fig.21 In the training method shown, the initial model is constrained to match the first training image and the second training image in the second tilt direction of the second training image, so that the disparity model has the ability to match the third image and the fourth image in the second tilt direction of the fourth image. Fig.21 In the example, the first training image and the second training image aligned in the second oblique direction are used as inputs of the initial model. The initial model is trained and the parameters of the initial model are continuously adjusted, so that the accuracy of the initial model in matching search and determining disparity information in the second oblique direction is improved. The adjusted parameters can be understood as Fig. 20 Parameters for feature extraction, disparity calculation, disparity optimization, and so on.

[0264] In some embodiments, the training sample set can be obtained by rotating and cropping an existing publicly available “row-aligned” or “column-aligned” data set, thus eliminating the need to create a new training sample set.

[0265] For example, please refer to Fig. 22 , Fig. 22 A schematic diagram of a process for obtaining a training sample set provided in an embodiment of the present application.

[0266] S2201, obtaining an original training sample set, where a single set of original training samples includes a first original image and a second original image, and the first original image and the second original image are aligned in rows or columns.

[0267] It should be noted that the current disparity model is trained based on row-aligned or column-aligned binocular images. Therefore, there are currently a large number of publicly available data sets formed by row-aligned or column-aligned binocular images, which can be used as the original training sample set here.

[0268] It should be understood that if the training samples include disparity labels, each group of original training samples also includes original disparity labels, such as original disparity maps.

[0269] S2202: Rotate the first original image and the second original image included in each group of original training samples in the same direction to obtain multiple groups of rotated training samples.

[0270] Each group of rotation training samples includes a first rotated image and a second rotated image, the first rotated image is an image obtained by rotating the first original image, the second rotated image is an image obtained by rotating the second original image, and the first rotated image and the second rotated image are aligned in the second tilt direction. Based on this, the angle and direction of rotation are based on the alignment of the first rotated image and the second rotated image in the second tilt direction.

[0271] It should be understood that if each group of original training samples also includes an original disparity map, then it is also necessary to rotate the original disparity map included in each group of original training samples in the same direction to obtain a rotated disparity map corresponding to each group of first rotated images and second rotated images. In this case, the rotated training samples also include a rotated disparity map.

[0272] S2203: Perform cropping in the effective area of ​​the first rotated image and the second rotated image included in each group of rotated training samples to obtain a training sample set including multiple groups of training samples.

[0273] The single group of training samples includes a first training image and a second training image, wherein the first training image is an image cropped from the first rotated image, and the second training image is an image cropped from the second rotated image, and the first training image and the second training image are also aligned in the second tilt direction.

[0274] It should be understood that if each set of rotation training samples also includes a rotation disparity map, the rotation disparity map included in each set of rotation training samples needs to be cropped to obtain a disparity map corresponding to each set of first training images and second training images. In this case, the training samples also include a disparity map.

[0275] For example, please refer to Fig.23 , Fig.23 A schematic diagram of the processing process of the original training samples provided in an embodiment of the present application.

[0276] Fig.23 The second tilt direction is 45°. Fig.23 The left and right images shown in (a) are the first original image and the second original image, respectively. The first original image and the second original image are aligned in rows. Fig.23 The first original image and the second original image shown in (a) are rotated 45° clockwise to obtain Fig.23 The left and right images shown in (b) are the first rotated image and the right image is the second rotated image. Fig.23 The first rotated image and the second rotated image shown in (b) are cropped at the same position in the effective area to obtain Fig.23In the left and right images shown in (c), the left image is the first training image, and the right image is the second training image. The first training image and the second training image are aligned at 45°.

[0277] Please refer to Fig.24 In some embodiments of the present application, after S804, the method further includes:

[0278] S805, blurring the first image according to the depth information of the first image to obtain a blurred image. It should be noted that there are a large number of algorithms related to image blurring in the relevant image processing technology, which will not be introduced in detail here.

[0279] This embodiment provides an application scenario of depth information, namely, an image blurring scenario. In the image blurring scenario, blurring the first image depends on the depth information of the first image. Based on this, when the amount of calculation and the complexity of calculation of the parallax information are reduced and the calculation efficiency is improved, the efficiency of obtaining the depth information of the first image is also improved accordingly. Naturally, the efficiency of blurring the first image to obtain the blurred image is improved.

[0280] It should be understood that the application scenarios of depth information are not limited to image blurring scenarios. In other application scenarios involving depth information, if the method provided in the embodiment of the present application is used to obtain depth information, it should be considered to be within the protection scope of the present application.

[0281] The embodiment of the present application also provides an electronic device. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the method described in any of the above embodiments when executing the computer program, and the method may be a method for image processing provided in any of the above embodiments.

[0282] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any of the above embodiments is implemented, and the method can be a method for image processing provided by any of the above embodiments.

[0283] The embodiment of the present application also provides a chip system. The chip system includes a processor and a memory; wherein the processor is coupled to the memory, and the memory is used to store programs or instructions. When the program or instruction is executed by the processor, the chip system implements the method described in any of the above embodiments, which may be the method for image processing provided in any of the above embodiments.

Claims

1. A method for image processing, characterized in that Applied in an electronic device, the electronic device comprises at least a first camera and a second camera, the first camera and the second camera are arranged in a first inclined direction, and the method comprises: When photographing the same photographing target, acquiring a first image based on the first camera and acquiring a second image based on the second camera; Correcting the first image and the second image to obtain a third image corresponding to the first image and a fourth image corresponding to the second image; wherein the third image and the fourth image are aligned in a second tilt direction; the difference between the extension angle of the second tilt direction and the extension angle of the first tilt direction is a+kπ / 2, where the absolute value of a is smaller than the tilt angle of the first tilt direction, and k is an integer; the extension angle of the tilt direction is an angle required to rotate with the 0° direction as the starting side and the tilt direction as the ending side in a plane rectangular coordinate system; matching the fourth image with the third image along the second tilt direction, and determining disparity information between feature points of the third image and corresponding points of the fourth image; Depth information of the first image is determined according to the disparity information.

2. The method according to claim 1, characterized in that: The correcting the first image and the second image to obtain a third image corresponding to the first image and a fourth image corresponding to the second image includes: The first image and the second image are respectively rotated by a+kπ / 2 to obtain the third image and the fourth image.

3. The method according to claim 1, characterized in that The correcting the first image and the second image to obtain a third image corresponding to the first image and a fourth image corresponding to the second image includes: Performing Bouguet stereoscopic correction on the first image and the second image so that the first image after Bouguet stereoscopic correction and the second image after Bouguet stereoscopic correction are aligned in a first direction; wherein the first direction is a row direction or a column direction, and the difference between an extension angle of the first oblique direction and an extension angle of the first direction is γ; The first image after Bouguet stereo correction and the second image after Bouguet stereo correction are rotated by a-γ+kπ / 2 to obtain the third image and the fourth image.

4. The method according to any one of claims 1 to 3, characterized in that: When the inclination angle of the first inclination direction is greater than 22.5°, the extension angle of the second inclination direction is 45°+kπ / 2.

5. The method according to any one of claims 1 to 3, characterized in that: The matching the fourth image with the third image along the second tilt direction to determine the disparity information between the feature points of the third image and the corresponding points of the fourth image includes: The third image and the fourth image are input into a disparity model to obtain the disparity information output by the disparity model; wherein the disparity model is used to match the fourth image with the third image along the second oblique direction and determine the disparity information.

6. The method according to claim 5, characterized in that The matching the fourth image with the third image along the second oblique direction and determining the disparity information includes: The fourth image is processed by using a convolution kernel, so that the fourth image is shifted by 1 to N steps in sequence along the second tilt direction, and N feature maps are obtained respectively; N is a positive integer greater than 1; Determine the correlation between each pixel pair of each feature map and the third image, and obtain N correlation maps; Based on N correlation graphs, the disparity information is determined; wherein the disparity information is n*d, d is the step size of one step offset of the fourth image, n is the serial number of the nth correlation graph, the nth correlation graph is the correlation graph with the highest pixel value among the pixel points with the same coordinates as the feature point in the N correlation graphs, and n is a positive integer less than or equal to N.

7. The method according to claim 6, characterized in that The convolution kernel is the matrix T (2m+1)×(2m+1) ; T (2m+1)×(2m+1) Meta ij = 1, T (2m+1)×(2m+1) Except t ij The rest of the elements are zero; among them, t ij Point to t mm The direction is the same as the second tilt direction, t ij T (2m+1)×(2m+1) The element in the i-th row and j-th column in t mm T (2m+1)×(2m+1) For the element in the mth row and mth column, 1≤i≤2m+1, 1≤j≤2m+1, where i, j, and m are all positive integers.

8. The method according to claim 7, characterized in that When the extension angle of the second tilt direction is 45°+2kπ, the convolution kernel is 9. The method according to any one of claims 6 to 8, characterized in that: Before inputting the third image and the fourth image into a disparity model to obtain the disparity information output by the disparity model, the method further includes: Acquire a training sample set, the training sample set comprising a plurality of groups of training samples, a single group of the training samples comprising a first training image and a second training image, the first training image and the second training image being aligned in the second tilt direction; The initial model is trained based on the training sample set to obtain the disparity model; wherein the initial model is used to match the second training image with the first training image along the second tilt direction, and determine the disparity information between the feature points of the first training image and the corresponding points of the second training image.

10. The method according to any one of claims 6 to 8, characterized in that: After determining the depth information of the first image, the method further includes: The first image is blurred according to the depth information to obtain a blurred image.

11. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 10 when executing the computer program.

12. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.

Citation Information

Patent Citations

  • Parallax-graph generation system and method and storage medium

    CN108230235A

  • 3D shooting device and method applied to mobile terminals

    CN108616741A