Image depth estimation method, electronic equipment and computer readable storage medium

By combining binocular and monocular depth calculation algorithms and selecting the appropriate algorithm based on scene information, the problem of inaccurate image depth information is solved, accurate background blur effects are achieved in different scenes, and the user experience is improved.

CN120765710AActive Publication Date: 2025-10-10HONOR DEVICE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411110628.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-13
Publication Date
2025-10-10
Estimated Expiration
2044-08-13

AI Technical Summary

Technical Problem

In the prior art, inaccurate image depth information leads to poor background blurring effect, such as blurred foreground or background.

Method used

A method combining binocular and monocular depth calculation algorithms is used to select the appropriate depth calculation algorithm based on the scene information. The binocular depth calculation algorithm is used when the conditions are met, and the monocular depth calculation algorithm is used when the conditions are not met to obtain more accurate depth information.

Benefits of technology

More accurate depth information can be obtained in different shooting scenarios, which improves the robustness of image depth estimation, makes the background blur effect more stable, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765710A_ABST
    Figure CN120765710A_ABST
Patent Text Reader

Abstract

The invention discloses an image depth estimation method, electronic equipment and a computer readable storage medium, relates to the technical field of image processing, and aims to solve the problem of how to accurately estimate image depth information. The image depth estimation method comprises the following steps: when scene information meets a binocular estimation condition, acquiring depth information by using a binocular depth calculation algorithm. And when the scene information does not meet the binocular estimation condition, obtaining depth information by using a monocular depth calculation algorithm. Wherein the scene information comprises environment information and / or image information. The environment information comprises respective hardware parameters, light sensitivity, exposure duration and multiplying power of the main camera and the auxiliary camera. The image information comprises texture information of the main path image, consistency information of the binocular image and accuracy information of binocular calibration. The binocular image comprises a main road image and an auxiliary road image, the main road image is an image collected by a main camera, and the auxiliary road image is an image collected by an auxiliary camera.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to an image depth estimation method, an electronic device, and a computer-readable storage medium. Background Art

[0002] When taking photos, electronic devices use the background blur feature to create a blurred image. In this blurred image, the foreground and subject appear sharp, while the background is blurred, giving the photo a more artistic and aesthetic quality. Background blur is achieved using an image blur algorithm that relies on image depth information. Inaccurate depth information can result in poor background blur effects, such as blurring the foreground or subject, or sharpening the background. Summary of the Invention

[0003] In view of this, embodiments of the present application provide an image depth estimation method, an electronic device, and a computer-readable storage medium, aiming to solve the problem of how to accurately estimate image depth information.

[0004] A first aspect of an embodiment of the present application provides an image depth estimation method, which is applied to an electronic device, the electronic device including a main camera and an auxiliary camera, and the method including: when the scene information meets the binocular estimation conditions, using a binocular depth calculation algorithm to obtain depth information. When the scene information does not meet the binocular estimation conditions, using a monocular depth calculation algorithm to obtain depth information. The scene information includes environmental information and / or image information. The environmental information includes the hardware parameters, sensitivity, exposure time and magnification of the main camera and the auxiliary camera respectively. The image information includes texture information of the main image, consistency information of the binocular image and accuracy information of the binocular positioning. The binocular image includes a main image and an auxiliary image, the main image is an image captured by the main camera, and the auxiliary image is an image captured by the auxiliary camera.

[0005] In this embodiment, the main camera and auxiliary camera are a combination of different cameras. For example, the main camera is a wide-angle camera and the auxiliary camera is an ultra-wide-angle camera. Another example is when the main camera is a telephoto camera and the auxiliary camera is a wide-angle camera. Scene information includes environmental information and / or image information. Environmental information can reflect the properties of the main and auxiliary cameras, while image information can reflect the quality of the binocular image. If the environmental information meets the binocular estimation conditions, it means that the properties of the main and auxiliary cameras meet the requirements of the binocular depth calculation algorithm. If the image information meets the binocular estimation conditions, it means that the binocular image quality meets the requirements of the binocular depth calculation algorithm. If the scene information meets the binocular estimation conditions, the binocular depth calculation algorithm has a low error rate. Because the depth information obtained by the binocular depth calculation algorithm represents the actual distance between the captured object and the camera, the depth information obtained using the binocular depth calculation algorithm is more accurate. If the scene information does not meet the binocular estimation conditions, it means that the binocular depth calculation algorithm has a high error rate, and the depth information obtained using the monocular depth calculation algorithm is more accurate. This allows for accurate depth information to be obtained in various shooting scenarios, thereby improving the robustness of image depth estimation.

[0006] In one embodiment, when the scene information includes image information, the scene information meeting the binocular estimation conditions includes: the texture information of the primary image includes a normal texture flag, the consistency information of the binocular images includes a consistent flag, and the accuracy information of the binocular calibration includes an accurate flag. The normal texture flag indicates that the texture of the primary image is normal, the consistent flag indicates that the primary image and the secondary image are consistent, and the accurate flag indicates that the calibration of the primary and secondary cameras is accurate.

[0007] In this embodiment, when the scene information includes image information, since the image information can reflect the quality of the input image of the binocular depth calculation algorithm, the error rate of the binocular depth calculation algorithm reflected by it is more accurate, thereby enabling a more accurate decision to use a monocular depth calculation algorithm or a binocular depth calculation algorithm.

[0008] In another embodiment, the method further includes: determining that the texture of the main road image is weak texture when the texture complexity of the main road image is less than a texture complexity threshold, and the texture information of the main road image includes a texture abnormality flag. The texture abnormality flag is used to indicate that the texture of the main road image is abnormal. Texture complexity is a numerical value obtained by statistically calculating texture features in the image using mathematical and statistical methods.

[0009] In another embodiment, the method further includes: when the texture complexity of the main image is greater than or equal to a texture complexity threshold, and the proportion of the number of target blocks in the main image is greater than or equal to a proportion threshold, determining that the texture of the main image is a repeated texture, and the texture information of the main image includes a texture anomaly flag. The similarity between the texture complexity of the target block and the texture complexity of the reference block is greater than a similarity threshold. The texture anomaly flag is used to indicate that the texture of the main image is abnormal. Texture complexity is a numerical value obtained by statistically calculating the texture features in the image using mathematical and statistical methods.

[0010] In another embodiment, the method further includes: determining that the primary image and the secondary image are consistent when the brightness difference between the primary image and the secondary image is within a target brightness range, the clarity difference is within a target clarity range, the color temperature difference is within a target color temperature range, and the exposure time difference is within a target time difference range, wherein the consistency information of the binocular images includes a consistency flag. The exposure time difference is the time difference between the exposure times of the primary image and the exposure times of the secondary image.

[0011] In another embodiment, the method further includes: when the error of the dual-target calibration is less than the error threshold, determining that the calibration of the main camera and the auxiliary camera is accurate, and the accuracy information of the dual-target calibration includes an accurate flag.

[0012] In another embodiment, the method further includes: detecting feature point pairs of the main road image and the auxiliary road image, the feature point pairs including feature points of the main road image and feature points of the auxiliary road image at corresponding positions. Calculating the internal and external parameters of the main camera and the auxiliary camera, the internal parameters including focal length, principal point, pixel size, and distortion coefficient, and the external parameters including rotation matrix and translation vector. Obtaining reprojected feature point pairs based on the internal and external parameters of the main camera and the auxiliary camera, the reprojected feature point pairs including reprojected feature points of the main road image and reprojected feature points of the auxiliary road image at corresponding positions. Calculating the error of binocular positioning based on the feature point pairs and the reprojected feature point pairs at corresponding positions.

[0013] In another embodiment, the method further includes: when the disparity value between the main road image and the auxiliary road image is less than 0, determining that the calibration of the main camera and the auxiliary camera is inaccurate, and the accuracy information of the dual-target calibration includes an inaccurate flag.

[0014] In another embodiment, the method also includes: when the number of feature point pairs of the main road image and the auxiliary road image is less than a number threshold, determining that the calibration of the main camera and the auxiliary camera is inaccurate, the accuracy information of the dual-target calibration includes an inaccurate flag, and the feature point pairs include feature points of the main road image and feature points of the auxiliary road image at corresponding positions.

[0015] In another embodiment, when the scene information includes environmental information, the scene information satisfies the binocular estimation conditions including: the hardware parameters of the primary camera and the secondary camera are within a target parameter range, the sensitivity is within a target sensitivity range, the exposure time is within a target time range, and the magnification is within a target magnification range. The hardware parameters include resolution, focal length, aperture size, and field of view.

[0016] In another embodiment, the hardware parameters of the main camera and the auxiliary camera are within the target parameter range, including: the resolution of the main camera and the auxiliary camera is within the target resolution range, the focal length is within the target focal length range, the aperture size is within the target aperture range, and the field of view angle is within the target viewing angle range.

[0017] In another embodiment, when the scene information includes environmental information and image information, the scene information satisfies the binocular estimation conditions including: the texture information of the main road image includes a normal texture flag, the consistency information of the binocular image includes a consistency flag, the accuracy information of the binocular positioning includes a consistency flag, the hardware parameters of the main camera and the auxiliary camera are within the target parameter range, the sensitivity is within the target sensitivity range, the exposure time is within the target time range, and the magnification is within the target magnification range.

[0018] A second aspect of an embodiment of the present application provides an electronic device, which includes a memory, a processor and a camera, the camera including a main camera and an auxiliary camera, the memory and the camera are electrically connected to the processor, and when the processor executes computer instructions stored in the memory, the image depth estimation method provided by the first aspect is implemented.

[0019] A third aspect of an embodiment of the present application provides a computer-readable storage medium having computer instructions stored thereon, and when a processor executes the computer instructions, the image depth estimation method provided in the first aspect is implemented.

[0020] A fourth aspect of an embodiment of the present application provides a computer program product, which includes computer instructions. When a processor executes the computer instructions, the image depth estimation method provided in the first aspect is implemented.

[0021] It can be understood that the beneficial effects of the electronic device provided in the second aspect, the computer-readable storage medium provided in the third aspect, and the computer program product provided in the fourth aspect of the embodiment of the present application are roughly the same as the beneficial effects of the image depth estimation method provided in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 The following is a schematic diagram of the software structure of an electronic device provided as an example.

[0023] Figure 2 The present invention is a schematic diagram of an interface of an electronic device provided as an example.

[0024] Figure 3 This is another example of an electronic device interface diagram.

[0025] Figure 4 This is a logical structure diagram of an image blurring algorithm provided as an example.

[0026] Figure 5 This is a logical structure diagram of a monocular depth calculation algorithm provided as an example.

[0027] Figure 6 This is a logical structure diagram of a binocular depth calculation algorithm provided as an example.

[0028] Figure 7 This is a logical structure diagram of a depth calculation algorithm provided as an example.

[0029] Figure 8 The present invention is a flowchart of an electronic device implementing a background blur function provided by an example.

[0030] Figure 9 This is another example of an electronic device interface diagram.

[0031] Figure 10 This is another example of an electronic device interface diagram.

[0032] Figure 11 This is a schematic diagram of a blurred image provided by an example.

[0033] Figure 12 The present invention is a flowchart of an electronic device implementing an image blurring algorithm provided by an example.

[0034] Figure 13 The following is a schematic diagram of the hardware structure of an electronic device provided as an example. DETAILED DESCRIPTION

[0035] It should be noted that in the embodiments of this application, "multiple" and "several" refer to two or more than two. The terms "first," "second," "third," "fourth," etc. in the specification, claims, and drawings of this application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0036] It should also be noted that the method disclosed in the embodiments of the present application or the method shown in the flowchart includes one or more steps for implementing the method. Without departing from the scope of the claims, the execution order of multiple steps can be interchanged with each other, and some steps can also be deleted.

[0037] The electronic devices of the embodiments of the present application include smartphones with shooting functions, tablet computers, PDAs, laptop computers, mobile Internet devices (MIDs), virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, cellular phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), handheld devices with wireless communication functions, computing devices or other processing devices connected to wireless modems, vehicle-mounted devices, wearable devices, terminal devices in 5G networks, or terminal devices in public land mobile networks (PLMNs), etc.

[0038] The software system of the electronic device is described in detail below.

[0039] The software system of electronic equipment can adopt a layered architecture, which divides the software into several layers. Each layer has a clear role and division of labor, and the layers communicate with each other through software interfaces. Figure 1 As shown in FIG, the software system of the electronic device is divided into four layers, namely, the application (APP) layer, the framework (FWK) layer, the hardware abstraction layer (HAL) and the hardware layer from top to bottom.

[0040] The application layer includes a series of application packages. The application package includes a camera application. The camera application has a background blur function. For example, Figure 2 As shown, when the user clicks on the camera application, the screen of the electronic device displays the shooting interface. On the shooting interface, when the user selects the large aperture mode, the camera application turns on the background blur function. For another example, Figure 3As shown, on the shooting interface, when the user selects portrait mode, the camera application turns on the background blur function.

[0041] The background blur feature blurs the background in an image, enhancing the subject and foreground, creating a more layered image. An image consists of a subject, foreground, and background. The subject is the object the camera is focused on, such as a person, animal, or landscape. The image in front of the subject is called the foreground, and the image behind the subject is called the background.

[0042] The framework layer provides the application programming interface (API) and programming framework for various apps in the application layer. The framework layer includes the camera framework (CameraFWK), which provides the camera API (CameraAPI) for camera applications.

[0043] The hardware abstraction layer (HAL) is used to execute various app functions and process data. It includes the camera hardware abstraction layer (CameraHAL), which includes the processor driver, camera driver, and camera algorithm library.

[0044] The processor driver is a control node for the processor, which is used to control the processor to implement image processing algorithms.

[0045] The camera driver is a control node for the camera, which is used to control the camera to capture images.

[0046] The camera algorithm library is used to store a series of image processing algorithms, such as image blurring (Image Blurring) algorithm, automatic exposure (AE) algorithm, automatic focus (AF) algorithm, electric image stabilization (EIS) algorithm, picture quality (PQ) algorithm, etc.

[0047] The hardware layer provides hardware support for the functional applications and data processing of various APPs in the hardware abstraction layer. The hardware layer includes processors and cameras. The processors include a central processing unit (CPU), a graphics processing unit (GPU), and a neural network processor (NPU). The CPU is used to process control logic and serial computing tasks. The GPU is used to process image algorithm logic and parallel computing tasks. The NPU is used to process neural network computing tasks. The cameras include wide-angle cameras, ultra-wide-angle cameras, and telephoto cameras. Various cameras have corresponding field of view (FOV) and focal length.

[0048] For example, when using an electronic device to shoot a distant subject, the zoom ratio can be increased to shorten the distance between the electronic device and the subject. The electronic device uses cameras with different focal lengths to take photos / videos according to different zoom ratios and different object distances, which can obtain higher-definition images. During the zoom process, when the zoom ratio is greater than or equal to the ratio threshold, the electronic device switches to a different camera. For example, as the zoom ratio increases from the minimum value, the electronic device uses the ultra-wide-angle camera, the wide-angle camera, and the telephoto camera in sequence.

[0049] In different shooting modes, electronic devices can use different camera combinations. For example, in large aperture mode or portrait mode, when the zoom ratio is less than the magnification threshold, the electronic device uses the wide-angle camera and the ultra-wide-angle camera, with the wide-angle camera being the primary camera and the ultra-wide-angle camera being the secondary camera. When the zoom ratio is greater than or equal to the magnification threshold, the electronic device uses the telephoto camera and the wide-angle camera, with the telephoto camera being the primary camera and the wide-angle camera being the secondary camera. The image captured by the primary camera is called the primary image, and the image captured by the secondary camera is called the secondary image.

[0050] exist Figure 1 In the camera algorithm library shown, the image blur algorithm obtains image depth information and then blurs the background of the main image based on this depth information. Depth information indicates the distance of the subject from the camera in three-dimensional space. Depth information is also called a depth value, and an image containing this depth information is called a depth map.

[0051] For example, Figure 4As shown in Figure 1, the logical structure of the image blur algorithm includes a depth calculation module and a blur rendering module. The depth calculation module is used to calculate the depth of the image to obtain a depth map. The blur rendering module is used to blur the background of the main image based on the depth map to obtain a blurred image.

[0052] The depth calculation module includes a monocular depth calculation module and a binocular depth calculation module. The monocular depth calculation module is used to perform depth calculation on a monocular image to obtain a depth map, where the monocular image is the primary image. The binocular depth calculation module is used to perform depth calculation on a binocular image to obtain a depth map, where the binocular image includes the primary image and the secondary image.

[0053] For example, Figure 5 As shown, the monocular depth calculation module includes a front-end processing module and a monocular depth estimation model. The front-end processing module is used to preprocess the main image to obtain a preprocessed image. Preprocessing may include image cropping, so that the size of the cropped preprocessed image meets the input image size requirements of the monocular depth estimation model. Preprocessing may also include image enhancement, noise removal, geometric transformation, and other functions to meet different image processing requirements.

[0054] The monocular depth estimation model is used to calculate the depth of the preprocessed image and obtain a depth map. Monocular depth estimation models can use the Dense Prediction Transformer (DPT) model, the MiDaS model, the Adabins model, the PackNet-SfM model, and other models.

[0055] For example, Figure 6 As shown, the binocular depth calculation module includes a front-end processing module, an epipolar correction module, and a binocular depth estimation model. The front-end processing module is used to preprocess the primary and secondary images to generate preprocessed image pairs. These preprocessed image pairs contain the primary and secondary preprocessed images. Preprocessing may include image cropping, where the size of the cropped preprocessed image pairs meets the input image size requirements of the binocular depth estimation model. Preprocessing may also include image enhancement, noise removal, geometric transformation, and other techniques to meet diverse image processing needs.

[0056] The epipolar correction module is used to perform epipolar correction on the preprocessed image pair to obtain a corrected image pair, which includes a main corrected image and an auxiliary corrected image. Epipolar correction is to align the scan lines of the image with the epipolar lines. The epipolar correction module includes a static correction module and a dynamic correction module. The static correction module is used to calculate the correspondence between the preprocessed image pairs through the fundamental matrix (Fundamental Matrix) or the essential matrix (Essential Matrix) based on the results of binocular calibration, so that the pixel pairs at corresponding positions in the main preprocessed image and the auxiliary preprocessed image are aligned on the same horizontal line. The dynamic correction module is used to update the correspondence between the preprocessed image pairs based on the feature point pairs at corresponding positions in the preprocessed image pairs, so that the pixel pairs at corresponding positions in the main preprocessed image and the auxiliary preprocessed image are further aligned on the same horizontal line, so that the correspondence between the preprocessed image pairs adapts to the position changes of the photographed objects, thereby meeting the requirements of epipolar correction.

[0057] The binocular depth estimation model is used to calculate the depth of the rectified image pair and obtain a depth map. Examples of binocular depth estimation models include the DispNetC model, the GC-Net model, the Pyramid Stereo Matching Network (PSMNet), the StereoDRNet model, the LW-Stereo model, and the DSMNet model.

[0058] It can be understood that the depth information obtained by the monocular depth calculation module represents the distance between the photographed object and the camera, and the depth information obtained by the binocular depth calculation module represents the actual distance between the photographed object and the camera.

[0059] In this embodiment, if Figure 7 As shown, the depth calculation module includes a decision module, and the decision module includes a monocular depth calculation module and a binocular depth calculation module. The decision module is used to determine whether to use a monocular depth calculation module or a binocular depth calculation module according to the scene information. Specifically, when the scene information meets the binocular estimation conditions, the binocular depth calculation module is used. When the scene information does not meet the binocular estimation conditions, the monocular depth calculation module is used. Scene information includes environmental information and / or image information. Environmental information includes the hardware parameters, sensitivity (also known as ISO value), exposure time and magnification of the main camera and auxiliary camera. Image information includes texture information of the main road image, consistency information of the binocular image and accuracy information of the binocular positioning.

[0060] The texture information of the primary image includes a normal texture flag or an abnormal texture flag. The normal texture flag indicates that the texture of the primary image is normal. The abnormal texture flag indicates that the texture of the primary image is abnormal. The specific forms of the normal texture flag and the abnormal texture flag can be set as needed. For example, the normal texture flag is "true" and the abnormal texture flag is "false". In another example, the normal texture flag is "1" and the abnormal texture flag is "0" or "null".

[0061] When it is determined that the texture of the main road image is an abnormal texture, the texture information of the main road image includes a texture abnormality flag, otherwise the texture information of the main road image includes a texture normal flag. Exemplarily, abnormal textures include weak textures and repeated textures, repeated textures are textures with regularity in the image, and weak textures are textures without obvious regularity in the image. When the texture complexity of the main road image is less than the texture complexity threshold, the texture of the main road image is determined to be a weak texture. When the texture complexity of the main road image is greater than or equal to the texture complexity threshold, and the number of target blocks (patches) in the main road image accounts for a greater than or equal proportion threshold, the texture of the main road image is determined to be a repeated texture. When the similarity between the texture complexity of the i-th block and the texture complexity of the reference block is greater than the similarity threshold, the i-th block is determined to be the target block, where i is a positive integer. The reference block can be any block in the main road image. Texture complexity is a value obtained by statistically calculating the texture features in the image through mathematical and statistical methods. Texture complexity is used to measure the richness and detail of the texture in the image.

[0062] It is understood that the texture complexity threshold, ratio threshold, and similarity threshold can be set as needed. Image texture detection methods may include using a local binary pattern (LBP) histogram to detect texture features, using a wavelet transform / Fourier transform to detect texture features, or using a Gabor filter to detect texture features.

[0063] The consistency information for binocular images includes a consistency flag and an inconsistency flag. The consistency flag indicates that the primary and secondary images are consistent. The inconsistency flag indicates that the primary and secondary images are inconsistent. The specific forms of the consistency flag and inconsistency flag can be customized. For example, the consistency flag is "true" and the inconsistency flag is "false." Another example is the consistency flag is "1" and the inconsistency flag is "0" or "null."

[0064] When the primary and secondary images are determined to be consistent, the binocular image consistency information includes a consistent flag; otherwise, the binocular image consistency information includes an inconsistent flag. For example, consistency is determined when the brightness difference between the primary and secondary images is within a target brightness range, the clarity difference is within a target clarity range, the color temperature difference is within a target color temperature range, and the exposure time difference is within a target time difference range. The exposure time difference is the time difference between the exposure times of the primary and secondary images.

[0065] It is understood that the target brightness range, target sharpness range, target color temperature range, and target time difference range can be set as needed. The brightness value of the image can be calculated using the calculate_brightness function in OpenCV, and the image sharpness can be calculated using the variance_of_laplacian function. The color temperature of the image can be estimated using a pre-trained regression model, and the exposure time can be obtained through the camera's photosensitive element. The sharpness of the image can also be calculated using the No-Reference Statistical Sharpness (NRSS) function, and the color temperature of the image can be estimated using the Correlated Color Temperature (CCT) function.

[0066] The accuracy information for dual-target calibration includes an accurate flag or an inaccurate flag. The accurate flag indicates that the calibration of the primary and secondary cameras is accurate. The inaccurate flag indicates that the calibration of the primary and secondary cameras is inaccurate. The specific forms of the accurate and inaccurate flags can be customized. For example, the accurate flag is "true" and the inaccurate flag is "false." Another example is the accurate flag is "1" and the inaccurate flag is "0" or "null."

[0067] When the calibrations of the primary camera and the secondary camera have accuracy, the accuracy information of the binocular calibration includes an accurate mark, otherwise the accuracy information of the binocular calibration includes an inaccurate mark. Illustratively, when the error of the binocular calibration is less than an error threshold, it is determined that the calibrations of the primary camera and the secondary camera have accuracy. Calculating the error of the binocular calibration includes detecting feature point pairs of the primary image and the secondary image, calculating intrinsic parameters and extrinsic parameters of the primary camera and the secondary camera, obtaining re-projection feature point pairs according to the intrinsic parameters and the extrinsic parameters of the primary camera and the secondary camera, and calculating the error of the binocular calibration according to the feature point pairs and the re-projection feature point pairs at the corresponding positions. The feature point pairs include a feature point of the primary image and a feature point of the secondary image at a corresponding position. Re-projection is to re-project a reference feature point known in a world coordinate system to an image coordinate system, and the re-projection feature point pairs include a re-projection feature point of the primary image and a re-projection feature point of the secondary image at a corresponding position. The intrinsic parameters represent optical and geometric characteristics inside the camera, and the intrinsic parameters include focal length, principal point, pixel size, and distortion coefficient, and the principal point is the intersection of the optical axis and the image plane. The extrinsic parameters represent the relative position and attitude between the primary camera and the secondary camera, and the extrinsic parameters include a rotation matrix and a translation vector, the rotation matrix represents the rotation relationship between the primary camera coordinate system and the secondary camera coordinate system, and the translation vector represents the translation relationship between the primary camera coordinate system and the secondary camera coordinate system.

[0068] In an embodiment, when the disparity value of the primary image and the secondary image is less than 0, it is determined that the calibrations of the primary camera and the secondary camera do not have accuracy. The disparity value is the difference between the column coordinate of a pixel in the primary image and the column coordinate of the corresponding pixel in the secondary image.

[0069] In another embodiment, when the number of the feature point pairs of the primary image and the secondary image is less than a number threshold, it is determined that the calibrations of the primary camera and the secondary camera do not have accuracy.

[0070] It can be understood that the error threshold and the number threshold of the above embodiments can be set as required.

[0071] In the embodiment, when the scene information includes image information, the scene information satisfying the binocular estimation condition includes that the texture information of the primary image includes a texture normal mark, the consistency information of the binocular image includes a consistent mark, and the accuracy information of the binocular calibration includes a consistent mark.

[0072] When the scene information includes environment information, the scene information satisfying the binocular estimation condition includes that the hardware parameters of the primary camera and the secondary camera are respectively within a target parameter range, the sensitivity is within a target sensitivity range, the exposure time is within a target time range, and the magnification is within a target magnification range.

[0073] It is understood that the target parameter range, target sensitivity range, target duration range, and target magnification range can be set as needed. For example, the hardware parameters of the camera include resolution, focal length, aperture size, and field of view angle. The hardware parameters of the camera within the target parameter range include: resolution within the target resolution range, focal length within the target focal length range, aperture size within the target aperture range, and field of view angle within the target viewing angle range. The target resolution range, target focal length range, target aperture range, and target viewing angle range can be set as needed.

[0074] When the scene information includes environmental information and image information, the scene information satisfies the binocular estimation conditions including: texture information of the main road image includes a texture normal flag, consistency information of the binocular image includes a consistency flag, accuracy information of the binocular positioning includes a consistency flag, hardware parameters of the main camera and the auxiliary camera are within the target parameter range, sensitivity is within the target sensitivity range, exposure time is within the target time range, and magnification is within the target magnification range.

[0075] The logical structure of the image blurring algorithm has been specifically described above. In this embodiment, a decision module is added to the image blurring algorithm to determine whether to use a monocular depth calculation module or a binocular depth calculation module based on scene information. When the scene information meets the binocular estimation conditions, the binocular depth calculation module has a low error rate. Since the depth information obtained by the binocular depth calculation module represents the actual distance between the subject and the camera, the depth information obtained by the binocular depth calculation module is more accurate, thereby achieving a better background blur effect. When the scene information does not meet the binocular estimation conditions, the binocular depth calculation module has a high error rate, and the depth information obtained by the monocular depth calculation module is more accurate. This allows for more accurate depth information to be obtained in various shooting scenarios, thereby improving the robustness of image depth estimation, making the background blurring function of the camera application more stable, and thus enhancing the user experience. Furthermore, when the scene information includes image information, since the image information can reflect the quality of the input image to the binocular depth calculation module, it reflects the error rate of the binocular depth calculation module more accurately, thereby enabling a more accurate decision on whether to use the monocular depth calculation module or the binocular depth calculation module.

[0076] The following combination Figure 1 The software system of the electronic device shown specifically describes the implementation process of the background blur function.

[0077] like Figure 8 As shown in FIG, the implementation process of the background blur function includes the following steps:

[0078] S101: When the shooting mode of the camera application is the target mode and the shooting control is triggered, the camera application sends a shooting instruction to the Camera API.

[0079] Target mode is a shooting mode that turns on the background blur function. The camera application can provide multiple shooting modes, such as large aperture mode, portrait mode, night scene mode, etc. For example, the target mode includes large aperture mode and portrait mode. Figure 2 and Figure 3 As shown in FIG, when a user clicks on the camera application, the screen of the electronic device displays a shooting interface. On the shooting interface, when the user selects a large aperture mode or a portrait mode, the camera application turns on a background blur function.

[0080] The shooting control is a control for controlling shooting. For example, Figure 9 As shown, the user clicks the shooting control 11 on the shooting interface to trigger the shooting control. Figure 10 As shown, the user presses the volume button 12 of the electronic device to trigger the shooting control. The volume button 12 includes a volume up button and a volume down button.

[0081] The shooting command is used to instruct shooting in target mode.

[0082] S102. CameraAPI sends a shooting instruction to the camera driver.

[0083] S103: The camera driver controls the camera to capture images in response to the shooting instruction.

[0084] The cameras include a main camera and an auxiliary camera, and the images include a main road image and an auxiliary road image.

[0085] S104 , the processor driver obtains an image from the camera driver, calls an image blurring algorithm from a camera algorithm library, controls the CPU, GPU, and the GPU to process the image to obtain a blurred image.

[0086] In this embodiment, the image blurring algorithm is nested in the depth calculation algorithm and the blurring rendering algorithm, the depth calculation algorithm is nested in the decision algorithm, and the decision algorithm is nested in the monocular depth calculation algorithm and the binocular depth calculation algorithm.

[0087] It is understandable that the depth calculation algorithm is Figure 4 The depth calculation module shown corresponds to the virtual rendering algorithm. Figure 4 The virtual rendering module shown corresponds to the decision algorithm Figure 7 The decision module shown corresponds to the monocular depth calculation algorithm. Figure 5 The monocular depth calculation module shown corresponds to the binocular depth calculation algorithm. Figure 6 The binocular depth calculation module shown corresponds to.

[0088] S105: The processor drives the camera to send the blurred image to the camera API.

[0089] S106 . Camera API sends the blurred image to the camera application.

[0090] S107: The camera application stores the blurred image in the gallery.

[0091] For example, Figure 11 As shown, in the blurred image, the subject and foreground are clearer, while the background is blurred.

[0092] The above describes in detail the implementation process of the background blur function. The implementation of the background blur function depends on the image blur algorithm. The following describes in detail the implementation process of the image blur algorithm.

[0093] like Figure 12 As shown, the implementation of the image blurring algorithm includes the following steps:

[0094] S201. Obtain scene information.

[0095] Scene information includes environmental information and / or image information. Environmental information includes the hardware parameters, sensitivity, exposure time, and magnification of the main and auxiliary cameras. Image information includes the texture information of the main image, the consistency of the binocular images, and the accuracy of binocular positioning.

[0096] The texture information of the main image includes a normal texture flag or an abnormal texture flag. The normal texture flag is used to indicate that the texture of the main image is normal. The abnormal texture flag is used to indicate that the texture of the main image is abnormal.

[0097] The consistency information of the binocular images includes a consistency flag or an inconsistency flag. The consistency flag is used to indicate that the primary image and the secondary image are consistent. The inconsistency flag is used to indicate that the primary image and the secondary image are inconsistent.

[0098] The accuracy information for dual-camera calibration includes an accurate flag or an inaccurate flag. The accurate flag indicates that the calibration of the primary and secondary cameras is accurate. The inaccurate flag indicates that the calibration of the primary and secondary cameras is inaccurate.

[0099] S202: Determine whether the scene information meets binocular estimation conditions.

[0100] If yes, execute step S203; if no, execute step S204.

[0101] After executing step S203 or S204, execute step S205.

[0102] S203: When the scene information meets the binocular estimation conditions, a binocular depth calculation algorithm is used to obtain a depth map.

[0103] When the scene information includes image information, the scene information satisfies the binocular estimation condition including: texture information of the main road image includes a texture normal flag, consistency information of the binocular image includes a consistency flag, and accuracy information of the binocular positioning includes an accurate flag.

[0104] When the scene information includes environmental information, the scene information satisfies the binocular estimation conditions including: the hardware parameters of the main camera and the auxiliary camera are within the target parameter range, the sensitivity is within the target sensitivity range, the exposure time is within the target time range, and the magnification is within the target magnification range.

[0105] When the scene information includes environmental information and image information, the scene information satisfies the binocular estimation conditions including: texture information of the main road image includes a texture normal flag, consistency information of the binocular image includes a consistency flag, accuracy information of the binocular positioning includes an accurate flag, hardware parameters of the main camera and the auxiliary camera are within the target parameter range, sensitivity is within the target sensitivity range, exposure time is within the target time range, and magnification is within the target magnification range.

[0106] Implementing a binocular depth calculation algorithm involves preprocessing the primary and secondary images to obtain a preprocessed image pair, which includes the primary and secondary preprocessed images. Performing epipolar correction on the preprocessed image pair to obtain a corrected image pair, which includes the primary and secondary corrected images. Depth calculation is performed on the corrected image pair to obtain a depth map.

[0107] Among them, epipolar correction includes static correction and dynamic correction. Static correction includes: based on the results of binocular calibration, calculating the correspondence between the preprocessed image pairs through the basic matrix or the essential matrix, so that the pixel pairs at corresponding positions in the main preprocessed image and the auxiliary preprocessed image are aligned on the same horizontal line. Dynamic correction includes: updating the correspondence between the preprocessed image pairs based on the feature point pairs at corresponding positions in the preprocessed image pairs, so that the pixel pairs at corresponding positions in the main preprocessed image and the auxiliary preprocessed image are further aligned on the same horizontal line, so that the correspondence between the preprocessed image pairs adapts to the position changes of the photographed objects, thereby meeting the requirements of epipolar correction.

[0108] S204: When the scene information does not meet the binocular estimation conditions, a monocular depth calculation algorithm is used to obtain a depth map.

[0109] When the scene information includes image information, the scene information does not meet the binocular estimation conditions including: texture information of the main road image includes a texture abnormality flag, and / or consistency information of the binocular image includes an inconsistent flag, and / or accuracy information of the binocular positioning includes an inaccurate flag.

[0110] When the scene information includes environmental information, the scene information does not meet the binocular estimation conditions, including: the hardware parameters of the main camera and the auxiliary camera are not within the target parameter range, and / or the sensitivity is not within the target sensitivity range, and / or the exposure time is not within the target time range, and / or the magnification is not within the target magnification range.

[0111] When the scene information includes environmental information and image information, the scene information does not meet the binocular estimation conditions, including: the texture information of the main road image includes a texture abnormality flag, and / or the consistency information of the binocular image includes an inconsistency flag, and / or the accuracy information of the binocular positioning includes an inaccurate flag, and / or the respective hardware parameters of the main camera and the auxiliary camera are not within the target parameter range, and / or the sensitivity is not within the target sensitivity range, and / or the exposure time is not within the target time range, and / or the magnification is not within the target magnification range.

[0112] Implementing a monocular depth calculation algorithm includes: preprocessing the main image to obtain a preprocessed image; performing depth calculation on the preprocessed image to obtain a depth map.

[0113] S205 , blurring the background of the main road image according to the depth map to obtain a blurred image.

[0114] For example, the blurring process may be performed by using Gaussian blur.

[0115] The above specifically describes the implementation process of the image blur algorithm. The above steps S201-S204 are the image depth estimation method of this embodiment. In this embodiment, by adding a decision algorithm to the image blur algorithm, it is determined whether to use a monocular depth calculation algorithm or a binocular depth calculation algorithm based on the scene information. When the scene information meets the binocular estimation conditions, it means that the error rate of the binocular depth calculation algorithm is low. Since the depth information obtained by the binocular depth calculation algorithm represents the true distance between the photographed object and the camera, the depth information obtained by the binocular depth calculation algorithm is more accurate, thereby being able to present a better background blur effect. When the scene information does not meet the binocular estimation conditions, it means that the error rate of the binocular depth calculation algorithm is high, and the depth information obtained by the monocular depth calculation algorithm is more accurate. As a result, more accurate depth information can be obtained in different shooting scenes, thereby improving the robustness of the image depth estimation, making the background blur function of the camera application more stable, and thus improving the user experience. Furthermore, when the scene information includes image information, since the image information can reflect the quality of the input image of the binocular depth calculation algorithm, the error rate of the binocular depth calculation algorithm reflected by it is more accurate, thereby enabling a more accurate decision to use a monocular depth calculation algorithm or a binocular depth calculation algorithm.

[0116] The hardware structure of the electronic device is described in detail below.

[0117] like Figure 13 As shown, the electronic device includes a processor 110, an external memory interface 120, an internal memory 121, a sensor module 130, a display screen 140, and a camera module 150. The sensor module 130 includes a touch sensor 131 and an image sensor 132. The camera module 150 may include a wide-angle camera 151, an ultra-wide-angle camera 152, and a telephoto camera 153.

[0118] The processor 110 is used to execute the various functions or steps performed by the electronic device in the above embodiments. The processor 110 includes a CPU, a GPU, and a GPU. The processor 110 may also include an application processor (AP), an image signal processor (ISP), a digital signal processor (DSP), a modem processor, and a video codec.

[0119] The external memory interface 120 is used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor 110 through the external memory interface 120 to implement data storage.

[0120] The internal memory 121 is used to store computer executable program code, and the executable program code includes instructions. The processor 120 executes the various functions or steps performed by the electronic device in the above embodiment by running the instructions stored in the internal memory 121. The internal memory 121 includes a program storage area and a data storage area. The program storage area can store an operating system, an APP required for at least one function, etc. The data storage area can store data created during the use of the electronic device. The internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, and a universal flash memory (UFS).

[0121] The display screen 140 is used to display an interface. The display screen 140 includes a display panel. The display panel can be a liquid crystal display (LCD), a light-emitting diode (LED), or an organic light-emitting diode (OLED).

[0122] If the display screen 140 integrates the touch sensor 131, the display screen 140 may be referred to as a touch screen. The touch sensor 131 may also be referred to as a "touch panel." That is, the display screen 140 may include both a display panel and a touch panel. The touch sensor 131 is used to detect touch operations applied to or near the display screen. After the touch sensor 131 detects a touch operation, it may trigger the driver of the hardware abstraction layer of the electronic device to periodically scan the touch parameters generated by the touch operation. The driver of the hardware abstraction layer then transmits the touch parameters to the relevant modules in the upper layer, so that the relevant modules can determine the touch event corresponding to the touch parameters.

[0123] The electronic device can realize the display function through the GPU, display screen 140 and AP, and can realize the shooting function through the GPU, ISP, camera module 150, image sensor 132, video codec, display screen 140 and AP, etc.

[0124] The electronic device can capture light through any one or more cameras in the camera module 150 and transmit the light signal to the image sensor 132. The light signal is converted into a raw image (or RAW image) by the image sensor 132. Then, the RAW image is converted into a YUV image through the ISP, and the YUV image is converted into an RGB image, thereby completing the image acquisition. Among them, RAW is a raw image data format that has not been compressed or modified. YUV is a color coding format, Y represents luminance, and U and V represent chrominance. The RGB image is an image composed of three color channels: red, green, and blue. The RGB image can be displayed directly on the display screen 140, or it can be subsequently edited and processed.

[0125] It is understood that the structure shown in this embodiment does not constitute a specific limitation on the electronic device. In other embodiments, the electronic device may include more or fewer components than shown, or combine or separate some components, or arrange the components differently.

[0126] The various functions or steps performed by the electronic device in the above embodiments may also be applied to a chip, a computer-readable storage medium or a computer program product.

[0127] The chip includes a processor and an interface circuit, which are electrically connected to each other. The interface circuit can read computer instructions stored in the memory and send the computer instructions to the processor. When the processor executes the computer instructions, the various functions or steps performed by the electronic device in the above embodiments are realized.

[0128] The computer-readable storage medium stores computer instructions, and when the processor executes the computer instructions, the various functions or steps performed by the electronic device in the above embodiments are implemented.

[0129] Computer-readable storage media include volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media include, but are not limited to, random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer.

[0130] The computer program product includes computer instructions, and when a processor executes the computer instructions, the various functions or steps performed by the electronic device in the above embodiments are implemented.

[0131] The embodiments of the present application are described in detail above in conjunction with the accompanying drawings, but the present application is not limited to the above embodiments. Various changes can be made within the scope of knowledge possessed by ordinary technicians in the relevant technical field without departing from the purpose of the present application.

Claims

1. A method for image depth estimation, applied to an electronic device, wherein the electronic device includes a main camera and an auxiliary camera, characterized in that: The method comprises: When the scene information meets the binocular estimation conditions, the binocular depth calculation algorithm is used to obtain the depth information; When the scene information does not meet the binocular estimation conditions, a monocular depth calculation algorithm is used to obtain depth information; Among them, the scene information includes environmental information and / or image information; the environmental information includes the hardware parameters, sensitivity, exposure time and magnification of the main camera and the auxiliary camera respectively; the image information includes the texture information of the main image, the consistency information of the binocular image and the accuracy information of the binocular positioning; the binocular image includes the main image and the auxiliary image, the main image is the image captured by the main camera, and the auxiliary image is the image captured by the auxiliary camera.

2. The image depth estimation method according to claim 1, wherein: When the scene information includes the image information, the scene information meeting the binocular estimation condition includes: The texture information of the main road image includes a texture normal flag, the consistency information of the binocular image includes a consistency flag, and the accuracy information of the binocular positioning includes an accurate flag; Among them, the normal texture flag is used to indicate that the texture of the main road image is normal, the consistent flag is used to indicate that the main road image and the auxiliary road image are consistent, and the accurate flag is used to indicate that the calibration of the main camera and the auxiliary camera is accurate.

3. The image depth estimation method according to claim 2, wherein: The method further comprises: When the texture complexity of the main road image is less than a texture complexity threshold, determining that the texture of the main road image is a weak texture, and the texture information of the main road image includes a texture abnormality flag; The texture abnormality flag is used to indicate that the texture of the main road image is abnormal; the texture complexity is a value obtained by statistically calculating the texture features in the image using mathematical and statistical methods.

4. The image depth estimation method according to claim 2 or 3, wherein: The method further comprises: When the texture complexity of the main road image is greater than or equal to a texture complexity threshold, and the number of target blocks in the main road image is greater than or equal to a ratio threshold, it is determined that the texture of the main road image is a repeated texture, and the texture information of the main road image includes a texture abnormality flag; Among them, the similarity between the texture complexity of the target block and the texture complexity of the reference block is greater than a similarity threshold; the texture anomaly flag is used to indicate that the texture of the main road image is abnormal; the texture complexity is a numerical value obtained by statistically calculating the texture features in the image through mathematical and statistical methods.

5. The image depth estimation method according to any one of claims 2 to 4, characterized in that: The method further comprises: When the brightness difference between the primary image and the secondary image is within a target brightness range, the clarity difference is within a target clarity range, the color temperature difference is within a target color temperature range, and the exposure time difference is within a target time difference range, it is determined that the primary image and the secondary image are consistent, and the consistency information of the binocular images includes a consistency flag; The exposure time difference is the time difference between the exposure time of the main image and the exposure time of the auxiliary image.

6. The image depth estimation method according to any one of claims 2 to 5, wherein: The method further comprises: When the error of the dual-target calibration is less than the error threshold, it is determined that the calibration of the main camera and the auxiliary camera is accurate, and the accuracy information of the dual-target calibration includes an accurate flag.

7. The image depth estimation method according to claim 6, wherein: The method further comprises: Detecting feature point pairs of the main road image and the auxiliary road image, the feature point pairs comprising feature points of the main road image and feature points of the auxiliary road image at corresponding positions; Calculating intrinsic and extrinsic parameters of the main camera and the auxiliary camera, the intrinsic parameters including focal length, principal point, pixel size, and distortion coefficient, and the extrinsic parameters including rotation matrix and translation vector; Obtaining a reprojection feature point pair according to the intrinsic parameters and extrinsic parameters of the main camera and the auxiliary camera, the reprojection feature point pair comprising a reprojection feature point of the main road image and a reprojection feature point of the auxiliary road image at a corresponding position; The error of the binocular positioning is calculated according to the feature point pair and the reprojected feature point pair at the corresponding position.

8. The image depth estimation method according to any one of claims 2 to 7, wherein: The method further comprises: When the disparity value between the main road image and the auxiliary road image is less than 0, it is determined that the calibration of the main camera and the auxiliary camera is inaccurate, and the accuracy information of the dual-target calibration includes an inaccurate flag.

9. The image depth estimation method according to any one of claims 2 to 8, wherein: The method further comprises: When the number of feature point pairs of the main road image and the auxiliary road image is less than a number threshold, it is determined that the calibration of the main camera and the auxiliary camera is inaccurate, the accuracy information of the dual-target calibration includes an inaccurate flag, and the feature point pairs include the feature points of the main road image and the feature points of the auxiliary road image at corresponding positions.

10. The image depth estimation method according to any one of claims 1 to 9, wherein: When the scene information includes the environmental information, the scene information meeting the binocular estimation condition includes: The hardware parameters of the main camera and the auxiliary camera are within the target parameter range, the sensitivity is within the target sensitivity range, the exposure time is within the target time range, and the magnification is within the target magnification range; The hardware parameters include resolution, focal length, aperture size and field of view.

11. The image depth estimation method according to claim 10, wherein: The hardware parameters of the main camera and the auxiliary camera respectively include: The resolution of each of the main camera and the auxiliary camera is within the target resolution range, the focal length is within the target focal length range, the aperture size is within the target aperture range, and the field of view angle is within the target viewing angle range.

12. An electronic device, characterized in that: The electronic device includes a memory, a processor and a camera, the camera includes a main camera and an auxiliary camera, the memory and the camera are electrically connected to the processor, and when the processor executes computer instructions stored in the memory, it implements the image depth estimation method according to any one of claims 1 to 11.

13. A computer-readable storage medium, characterized in that Computer instructions are stored thereon, and when a processor executes the computer instructions, the image depth estimation method according to any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and computer readable storage medium

    CN111383255A

  • Image depth estimation method and electronic equipment

    CN117560480A

  • Image processing method, electronic equipment and storage medium

    CN118075607A

  • Image processing method and electronic equipment

    CN118447071A

  • Methods and systems for selective sensor fusion

    US20190273909A1