Image depth estimation method, electronic device, and computer-readable storage medium
Patent Information
- Application Number
- CN202411110628.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-13
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2044-08-13
AI Technical Summary
[0003]鉴于此,本申请实施例提供一种图像深度估计方法、电子设备及计算机可读存储介质,旨在解决如何准确估计图像深度信息的问题
[0021] It is understood that the beneficial effects of the electronic device provided in the second aspect of the embodiments of this application, the computer-readable storage medium provided in the third aspect, and the computer program product provided in the fourth aspect are substantially the same as the beneficial effects of the image depth estimation method provided in the first aspect, and will not be repeated here.
Smart Images

Figure CN120765710B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, specifically to an image depth estimation method, an electronic device, and a computer-readable storage medium. Background Technology
[0002] When electronic devices take photos, the background blur function can create a blurred image. In a blurred image, the foreground and subject are relatively clear, while the background is blurred, thus giving the photo a stronger artistic and aesthetic appeal. Background blur is achieved through image blur algorithms, which rely on the image's depth information. Inaccurate depth information can lead to poor background blur effects; for example, the foreground or subject may be blurry, or the background may be relatively clear. Summary of the Invention
[0003] In view of this, embodiments of this application provide an image depth estimation method, an electronic device, and a computer-readable storage medium, aiming to solve the problem of how to accurately estimate image depth information.
[0004] The first aspect of this application provides an image depth estimation method applied to an electronic device, which includes a main camera and an auxiliary camera. The method includes: when scene information meets the binocular estimation conditions, using a binocular depth calculation algorithm to obtain depth information; when scene information does not meet the binocular estimation conditions, using a monocular depth calculation algorithm to obtain depth information. The scene information includes environmental information and / or image information. The environmental information includes the hardware parameters, sensitivity, exposure time, and magnification of the main camera and the auxiliary camera, respectively. The image information includes texture information of the main path image, consistency information of the binocular images, and accuracy information of the binocular calibration. The binocular images include a main path image and an auxiliary path image, where the main path image is the image captured by the main camera, and the auxiliary path image is the image captured by the auxiliary camera.
[0005] In this embodiment, the main camera and the auxiliary camera are a combination of different cameras. For example, the main camera may be a wide-angle camera, and the auxiliary camera may be an ultra-wide-angle camera. Another example is that the main camera may be a telephoto camera, and the auxiliary camera may be a wide-angle camera. Scene information includes environmental information and / or image information. Environmental information reflects the attributes of the main camera and the auxiliary camera, while image information reflects the quality of the binocular images. If the environmental information satisfies the binocular estimation conditions, it means that the attributes of the main camera and the auxiliary camera meet the requirements of the binocular depth calculation algorithm. If the image information satisfies the binocular estimation conditions, it means that the quality of the binocular images meets the requirements of the binocular depth calculation algorithm. When the scene information satisfies the binocular estimation conditions, it indicates that the error rate of the binocular depth calculation algorithm is low. Since the depth information obtained by the binocular depth calculation algorithm represents the true distance between the subject and the camera, the depth information obtained using the binocular depth calculation algorithm is more accurate. When the scene information does not satisfy the binocular estimation conditions, it indicates that the error rate of the binocular depth calculation algorithm is high. The depth information obtained using a monocular depth calculation algorithm is more accurate, thus enabling the acquisition of more accurate depth information in different shooting scenarios, thereby improving the robustness of image depth estimation.
[0006] In one embodiment, when the scene information includes image information, and the scene information satisfies the binocular estimation conditions, the following are included: the texture information of the main path image includes a texture normality flag, the consistency information of the binocular images includes a consistency flag, and the accuracy information of the binocular calibration includes an accuracy flag. The texture normality flag indicates that the texture of the main path image is normal, the consistency flag indicates that the main path image and the auxiliary path image are consistent, and the accuracy flag indicates that the calibration of the main camera and the auxiliary camera is accurate.
[0007] In this embodiment, when the scene information includes image information, since the image information can reflect the quality of the input image of the binocular depth calculation algorithm, the error rate of the binocular depth calculation algorithm reflected by it is more accurate, thereby enabling a more accurate decision on whether to use a monocular depth calculation algorithm or a binocular depth calculation algorithm.
[0008] In another embodiment, the method further includes: when the texture complexity of the main road image is less than a texture complexity threshold, determining that the texture of the main road image is a weak texture, and the texture information of the main road image includes a texture anomaly flag. The texture anomaly flag is used to indicate texture anomalies in the main road image. Texture complexity is a value obtained by statistically calculating the texture features in the image using mathematical and statistical methods.
[0009] In another embodiment, the method further includes: when the texture complexity of the main path image is greater than or equal to a texture complexity threshold, and the proportion of target blocks in the main path image is greater than or equal to a proportion threshold, determining that the texture of the main path image is a repeating texture, and the texture information of the main path image includes a texture anomaly flag. Specifically, the similarity between the texture complexity of the target block and the texture complexity of the reference block is greater than a similarity threshold. The texture anomaly flag is used to indicate texture anomalies in the main path image. Texture complexity is a value obtained by statistically calculating the texture features in the image using mathematical and statistical methods.
[0010] In another embodiment, the method further includes: determining that the main image and the auxiliary image are consistent when the brightness difference between the main image and the auxiliary image is within a target brightness range, the sharpness difference is within a target sharpness range, the color temperature difference is within a target color temperature range, and the exposure time difference is within a target time difference range; wherein the consistency information of the binocular images includes a consistency flag. The exposure time difference is the time difference between the exposure time of the main image and the exposure time of the auxiliary image.
[0011] In another embodiment, the method further includes: when the error of the dual-camera calibration is less than an error threshold, determining that the calibration of the main camera and the auxiliary camera is accurate, wherein the accuracy information of the dual-camera calibration includes an accuracy flag.
[0012] In another embodiment, the method further includes: detecting feature point pairs of the main road image and the auxiliary road image, wherein the feature point pair includes feature points of the main road image and feature points of the corresponding location of the auxiliary road image; calculating the intrinsic and extrinsic parameters of the main camera and the auxiliary camera, wherein the intrinsic parameters include focal length, principal point, pixel size, and distortion coefficient, and the extrinsic parameters include rotation matrix and translation vector; obtaining reprojection feature point pairs based on the intrinsic and extrinsic parameters of the main camera and the auxiliary camera, wherein the reprojection feature point pairs include reprojection feature points of the main road image and reprojection feature points of the corresponding location of the auxiliary road image; and calculating the dual-target calibration error based on the feature point pairs and the corresponding reprojection feature point pairs.
[0013] In another embodiment, the method further includes: when the parallax value between the main road image and the auxiliary road image is less than 0, determining that the calibration of the main camera and the auxiliary camera is inaccurate, wherein the accuracy information of the dual-camera calibration includes an inaccuracy flag.
[0014] In another embodiment, the method further includes: when the number of feature point pairs between the main road image and the auxiliary road image is less than a number threshold, determining that the calibration of the main camera and the auxiliary camera is inaccurate, wherein the accuracy information of the dual-camera calibration includes an inaccurate flag, and the feature point pair includes feature points of the main road image and feature points of the corresponding location of the auxiliary road image.
[0015] In another embodiment, when the scene information includes environmental information, the scene information satisfies the binocular estimation conditions including: the hardware parameters of the main camera and the auxiliary camera are within the target parameter range, the sensitivity is within the target sensitivity range, the exposure time is within the target exposure time range, and the magnification is within the target magnification range. The hardware parameters include resolution, focal length, aperture size, and field of view.
[0016] In another embodiment, the hardware parameters of the main camera and the auxiliary camera within the target parameter range include: the resolution of the main camera and the auxiliary camera within the target resolution range, the focal length within the target focal length range, the aperture size within the target aperture range, and the field of view within the target viewing angle range.
[0017] In another embodiment, when the scene information includes environmental information and image information, the scene information satisfies the binocular estimation conditions including: the texture information of the main path image includes a texture normality indicator, the consistency information of the binocular image includes a consistency indicator, the accuracy information of the binocular estimation includes a consistency indicator, the hardware parameters of the main camera and the auxiliary camera are within the target parameter range, the sensitivity is within the target sensitivity range, the exposure time is within the target duration range, and the magnification is within the target magnification range.
[0018] The second aspect of this application provides an electronic device, which includes a memory, a processor, and a camera. The camera includes a main camera and an auxiliary camera. The memory and the camera are electrically connected to the processor. When the processor executes computer instructions stored in the memory, it implements the image depth estimation method provided in the first aspect.
[0019] A third aspect of this application provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the image depth estimation method provided in the first aspect.
[0020] A fourth aspect of this application provides a computer program product including computer instructions that, when executed by a processor, implement the image depth estimation method provided in the first aspect.
[0021] It is understood that the beneficial effects of the electronic device provided in the second aspect of the embodiments of this application, the computer-readable storage medium provided in the third aspect, and the computer program product provided in the fourth aspect are substantially the same as the beneficial effects of the image depth estimation method provided in the first aspect, and will not be repeated here. Attached Figure Description
[0022] Figure 1 This is a schematic diagram of the software structure of an electronic device provided as an example.
[0023] Figure 2 This is a schematic diagram of the interface of an electronic device provided as an example.
[0024] Figure 3 This is another example of an interface diagram for an electronic device.
[0025] Figure 4 This is a schematic diagram of the logical structure of an image blurring algorithm provided as an example.
[0026] Figure 5 This is a schematic diagram of the logical structure of a monocular depth calculation algorithm provided as an example.
[0027] Figure 6 This is a schematic diagram of the logical structure of a stereo depth calculation algorithm provided as an example.
[0028] Figure 7 This is a schematic diagram of the logical structure of a depth computing algorithm provided as an example.
[0029] Figure 8 This is a flowchart illustrating how an electronic device can implement a background blurring function.
[0030] Figure 9 This is another example of an interface diagram for an electronic device.
[0031] Figure 10 This is another example of an interface diagram for an electronic device.
[0032] Figure 11 This is a schematic diagram of a blurred image provided as an example.
[0033] Figure 12 This is a flowchart illustrating an image blurring algorithm implemented in an electronic device, as provided in this example.
[0034] Figure 13 This is a schematic diagram of the hardware structure of an electronic device provided as an example. Detailed Implementation
[0035] It should be noted that in the embodiments of this application, "multiple" and "several" refer to two or more. The terms "first," "second," "third," "fourth," etc., in the specification, claims, and drawings of this application are used to distinguish similar objects, not to describe a specific order or sequence.
[0036] It should also be noted that the methods disclosed in the embodiments of this application or the methods shown in the flowcharts include one or more steps for implementing the method. Without departing from the scope of the claims, the execution order of multiple steps can be interchanged, and some steps can also be deleted.
[0037] The electronic devices in this application include smartphones with camera functions, tablet computers, handheld computers, laptop computers, mobile internet devices (MID), virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, cellular phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), handheld devices with wireless communication functions, computing devices or other processing devices connected to a wireless modem, in-vehicle devices, wearable devices, terminal devices in 5G networks, or terminal devices in public land mobile networks (PLMNs), etc.
[0038] The software system of the electronic device is described in detail below.
[0039] The software system of electronic devices can adopt a layered architecture, which divides the software into several layers, each with a clear role and division of labor, and the layers communicate with each other through software interfaces. For example... Figure 1 As shown, the software system of an electronic device is divided into four layers, from top to bottom: Application (APP) layer, Framework (FWK) layer, Hardware Abstraction Layer (HAL) and Hardware layer.
[0040] The application layer comprises a series of application packages. One application package includes a camera application. This camera application features background blur functionality. For example, such as... Figure 2 As shown, when a user taps the camera app, the electronic device's screen displays the shooting interface. On this interface, when the user selects a large aperture mode, the camera app activates the background blur function. For example, as... Figure 3As shown, on the shooting interface, when the user selects portrait mode, the camera application enables the background blur function.
[0041] Background blur function can blur the background in an image, thereby highlighting the subject and foreground, and making the content of the image more layered. An image includes the subject, foreground, and background. The subject is the object that the camera is focusing on, such as a person, animal, or landscape. The part of the image that is in front of the subject is called the foreground, and the part that is behind the subject is called the background.
[0042] The framework layer provides application programming interfaces (APIs) and programming frameworks for various apps in the application layer. The framework layer includes the camera framework (CameraFWK), which provides the camera API (CameraAPI) for camera applications.
[0043] The hardware abstraction layer is used to execute various application functions and data processing in apps. The hardware abstraction layer includes the camera hardware abstraction layer (CameraHAL), which comprises the processor driver, camera driver, and camera algorithm library.
[0044] A processor driver is a processor-oriented control node used to control the processor to implement image processing algorithms.
[0045] A camera driver is a control node for cameras, used to control the camera to capture images.
[0046] The camera algorithm library stores a series of image processing algorithms, such as image blurring, automatic exposure (AE), automatic focus (AF), electronic image stabilization (EIS), and image quality (PQ).
[0047] The hardware layer provides hardware support for the various applications and data processing of the hardware abstraction layer. The hardware layer includes processors and cameras. Processors include a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), and a Neural Processing Unit (NPU). The CPU handles control logic and serial computation tasks. The GPU handles image algorithm logic and parallel computation tasks. The NPU handles neural network computation tasks. Cameras include wide-angle, ultra-wide-angle, and telephoto cameras. Each camera has a corresponding field of view (FOV) and focal length.
[0048] For example, when using an electronic device to photograph a distant subject, the zoom level can be increased to bring the electronic device closer to the subject. The electronic device uses cameras with different focal lengths to take photos / videos depending on the zoom level and the subject distance, resulting in higher image clarity. During the zoom process, the electronic device switches between different cameras when the zoom level is greater than or equal to a zoom threshold. For instance, as the zoom level increases from its minimum value, the electronic device might sequentially use an ultra-wide-angle camera, a wide-angle camera, and a telephoto camera.
[0049] In different shooting modes, electronic devices can use different combinations of cameras. For example, in large aperture mode or portrait mode, when the zoom level is less than a certain threshold, the electronic device uses a wide-angle camera and an ultra-wide-angle camera, with the wide-angle camera being the main camera and the ultra-wide-angle camera being the secondary camera. When the zoom level is greater than or equal to the threshold, the electronic device uses a telephoto camera and a wide-angle camera, with the telephoto camera being the main camera and the wide-angle camera being the secondary camera. The image captured by the main camera is called the primary path image, and the image captured by the secondary camera is called the secondary path image.
[0050] exist Figure 1 In the camera algorithm library shown, the image blurring algorithm acquires the depth information of the image and then blurs the background of the main image based on this depth information. Depth information represents the distance of the object being photographed from the camera in three-dimensional space. This depth information can be called a depth value, and an image containing this depth information is called a depth map.
[0051] For example, such as Figure 4As shown, the logical structure of the image blurring algorithm includes a depth calculation module and a blurring rendering module. The depth calculation module performs depth calculations on the image to obtain a depth map. The blurring rendering module blurs the background of the main image based on the depth map, resulting in a blurred image.
[0052] The depth calculation module includes a monocular depth calculation module and a binocular depth calculation module. The monocular depth calculation module performs depth calculations on the monocular image to obtain a depth map, where the monocular image is the main road image. The binocular depth calculation module performs depth calculations on the binocular image to obtain a depth map, where the binocular image includes both the main road image and the auxiliary road image.
[0053] For example, such as Figure 5 As shown, the monocular depth calculation module includes a front-end processing module and a monocular depth estimation model. The front-end processing module is used to preprocess the main path image to obtain a preprocessed image. Preprocessing may include image cropping, and the size of the cropped preprocessed image meets the requirements of the monocular depth estimation model for the input image size. Preprocessing may also include image enhancement, noise removal, geometric transformation, etc., to adapt to different image processing needs.
[0054] Monocular depth estimation models are used to calculate depth in preprocessed images to obtain depth maps. Monocular depth estimation models can employ Dense Prediction Transformer (DPT), MiDaS, Adabins, PackNet-SfM, and other models.
[0055] For example, such as Figure 6 As shown, the binocular depth calculation module includes a front-end processing module, an epipolar correction module, and a binocular depth estimation model. The front-end processing module preprocesses the main road image and the auxiliary road image to obtain preprocessed image pairs, which contain preprocessed images of the main road and auxiliary roads. Preprocessing may include image cropping, ensuring the size of the cropped preprocessed image pairs meets the input image size requirements of the binocular depth estimation model. Preprocessing may also include image enhancement, noise removal, and geometric transformations to adapt to different image processing needs.
[0056] The epipolar correction module performs epipolar correction on preprocessed image pairs, resulting in corrected image pairs, which include a main path corrected image and an auxiliary path corrected image. Epipolar correction aligns the scan lines of the image with the epipolar lines. The epipolar correction module includes a static correction module and a dynamic correction module. The static correction module calculates the correspondence between preprocessed image pairs based on the results of dual-target calibration using a fundamental matrix or an essential matrix, aligning corresponding pixel pairs in the main path and auxiliary path preprocessed images on the same horizontal line. The dynamic correction module updates the correspondence between preprocessed image pairs based on corresponding feature point pairs in the preprocessed image pairs, further aligning corresponding pixel pairs in the main path and auxiliary path preprocessed images on the same horizontal line. This adapts the correspondence between preprocessed image pairs to changes in the position of the photographed object, thus meeting the requirements of epipolar correction.
[0057] Binocular depth estimation models are used to calculate the depth of calibrated image pairs to obtain depth maps. Binocular depth estimation models can include DispNetC, GC-Net, Pyramid StereoMatching Network (PSMNet), StereoDRNet, LW-Stereo, and DSMNet.
[0058] It can be understood that the depth information obtained by the monocular depth calculation module represents the distance between the subject and the camera, while the depth information obtained by the binocular depth calculation module represents the actual distance between the subject and the camera.
[0059] In this embodiment, as Figure 7 As shown, the depth calculation module includes a decision module, which further comprises a monocular depth calculation module and a binocular depth calculation module. The decision module determines whether to use the monocular or binocular depth calculation module based on scene information. Specifically, when the scene information meets the binocular estimation conditions, the binocular depth calculation module is used. When the scene information does not meet the binocular estimation conditions, the monocular depth calculation module is used. Scene information includes environmental information and / or image information. Environmental information includes the hardware parameters, sensitivity (also known as ISO value), exposure time, and magnification of the main and auxiliary cameras. Image information includes texture information of the main image, consistency information of the binocular images, and accuracy information of the binocular calibration.
[0060] The texture information of the main path image includes a normal texture flag or a texture abnormality flag. The normal texture flag indicates that the texture of the main path image is normal. The texture abnormality flag indicates that the texture of the main path image is abnormal. The specific forms of the normal and abnormal texture flags can be set as needed. For example, the normal texture flag can be "true" and the abnormal texture flag can be "false". Another example is that the normal texture flag can be "1" and the abnormal texture flag can be "0" or "null".
[0061] When the texture of the main road image is determined to be an abnormal texture, the texture information of the main road image includes a texture abnormality flag; otherwise, the texture information of the main road image includes a texture normality flag. For example, abnormal textures include weak textures and repetitive textures. Repetitive textures are textures with regularity in the image, while weak textures are textures without obvious regularity in the image. When the texture complexity of the main road image is less than a texture complexity threshold, the texture of the main road image is determined to be a weak texture. When the texture complexity of the main road image is greater than or equal to a texture complexity threshold, and the proportion of target patches in the main road image is greater than or equal to a proportion threshold, the texture of the main road image is determined to be a repetitive texture. When the similarity between the texture complexity of the i-th patch and the texture complexity of the reference patch is greater than a similarity threshold, the i-th patch is determined to be a target patch, where i is a positive integer. The reference patch can be any patch in the main road image. Texture complexity is a numerical value obtained by statistically calculating the texture features in the image using mathematical and statistical methods. Texture complexity is used to measure the richness and detail of the texture in the image.
[0062] It is understandable that texture complexity thresholds, ratio thresholds, and similarity thresholds can be set as needed. Image texture detection methods may include using Local Binary Pattern (LBP) histograms to detect texture features, using wavelet transform / Fourier transform to detect texture features, or using Gabor filters to detect texture features, etc.
[0063] The consistency information for binocular images includes consistency flags and inconsistency flags. Consistency flags indicate that the main road image and the auxiliary road image are consistent. Inconsistency flags indicate that the main road image and the auxiliary road image are inconsistent. The specific forms of consistency and inconsistency flags can be set as needed. For example, the consistency flag can be "true" and the inconsistency flag can be "false". Another example is that the consistency flag can be "1" and the inconsistency flag can be "0" or "null".
[0064] When the main road image and the auxiliary road image are determined to be consistent, the consistency information of the binocular images includes a consistency flag; otherwise, the consistency information of the binocular images includes an inconsistency flag. For example, when the brightness difference between the main road image and the auxiliary road image is within the target brightness range, the sharpness difference is within the target sharpness range, the color temperature difference is within the target color temperature range, and the exposure time difference is within the target time difference range, the main road image and the auxiliary road image are determined to be consistent. The exposure time difference is the time difference between the exposure time of the main road image and the exposure time of the auxiliary road image.
[0065] It is understandable that the target brightness range, target sharpness range, target color temperature range, and target time difference range can be set as needed. The image brightness value can be calculated using the `calculate_brightness` function in OpenCV, and the image sharpness can be calculated using the `variance_of_laplacian` function. The image color temperature can be estimated using a pre-trained regression model, and the exposure time can be obtained from the camera's image sensor. Alternatively, the image sharpness can be calculated using the No-Reference Statistical Sharpness (NRSS) function, and the image color temperature can be estimated using the Correlated Color Temperature (CCT) function.
[0066] Dual-camera calibration accuracy information includes an accuracy flag or an inaccuracy flag. An accuracy flag indicates that the calibration of the main and auxiliary cameras is accurate. An inaccuracy flag indicates that the calibration of the main and auxiliary cameras is inaccurate. The specific format of the accuracy and inaccuracy flags can be set as needed. For example, the accuracy flag can be "true" and the inaccuracy flag can be "false". Another example is that the accuracy flag can be "1" and the inaccuracy flag can be "0" or "null".
[0067] When the calibration of the main camera and the auxiliary camera is accurate, the accuracy information of the dual-camera calibration includes an accuracy flag; otherwise, the accuracy information includes an inaccuracy flag. For example, when the dual-camera calibration error is less than an error threshold, the calibration of the main camera and the auxiliary camera is determined to be accurate. Calculating the dual-camera calibration error includes detecting feature point pairs in the main road image and the auxiliary road image, calculating the intrinsic and extrinsic parameters of the main camera and the auxiliary camera, obtaining reprojected feature point pairs based on the intrinsic and extrinsic parameters of the main camera and the auxiliary camera, and calculating the dual-camera calibration error based on the feature point pairs and the corresponding reprojected feature point pairs. The feature point pairs include feature points in the main road image and feature points in the corresponding auxiliary road image. Reprojection is the reprojection of known reference feature points in the world coordinate system to the image coordinate system; the reprojected feature point pairs include reprojected feature points in the main road image and reprojected feature points in the corresponding auxiliary road image. Intrinsic parameters represent the internal optical and geometric characteristics of the camera, including focal length, principal point, pixel size, and distortion coefficients. The principal point is the intersection of the optical axis and the image plane. The extrinsic parameters represent the relative position and orientation between the main camera and the auxiliary camera. The extrinsic parameters include the rotation matrix and the translation vector. The rotation matrix represents the rotation relationship between the main camera coordinate system and the auxiliary camera coordinate system, and the translation vector represents the translation relationship between the main camera coordinate system and the auxiliary camera coordinate system.
[0068] In one embodiment, when the disparity value between the main path image and the auxiliary path image is less than 0, it is determined that the calibration of the main camera and the auxiliary camera is inaccurate. The disparity value is the difference between the column coordinates of a pixel in the main path image and the column coordinates of the corresponding pixel in the auxiliary path image.
[0069] In another embodiment, when the number of feature point pairs in the main road image and the auxiliary road image is less than a threshold, it is determined that the calibration of the main camera and the auxiliary camera is inaccurate.
[0070] It is understood that the error threshold and number threshold in the above embodiments can be set as needed.
[0071] In this embodiment, when the scene information includes image information, the scene information satisfies the binocular estimation conditions as follows: the texture information of the main path image includes a texture normality flag, the consistency information of the binocular image includes a consistency flag, and the accuracy information of the binocular estimation includes a consistency flag.
[0072] When scene information includes environmental information, the scene information satisfies the binocular estimation conditions, including: the hardware parameters of the main camera and the auxiliary camera are within the target parameter range, the sensitivity is within the target sensitivity range, the exposure time is within the target duration range, and the magnification is within the target magnification range.
[0073] It is understood that the target parameter range, target sensitivity range, target duration range, and target magnification range can be set as needed. For example, the camera's hardware parameters include resolution, focal length, aperture size, and field of view. The camera's hardware parameters within the target parameter range include: resolution within the target resolution range, focal length within the target focal length range, aperture size within the target aperture range, and field of view within the target angle of view range. The target resolution range, target focal length range, target aperture range, and target angle of view range can be set as needed.
[0074] When scene information includes environmental information and image information, the scene information satisfies the binocular estimation conditions, including: the texture information of the main image includes texture normality markers, the consistency information of the binocular images includes consistency markers, the accuracy information of binocular estimation includes consistency markers, the hardware parameters of the main camera and the auxiliary camera are within the target parameter range, the sensitivity is within the target sensitivity range, the exposure time is within the target duration range, and the magnification is within the target magnification range.
[0075] The logical structure of the image blurring algorithm has been explained in detail above. In this embodiment, a decision module is added to the image blurring algorithm to determine whether to use a monocular depth calculation module or a binocular depth calculation module based on scene information. When the scene information meets the binocular estimation conditions, it indicates that the error rate of the binocular depth calculation module is low. Since the depth information obtained by the binocular depth calculation module represents the true distance between the subject and the camera, the depth information obtained by the binocular depth calculation module is more accurate, thus presenting a better background blurring effect. When the scene information does not meet the binocular estimation conditions, it indicates that the error rate of the binocular depth calculation module is high. The depth information obtained by the monocular depth calculation module is more accurate, thus enabling the acquisition of more accurate depth information in different shooting scenarios, thereby improving the robustness of image depth estimation, making the background blurring function of the camera application more stable, and thus improving the user experience. Furthermore, when the scene information includes image information, since the image information can reflect the quality of the input image of the binocular depth calculation module, it reflects the error rate of the binocular depth calculation module more accurately, thus enabling a more accurate decision on whether to use a monocular depth calculation module or a binocular depth calculation module.
[0076] The following is combined Figure 1 The software system of the electronic device shown is described in detail, outlining the implementation process of the background blurring function.
[0077] like Figure 8 As shown, the implementation process of the background blur function includes the following steps:
[0078] S101. When the camera application's shooting mode is target mode and the shooting control is triggered, the camera application sends a shooting command to the CameraAPI.
[0079] Target mode is a shooting mode that enables background blur. Camera apps offer various shooting modes, such as wide aperture mode, portrait mode, and night mode. For example, target mode includes wide aperture mode and portrait mode. Figure 2 and Figure 3 As shown, when a user taps the camera app, the electronic device's screen displays the shooting interface. On the shooting interface, when the user selects a large aperture mode or portrait mode, the camera app activates the background blur function.
[0080] The shooting control is a control that controls the shooting process. For example, such as... Figure 9 As shown, when a user clicks the shooting control 11 on the shooting interface, the shooting control is triggered. For example, as... Figure 10 As shown, when a user presses the volume button 12 on the electronic device, the shooting control is triggered. The volume button 12 includes a volume up button and a volume down button.
[0081] The shooting command is used to instruct the user to shoot in target mode.
[0082] S102, CameraAPI sends shooting commands to the camera driver.
[0083] S103, the camera driver responds to the shooting command and controls the camera to capture images.
[0084] The camera includes a main camera and an auxiliary camera, and the images include images of the main road and images of the auxiliary road.
[0085] S104: The processor driver acquires images from the camera driver, calls the image blurring algorithm from the camera algorithm library, and controls the CPU, GPU, and GPU to process the images to obtain a blurred image.
[0086] In this embodiment, the image blurring algorithm is nested with the depth calculation algorithm and the blurring rendering algorithm, the depth calculation algorithm is nested with the decision algorithm, and the decision algorithm is nested with the monocular depth calculation algorithm and the binocular depth calculation algorithm.
[0087] It is understandable that deep learning algorithms and Figure 4 The depth calculation module shown corresponds to the blurring rendering algorithm and Figure 4 The blurred rendering module shown corresponds to the decision algorithm and Figure 7 The decision module shown corresponds to the monocular depth calculation algorithm and Figure 5 The monocular depth calculation module shown corresponds to the binocular depth calculation algorithm. Figure 6 The binocular depth calculation module shown corresponds to this.
[0088] S105, The processor driver sends a blurred image to the Camera API.
[0089] S106, CameraAPI sends blurred images to camera applications.
[0090] S107, The camera application stores the blurred image to the gallery.
[0091] For example, such as Figure 11 As shown, in a blurred image, the subject and foreground are relatively clear, while the background is relatively blurry.
[0092] The above provides a detailed explanation of the background blurring function's implementation process. The implementation of the background blurring function relies on an image blurring algorithm; the following section describes the implementation process of this algorithm.
[0093] like Figure 12 As shown, the image blurring algorithm includes the following steps:
[0094] S201. Obtain scene information.
[0095] Scene information includes environmental information and / or image information. Environmental information includes the hardware parameters, sensitivity, exposure time, and magnification of the main and auxiliary cameras. Image information includes texture information of the main image, consistency information of the binocular images, and accuracy information of the binocular positioning.
[0096] The texture information of the main road image includes a normal texture flag or a texture abnormality flag. The normal texture flag indicates that the texture of the main road image is normal. The texture abnormality flag indicates that the texture of the main road image is abnormal.
[0097] The consistency information of binocular images includes consistency flags and inconsistency flags. Consistency flags indicate that the main road image and the auxiliary road image are consistent. Inconsistency flags indicate that the main road image and the auxiliary road image are inconsistent.
[0098] Dual-camera calibration accuracy information includes an accuracy flag or an inaccuracy flag. An accuracy flag indicates that the calibration of the main and auxiliary cameras is accurate. An inaccuracy flag indicates that the calibration of the main and auxiliary cameras is inaccurate.
[0099] S202. Determine whether the scene information meets the conditions for binocular estimation.
[0100] If yes, proceed to step S203; otherwise, proceed to step S204.
[0101] After completing step S203 or S204, proceed to step S205.
[0102] S203. When the scene information meets the conditions for binocular estimation, use the binocular depth calculation algorithm to obtain the depth map.
[0103] When scene information includes image information, the scene information satisfies the binocular estimation conditions, including: texture information of the main path image including texture normality markers, consistency information of the binocular images including consistency markers, and accuracy information of binocular estimation including accuracy markers.
[0104] When scene information includes environmental information, the scene information satisfies the binocular estimation conditions, including: the hardware parameters of the main camera and the auxiliary camera are within the target parameter range, the sensitivity is within the target sensitivity range, the exposure time is within the target duration range, and the magnification is within the target magnification range.
[0105] When scene information includes environmental information and image information, the scene information satisfies the binocular estimation conditions, including: texture information of the main image including texture normality markers, consistency information of the binocular images including consistency markers, accuracy information of binocular estimation including accuracy markers, hardware parameters of the main camera and auxiliary camera within the target parameter range, sensitivity within the target sensitivity range, exposure time within the target exposure time range, and magnification within the target magnification range.
[0106] The binocular depth calculation algorithm includes: preprocessing the main road image and the auxiliary road image to obtain preprocessed image pairs, each containing the main road preprocessed image and the auxiliary road preprocessed image; performing epipolar correction on the preprocessed image pairs to obtain corrected image pairs, each containing the main road corrected image and the auxiliary road corrected image; and performing depth calculation on the corrected image pairs to obtain a depth map.
[0107] Epipolar correction includes static correction and dynamic correction. Static correction involves calculating the correspondence between preprocessed image pairs using a fundamental or essential matrix based on the results of dual-target calibration, aligning corresponding pixel pairs in the main path and auxiliary path preprocessed images on the same horizontal line. Dynamic correction involves updating the correspondence between preprocessed image pairs based on corresponding feature point pairs, further aligning corresponding pixel pairs in the main path and auxiliary path preprocessed images on the same horizontal line. This adapts the correspondence between preprocessed image pairs to changes in the position of the subject, thus meeting the requirements of epipolar correction.
[0108] S204. When the scene information does not meet the conditions for binocular estimation, use a monocular depth calculation algorithm to obtain a depth map.
[0109] When scene information includes image information, the scene information that does not meet the binocular estimation conditions includes: texture information of the main path image including texture anomaly markers, and / or consistency information of the binocular images including inconsistency markers, and / or accuracy information of the binocular estimation including inaccuracy markers.
[0110] When scene information includes environmental information, the scene information does not meet the binocular estimation conditions, including: the hardware parameters of the main camera and the auxiliary camera are not within the target parameter range, and / or the sensitivity is not within the target sensitivity range, and / or the exposure time is not within the target duration range, and / or the magnification is not within the target magnification range.
[0111] When scene information includes environmental information and image information, the scene information that does not meet the binocular estimation conditions includes: texture information of the main path image including texture anomaly markers, and / or consistency information of the binocular images including inconsistency markers, and / or accuracy information of the binocular estimation including inaccuracy markers, and / or the hardware parameters of the main camera and the auxiliary camera are not within the target parameter range, and / or the sensitivity is not within the target sensitivity range, and / or the exposure time is not within the target duration range, and / or the magnification is not within the target magnification range.
[0112] The monocular depth calculation algorithm includes: preprocessing the main path image to obtain a preprocessed image; and performing depth calculation on the preprocessed image to obtain a depth map.
[0113] S205. Blur the background of the main road image based on the depth map to obtain a blurred image.
[0114] For example, the blurring process can be performed using Gaussian blur.
[0115] The implementation process of the image blurring algorithm has been explained in detail above. Steps S201-S204 described above constitute the image depth estimation method of this embodiment. In this embodiment, a decision algorithm is added to the image blurring algorithm to determine whether to use a monocular depth calculation algorithm or a binocular depth calculation algorithm based on scene information. When the scene information meets the binocular estimation conditions, it indicates that the error rate of the binocular depth calculation algorithm is low. Since the depth information obtained by the binocular depth calculation algorithm represents the true distance between the subject and the camera, the depth information obtained by using the binocular depth calculation algorithm is more accurate, thus presenting a better background blurring effect. When the scene information does not meet the binocular estimation conditions, it indicates that the error rate of the binocular depth calculation algorithm is high. The depth information obtained by using the monocular depth calculation algorithm is more accurate, thus enabling the acquisition of more accurate depth information in different shooting scenarios, thereby improving the robustness of image depth estimation, making the background blurring function of the camera application more stable, and ultimately improving the user experience. Furthermore, when scene information includes image information, since image information can reflect the quality of the input image for the binocular depth calculation algorithm, it reflects the error rate of the binocular depth calculation algorithm more accurately. Therefore, it is possible to make a more accurate decision on whether to use a monocular depth calculation algorithm or a binocular depth calculation algorithm.
[0116] The hardware structure of the electronic device is described in detail below.
[0117] like Figure 13 As shown, the electronic device includes a processor 110, an external memory interface 120, an internal memory 121, a sensor module 130, a display screen 140, and a camera module 150. The sensor module 130 includes a touch sensor 131 and an image sensor 132. The camera module 150 may include a wide-angle camera 151, an ultra-wide-angle camera 152, and a telephoto camera 153.
[0118] The processor 110 is used to execute the various functions or steps performed by the electronic device in the above embodiments. The processor 110 includes a CPU, a GPU, and a GPU. The processor 110 may also include an application processor (AP), an image signal processor (ISP), a digital signal processor (DSP), a modem processor, and a video codec, etc.
[0119] The external memory interface 120 is used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor 110 through the external memory interface 120 to perform data storage functions.
[0120] Internal memory 121 stores executable program code, including instructions. Processor 110 executes the various functions or steps performed by the electronic device in the above embodiments by running the instructions stored in internal memory 121. Internal memory 121 includes a program storage area and a data storage area. The program storage area may store the operating system, at least one application required for a function, etc. The data storage area may store data created during the use of the electronic device. Internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, and Universal Flash Storage (UFS), etc.
[0121] The display screen 140 is used to display an interface. The display screen 140 includes a display panel. The display panel may be a liquid crystal display (LCD), a light-emitting diode (LED), or an organic light-emitting diode (OLED), etc.
[0122] If the display screen 140 integrates the touch sensor 131, then the display screen 140 can be referred to as a touch screen. The touch sensor 131 can also be referred to as a "touch panel". That is, the display screen 140 may include a display panel and a touch panel. The touch sensor 131 is used to detect touch operations applied to or near it. After the touch sensor 131 detects a touch operation, it can trigger the driver of the hardware abstraction layer of the electronic device to periodically scan the touch parameters generated by the touch operation. Then, the driver of the hardware abstraction layer sends the touch parameters to the relevant modules in the upper layer so that the relevant modules can determine the touch event corresponding to the touch parameters.
[0123] Electronic devices can achieve display functions through GPU, display screen 140 and AP, and can achieve shooting functions through GPU, ISP, camera module 150, image sensor 132, video codec, display screen 140 and AP.
[0124] Electronic devices can capture light through any one or more cameras in camera module 150 and transmit the light signal to image sensor 132. The light signal is converted into a raw image (or RAW image) by image sensor 132. Then, the RAW image is converted into a YUV image by ISP, and then the YUV image is converted into an RGB image, thus completing the image acquisition. Here, RAW is a raw image data format that has not been compressed or modified. YUV is a color encoding format, where Y represents luminance, and U and V represent chrominance. An RGB image is an image composed of three color channels: red, green, and blue. RGB images can be directly displayed on display screen 140 or subjected to subsequent editing and processing.
[0125] It is understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device. In other embodiments, the electronic device may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements.
[0126] The functions or steps performed by the electronic device in the above embodiments can also be applied to chips, computer-readable storage media, or computer program products.
[0127] The chip includes a processor and an interface circuit, with the processor and interface circuit electrically connected. The interface circuit can read computer instructions stored in the memory and send the computer instructions to the processor. When the processor executes the computer instructions, it implements the various functions or steps performed by the electronic device in the above embodiments.
[0128] The computer-readable storage medium stores computer instructions, which, when executed by a processor, implement the various functions or steps performed by the electronic device in the above embodiments.
[0129] Computer-readable storage media include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules or other data). Computer-readable storage media include, but are not limited to, random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other storage, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tapes, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer.
[0130] The computer program product includes computer instructions, which, when executed by a processor, implement the various functions or steps performed by the electronic device in the above embodiments.
[0131] The embodiments of this application have been described in detail above with reference to the accompanying drawings. However, this application is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of this application.
Claims
1. An image depth estimation method applied to an electronic device, the electronic device comprising a main camera and an auxiliary camera, characterized in that, The method includes: When the scene information meets the conditions for stereo estimation, use a stereo depth calculation algorithm to obtain depth information. When the scene information does not meet the conditions for binocular estimation, a monocular depth calculation algorithm is used to obtain depth information. The scene information includes image information; the image information includes texture information of the main path image, consistency information of the binocular images, and accuracy information of the binocular positioning; the binocular images include the main path image and the auxiliary path image, the main path image being the image captured by the main camera, and the auxiliary path image being the image captured by the auxiliary camera; The texture information of the main road image includes a normal texture flag or an abnormal texture flag. When the texture of the main road image is determined to be an abnormal texture, the texture information of the main road image includes the abnormal texture flag; otherwise, the texture information of the main road image includes the normal texture flag. The abnormal texture includes weak textures and repetitive textures. When the texture complexity of the main road image is less than the texture complexity threshold, the texture of the main road image is determined to be a weak texture. When the texture complexity of the main road image is greater than or equal to the texture complexity threshold, and the proportion of the number of target blocks in the main road image is greater than or equal to the proportion threshold, the texture of the main road image is determined to be a repetitive texture. The consistency information of the binocular images includes consistency flags or inconsistency flags; The accuracy information of the dual-target calibration includes an accuracy indicator or an inaccuracy indicator; The scene information that satisfies the binocular estimation conditions includes: The texture information of the main road image includes the texture normality flag, the consistency information of the binocular image includes the consistency flag, and the accuracy information of the binocular calibration includes the accuracy flag; the texture normality flag is used to indicate that the texture of the main road image is normal, the consistency flag is used to indicate that the main road image and the auxiliary road image are consistent, and the accuracy flag is used to indicate that the calibration of the main camera and the auxiliary camera is accurate.
2. The image depth estimation method as described in claim 1, characterized in that, The texture complexity is a value obtained by statistically calculating the texture features in the image using mathematical and statistical methods.
3. The image depth estimation method as described in claim 2, characterized in that, The similarity between the texture complexity of the target block and the texture complexity of the reference block is greater than a similarity threshold.
4. The image depth estimation method according to any one of claims 1-3, characterized in that, The method further includes: When the brightness difference between the main image and the auxiliary image is within the target brightness range, the sharpness difference is within the target sharpness range, the color temperature difference is within the target color temperature range, and the exposure time difference is within the target time difference range, it is determined that the main image and the auxiliary image are consistent. The consistency information of the binocular image includes the consistency flag. The exposure time difference is the time difference between the exposure time of the main path image and the exposure time of the auxiliary path image.
5. The image depth estimation method according to any one of claims 1-3, characterized in that, The method further includes: When the calibration error of the dual-camera system is less than the error threshold, it is determined that the calibration of the main camera and the auxiliary camera is accurate, and the accuracy information of the dual-camera system includes the accuracy flag.
6. The image depth estimation method as described in claim 5, characterized in that, The method further includes: Detect feature point pairs between the main road image and the auxiliary road image, wherein the feature point pair includes feature points of the main road image and feature points of the auxiliary road image at the corresponding positions; Calculate the intrinsic and extrinsic parameters of the main camera and the auxiliary camera. The intrinsic parameters include focal length, principal point, pixel size and distortion coefficient. The extrinsic parameters include rotation matrix and translation vector. Reprojection feature point pairs are obtained based on the intrinsic and extrinsic parameters of the main camera and the auxiliary camera. The reprojection feature point pairs include the reprojection feature points of the main road image and the reprojection feature points of the auxiliary road image at the corresponding positions. The error of the bi-target calibration is calculated based on the feature point pair and the corresponding reprojected feature point pair.
7. The image depth estimation method according to any one of claims 1-3 and 6, characterized in that, The method further includes: When the parallax value between the main road image and the auxiliary road image is less than 0, it is determined that the calibration of the main camera and the auxiliary camera is inaccurate, and the accuracy information of the dual-camera calibration includes the inaccuracy flag.
8. The image depth estimation method according to any one of claims 1-3 and 6, characterized in that, The method further includes: When the number of feature point pairs in the main road image and the auxiliary road image is less than a threshold, it is determined that the calibration of the main camera and the auxiliary camera is inaccurate. The accuracy information of the dual-camera calibration includes the inaccuracy flag. The feature point pair includes the feature point of the main road image and the feature point of the corresponding position of the auxiliary road image.
9. The image depth estimation method according to any one of claims 1-3 and 6, characterized in that, The scene information also includes environmental information, which includes the hardware parameters, sensitivity, exposure time and magnification of the main camera and the auxiliary camera, respectively. The scene information that satisfies the binocular estimation conditions also includes: The hardware parameters of the main camera and the auxiliary camera are within the target parameter range, the sensitivity is within the target sensitivity range, the exposure time is within the target duration range, and the magnification is within the target magnification range. The hardware parameters include resolution, focal length, aperture size, and field of view.
10. The image depth estimation method as described in claim 9, characterized in that, The hardware parameters of the main camera and the auxiliary camera, within the target parameter range, include: The resolution of the main camera and the auxiliary camera are within the target resolution range, the focal length is within the target focal length range, the aperture size is within the target aperture range, and the field of view is within the target angle of view range.
11. An electronic device, characterized in that, The electronic device includes a memory, a processor, and a camera. The camera includes a main camera and an auxiliary camera. The memory and the camera are electrically connected to the processor. When the processor executes computer instructions stored in the memory, it implements the image depth estimation method as described in any one of claims 1-10.
12. A computer-readable storage medium, characterized in that, It stores computer instructions, which, when executed by the processor, implement the image depth estimation method as described in any one of claims 1-10.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and computer readable storage medium
CN111383255A
Image depth estimation method and electronic equipment
CN117560480A
Image processing method and electronic equipment
CN118447071A
Methods and systems for selective sensor fusion
US20190273909A1