Depth of field simulation method via optical-based depth extraction

By combining quantized cross-correlation and cost-capacity resampling techniques with filtering, the problem of poor depth-of-field simulation in mobile phone cameras was solved, achieving high-quality depth-of-field simulation and accelerated image processing.

CN113920181BActive Publication Date: 2025-11-11BLACK SESAME TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111179231.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-03
Filing Date
2021-10-08
Publication Date
2025-11-11
Estimated Expiration
2041-10-08

AI Technical Summary

Technical Problem

Mobile phone cameras do not simulate depth of field as well as digital SLR cameras, resulting in lower image quality. Existing technologies also affect the quality of depth maps and are slow in processing color deviations and lighting changes.

Method used

By receiving and processing two images, utilizing quantized cross-correlation operations and cost-capacity resampling, combined with edge filtering and smoothing filtering, the fully focused and out-of-focus layers are separated to achieve depth estimation and depth-of-field simulation.

Benefits of technology

It improves the depth-of-field simulation effect of mobile phone cameras, making the image quality close to that of digital SLR cameras, preserving image details and fine structure, while speeding up image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113920181B_ABST
    Figure CN113920181B_ABST
Patent Text Reader

Abstract

A depth-of-field simulation method includes receiving multiple images, predicting the mask layers of interest for the multiple images, determining multiple average brightness anchor values ​​for the window arrays of the corresponding multiple layers of interest, setting a set of binary codes for the multiple layers of interest, determining the Hamming distance between the set of binary codes for the multiple layers of interest, determining the cost capacity based on the Hamming distance, resampling the vertical cost based on the vertical ordinal direction of the cost capacity, resampling the horizontal cost based on the horizontal ordinal direction of the cost capacity, determining the full-focus layer based on the vertical cost and the horizontal cost, determining the defocus layer based on the vertical cost and the horizontal cost, and determining the depth of the full-focus layer and the depth of the defocus layer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to depth-of-field simulation, and more particularly to depth-of-field simulation via optical-based depth extraction. Background Technology

[0002] Currently, mobile phone cameras have become the primary method of taking pictures, surpassing previous digital single-lens reflex (DSLR) cameras. The image quality of mobile phone cameras may be considered inferior to that of DSLR cameras, largely due to the difference in depth of field (DOF) between the two. DSLR cameras are equipped with lenses with long focal lengths, which naturally blur out-of-focus areas and draw the viewer's attention to the subject of the image. Currently, mobile phones use pinhole cameras to capture images, which preserve high-frequency components in the background, resulting in less appealing photos. Summary of the Invention

[0003] An example method for depth-of-field simulation includes receiving a first image; receiving a second image; predicting an interest mask layer of the first image; performing a quantized cross-correlation operation on the first image, including determining a first average brightness anchor value of a first window array of the interest mask layer of the first image, and setting a first set of binary codes for the first window array based on the first average brightness anchor value; predicting an interest mask layer of the second image; performing a quantized cross-correlation operation on the second image, including determining a second average brightness anchor value of a second window array of the interest mask layer of the second image, and setting a second set of binary codes for the second window array based on the second average brightness anchor value; determining a Hamming distance between the first set of binary codes and the second set of binary codes; determining a cost volume based on the Hamming distance; resampling the vertical cost based on the vertical ordinal direction of the cost volume; resampling the horizontal cost based on the horizontal ordinal direction of the cost volume; determining a full-focus layer based on the vertical cost and the horizontal cost; determining a defocus layer based on the vertical cost and the horizontal cost; and determining the depth of the full-focus layer and the depth of the defocus layer.

[0004] Another depth-of-field simulation method includes receiving an initial cost capacity of a first image, determining an interest mask layer of the first image, determining the average cost capacity of the interest mask layer, separating the interest mask layer from the remaining layers, performing edge filtering on the interest mask layer, performing smoothing filtering on the remaining layers, and recombining the edge-filtered interest mask layer and the smoothed remaining layers.

[0005] Another depth-of-field simulation method includes: receiving multiple images; predicting an interest mask layer for each of the multiple images; performing a quantized cross-correlation operation on each of the multiple images, including determining an average brightness anchor value for the window array of the interest mask layer for each of the multiple images; setting a set of binary codes for the window array of the interest mask layer for each image based on the average brightness anchor value; determining a Hamming distance between the set of binary codes of the window array of the interest mask layer for each image; determining a cost capacity based on the Hamming distance between the set of binary codes of the window array of the interest mask layer for each image; resampling a vertical cost based on the vertical ordinal direction of the cost capacity; resampling a horizontal cost based on the horizontal ordinal direction of the cost capacity; determining a full-focus layer based on the vertical cost and the horizontal cost; determining a defocus layer based on the vertical cost and the horizontal cost; and determining the depth of the full-focus layer and the depth of the defocus layer. Attached Figure Description

[0006] In the attached diagram:

[0007] Figure 1 This is a first example system diagram based on an embodiment of the present disclosure;

[0008] Figure 2 This is a second example system diagram based on an embodiment of the present disclosure;

[0009] Figure 3 This is an example method workflow based on an embodiment of the present disclosure;

[0010] Figure 4 This is an example of a quantized cross-correlation workflow based on an embodiment of this disclosure;

[0011] Figure 5 This is another example of cost volume partitioning and adaptive aggregation based on an embodiment of the present disclosure;

[0012] Figure 6 This is a first example method based on an embodiment of the present disclosure;

[0013] Figure 7 This is a second example method based on an embodiment of the present disclosure; and

[0014] Figure 8 This is a third example method based on an embodiment of the present disclosure. Detailed Implementation

[0015] The embodiments listed below are for illustrative purposes only and are not intended to limit the scope of the apparatus and method. Equivalent modifications to the apparatus and method should fall within the scope of the claims.

[0016] Certain terms are used throughout the following description and claims to refer to specific system components. As those skilled in the art will understand, different companies may use different names to refer to components and / or methods. This document is not intended to distinguish between components and / or methods that have different names but the same function.

[0017] In the following discussion and claims, the terms “comprising” and “including” are used in an open-ended manner and can therefore be interpreted as meaning “including, but not limited to…”. Furthermore, the terms “coupled” or “coupled” are intended to mean either an indirect or direct connection. Thus, if a first device is coupled to a second device, the coupling can be a direct connection or an indirect connection via other devices and connectors.

[0018] Figure 1 An example hybrid computing system 100 is depicted, which can be used to implement and Figure 6-8 The neural network associated with the operation of one or more parts or steps of the process described herein. In this example, the processors associated with the hybrid system include a Field Programmable Gate Array (FPGA) 122, a Graphical Processor Unit (GPU) 120, and a Central Processing Unit (CPU) 118.

[0019] CPU118, GPU120, and FPGA122 all possess the capability to provide neural networks. A CPU is a general-purpose processor capable of performing many different functions; its versatility results in the ability to execute multiple different tasks. However, its processing of multiple data streams is limited, and its capabilities relative to neural networks are also limited. A GPU is a graphics processing unit with many small processing cores capable of processing parallel tasks sequentially. An FPGA is a field-programmable device that has the ability to be reconfigured and execute any function programmable into a CPU or GPU in a hard-wired circuit style. Because FPGA programming is in the form of circuits, it is many times faster than a CPU and significantly faster than a GPU.

[0020] The system can include other types of processors, such as accelerated processing units (APUs) (which include CPUs with on-chip GPU elements) and digital signal processors (DSPs) designed to perform high-speed digital data processing. Application-specific integrated circuits (ASICs) can also perform the hard-wired functions of FPGAs; however, the lead time for designing and manufacturing ASICs is approximately several quarters of a year, rather than the rapid turnaround implementation time available in programming FPGAs.

[0021] The graphics processing unit 120, central processing unit 118, and field-programmable gate array 122 are connected and linked to the memory interface and controller 112. The FPGA is connected to the memory interface via a programmable logic-to-memory interconnect 130. This additional device is utilized because the FPGA operates with a very high bandwidth and to minimize the electronics used by the FPGA for performing memory tasks. The memory interface and controller 112 are also connected to a permanent storage disk 110, system memory 114, and read-only memory (ROM) 116.

[0022] Can be used Figure 1 The system programs and trains the FPGA. The GPU can handle unstructured data well and can be used for training. Once the data has been trained, a deterministic inference model can be found, and the CPU can use the model data determined by the GPU to program the FPGA.

[0023] The memory interface and controller are connected to a central interconnect 124, which is also connected to a GPU 120, a CPU 118, and an FPGA 122. The central interconnect 124 is also connected to an input / output interface 128 and a network interface 126.

[0024] Figure 2 A second example hybrid computing system 200 is depicted, which can be used to implement neural networks associated with the operation of one or more parts or steps of process 1000. In this example, the processors associated with the hybrid system include a field-programmable gate array (FPGA) 210 and a central processing unit (CPU) 220.

[0025] The FPGA is electrically connected to FPGA controller 212, which is coupled to Direct Memory Access (DMA) 218. The DMA is connected to input buffer 214 and output buffer 216, which are coupled to the FPGA to buffer data into and out of the FPGA, respectively. DMA 218 includes two First-In-First-Out (FIFO) buffers, one for the main CPU and the other for the FPGA. The DMA allows data to be written to and read from the appropriate buffers.

[0026] On the CPU side of the DMA is a master switch 228, which transfers data and commands to and from the DMA. The DMA is also connected to an SDRAM controller 224, which allows data to be transferred from the CPU 220 to the FPGA and from the FPGA to the CPU 220. The SDRAM controller is also connected to external SDRAM 226 and the CPU 220. The master switch 228 is connected to a peripheral interface 230. A flash memory controller 222 controls persistent memory and is connected to the CPU 220.

[0027] To simulate the depth of field of a DSLR using a moving camera, two horizontally placed cameras can be used. When the shutter is pressed, two images are captured simultaneously. After the images are corrected, pixels in the left image will have corresponding counterparts in the right image along a horizontal line (i.e., the epipolar line). The displacement of corresponding pixels in the two images is called parallax, which is inversely proportional to the distance between the object and the camera. Using the pixel parallax, the depth of the captured objects in the scene can be found, which can then be used to find the point spread kernel function for each pixel in the simulated DSLR image. By applying the point spread function as a post-processing method, images with DSLR-like depth of field can be simulated with high fidelity.

[0028] Currently, disparity can be estimated by using a minimization function to compare the colors of pixels to the left and right of the epipolar line. Depth maps generated by simple color difference minimization techniques can provide images with noise and poor color characteristics.

[0029] Current cost-capacity construction methods allow color discrepancies between two cameras to negatively impact image quality by causing significant degradation in the depth map. In this case, when there is a systematic color discrepancy between the two cameras (which can occur in mobile devices), the resulting depth map will be significantly degraded.

[0030] Another negative impact of current cost-capacity construction methods may be attributed to misaligned camera sensors. Various image filtering techniques have been employed to extract features within small windows to minimize differences in illumination and color variations, such as those utilizing image gradient information for cost-capacity determination. However, filtering within small windows using image gradient information significantly slows down image processing.

[0031] Other current cost-capacity building methods aggregate volumes, assuming that the cost of pixels with the same depth on average will reduce noise and improve accuracy, but these techniques have significant drawbacks. Box filtering leads to the loss of detail and fine structure in the fully focused region. Bilateral filtering causes over-segmentation of the background region, resulting in rendering artifacts and a significant slowdown in image processing.

[0032] The proposed scheme for constructing cost capacity includes quantized cross-correlation that may be invariant to color and illumination biases, and applying different filters to different layers of cost capacity based on depth estimation of the full-focus region.

[0033] The quantized cross-correlation determines the average brightness on the small window to be used as the anchor value. The original brightness values ​​in the surrounding small windows can be compared with the anchor value. The comparison result can be encoded as a binary string, where values ​​greater than the anchor value are encoded as one (1), and values ​​less than the anchor value are encoded as zero (0). In this way, the surrounding pixels in the stereo pair can be described as eight (8) bits of binary.

[0034] By utilizing quantized cross-correlation, cost capacity can be determined by finding the Hamming distance between the binary codes of the left and right image pixels along the epipolar line of the layer.

[0035] Currently, cost capacity complexity is driven by the size of the input image, and resampling is rarely performed once found to avoid invalidating previously determined disparity values. In the proposed scheme, only the ordinal information of the cost capacity can be resampled, since the absolute physical depth may have a limited impact on the final result. Therefore, in the proposed scheme, cost capacity can be resampled along the horizontal and vertical dimensions while preserving the depth dimension. This cost capacity resampling maintains depth resolution.

[0036] The depth of the fully focused region can be found before the depth map of the entire image. For example, to refocus the optical system of a mobile device, the system can ask the user to indicate the region and / or object to be fully focused. By determining the depth of the indicated region, the layer of interest (ROI) can be determined. In another example, a deep learning segmentation model can predict likely regions or objects of interest. The cost capacity within the aggregate mask, with the disparity value having the minimum total cost, can indicate one or more ROIs. In yet another example, the user can place the object of interest within a specific depth range, thus fixing the ROI.

[0037] Rendering depth of field in the fully focused area allows spatial resolution to preserve detail and fine structure, and the rendered depth map of the out-of-focus area should be smooth in the spatial domain, giving the resulting image a natural look.

[0038] Figure 3 An example overall workflow 300 is depicted. A first image 310 on the left and a second image 312 on the right are received, respectively. The interest mask layers for the first image 310 and the second image 312 are predicted, respectively. The first image 310 and the second image 312 are then subjected to quantized cross-correlation operations 314 and 316, respectively. The quantized cross-correlation operation for the first image 310 includes: determining a first average brightness anchor value for a first window array of the interest mask layer of the first image, and setting a first set of binary codes for the first window array based on the first average brightness anchor value. The quantized cross-correlation operation for the second image 312 includes: determining a second average brightness anchor value for a second window array of the interest mask layer of the second image, and setting a second set of binary codes for the second window array based on the second average brightness anchor value. A Hamming distance 318 is determined between the first set of binary codes and the second set of binary codes, cost capacity 320 is resampled, and semantic partitioning 322 is performed. Semantic partitioning allows the image to be divided into a full-focus layer 324 and a defocus layer 328, which replicates the focal length of the DSLR. An edge-preserving filter can be used to filter the full-focus layer 326 to maintain sharp edges. A smoothing filter combined with semantic segmentation can be used to filter the defocus layer 330 to allow natural out-of-focus portions of the image. The filtered full-focus layer and the filtered defocus layer can be recombine 332, and depth inference can be performed 334 before rendering the resulting image 336.

[0039] Figure 4The quantized cross-correlation 400 of the example is depicted. If array 410 is the left image window and array 412 is the right image window, the average value of the cells within arrays 410 and 412 is 46, given by 414 and 416. If the average value is located at the center of each of the arrays 418 and 420, it can be determined whether each cell in the surrounding cells is less than or greater than the average value in the center cell. If the surrounding cells have an original value greater than the average, the cell is refilled with a - (1), and if the original value is less than the average, the cell is refilled with zero (0). After removing the center of the average, the remaining eight (8) values ​​indicate the quantized cross-correlation of the window. The Hamming distance 430 can be determined from the two quantized cross-correlations 426 and 428.

[0040] Different aggregation strategies can be applied to different layers of interest. For example, edge-preserving filters such as bilateral filters, guided filters, or domain transform filters can be applied to depict fine details. In the defocus depth layer, cost capacities can be separated and spatially downsampled. Subsampling reduces subsequent post-processing because it reduces horizontal and vertical dimensions and maintains depth resolution since subsampling is not performed in the parallax dimension. Smoothing filters can be applied. For finer detail, bilateral filtering with large spatial and color sigma values ​​can be applied. Alternatively, similar operations can be applied in the parallax dimension to enhance depth resolution. Cost capacities can be recombined to form aggregated cost capacities, from which the final depth map can be determined.

[0041] One possible approach to depth-of-field simulation could include cost capacity construction, decoupling of the spatial and depth domains, and resampling of the cost capacity. The depth of the fully focused region can be estimated and used to divide the cost capacity into layers of interest and out-of-focus layers. Adaptive processing of components can be performed before recombination and determination of the final depth.

[0042] Figure 5 An example cost capacity partitioning and adaptive aggregation 500 is depicted. In this example, an input image is used, and a mask 510 is created for the region of interest. The original cost capacity 512 and the mask can be used to partition 514 of the original cost capacity based on the average cost capacity of the mask 516. This partitioning divides the cost capacity into a fully focused layer of interest 520 and a defocused remaining layer 518. An edge-preserving filter 522 is applied to the fully focused layer of interest, and a smoothing filter 524 is applied to the defocused remaining layer. The filtered layer of interest and the filtered remaining layer can be recombine to form a recombined cost capacity 526.

[0043] Figure 6An example method for depth-of-field simulation 600 is described, including receiving 610 a first image, receiving 612 a second image, predicting 614 a mask layer of interest for the first image, determining 616 a first average brightness anchor value for a first window array of the mask layer of interest, and determining 618 a second average brightness anchor value for a second window array of the mask layer of interest. The method also includes setting 620 a first set of binary codes for the first window array, setting 622 a second set of binary codes for the second window array, determining 624 a Hamming distance between the first and second sets of binary codes, and determining 626 a cost capacity based on the Hamming distance. The method further includes resampling 628 vertical cost based on the vertical ordinal direction of the cost capacity, resampling 630 horizontal cost based on the horizontal ordinal direction of the cost capacity, determining 632 a fully focused layer based on the vertical and horizontal costs, determining 634 a defocused layer based on the vertical and horizontal costs, and determining 636 the depth of the fully focused layer and the depth of the defocused layer.

[0044] This method may include semantic partitioning of cost capacity, edge filtering of the full-focus layer, and smoothing filtering of the off-focus layer. The method may also include recombining the edge-filtered full-focus layer and the smoothing-filtered off-focus layer, decoupling the spatial and depth domains of the first and second images, estimating the full-focus depth of the full-focus layer, and separating the full-focus layer and the off-focus layer.

[0045] Figure 7 Another example method for depth-of-field simulation 700 is described, including receiving an initial cost capacity of a first image 710, determining an interest mask layer of the first image 712, and determining the average cost capacity of the interest mask layer 714. The method also includes separating the interest mask layer from the remaining layers 716, performing edge filtering on the interest mask layer 718, performing smoothing filtering on the remaining layers 720, and recombining the edge-filtered interest mask layer with the smoothed remaining layers 722. The method may further include determining a final depth map based on the recombined cost capacity.

[0046] Figure 8Another example method for depth-of-field simulation 800 is described, including receiving 810 or more images, predicting the interest mask layer for each of the 812 or more images, and performing a quantized cross-correlation operation on each of the 810 or more images, including: determining an average brightness anchor value for the window array of the interest mask layer for each of the 814 or more images, and setting a set of binary codes for the window array of the interest mask layer for each of the 816 images based on the average brightness anchor value. The method includes determining a Hamming distance between the set of binary codes of the window array of the interest mask layer for each of the 818 images, determining a cost capacity based on the Hamming distance between the set of binary codes of the window array of the interest mask layer for each of the 820 images, resampling the vertical cost based on the vertical ordinal direction of the cost capacity 822, and resampling the horizontal cost based on the horizontal ordinal direction of the cost capacity 824. The method also includes determining a full-focus layer based on the vertical and horizontal costs 826, determining a defocus layer based on the vertical and horizontal costs 828, and determining the depth of the full-focus layer and the depth of the defocus layer 830.

[0047] The method may further include semantic partitioning of cost capacity, edge filtering of the full-focus layer, and smoothing filtering of the defocus layer. The method may include recombining the edge-filtered full-focus layer and the smoothing-filtered defocus layer, decoupling the spatial and depth domains of the first and second images, estimating the full-focus depth of the full-focus layer, and separating the full-focus layer and the defocus layer.

[0048] Those skilled in the art will understand that the various illustrative blocks, modules, elements, components, methods, and algorithms described herein can be implemented as electronic hardware, computer software, or a combination of both. To illustrate this interchangeability between hardware and software, the various illustrative blocks, modules, elements, components, methods, and algorithms have been broadly described above in terms of their functionality. Whether this functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the system. Those skilled in the art can implement the described functionality in different ways for each specific application. Without departing from the scope of the subject matter, the various components and blocks can be arranged differently (e.g., in different orders or divided in different ways).

[0049] It should be understood that the specific order or hierarchy of steps in the disclosed process is illustrative of the method. Based on design preferences, it is understood that the specific order or hierarchy of steps in the process may be rearranged. Some steps may be performed simultaneously. The appended method claims present the elements of various steps in a sample order and are not intended to limit one to the specific order or hierarchy presented.

[0050] The foregoing description is provided to enable any person skilled in the art to practice the various aspects described herein. The foregoing description provides various examples of the subject matter, and the subject matter is not limited to these examples. Various modifications to these aspects may be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Therefore, the claims are not intended to limit them to the aspects shown herein, but rather to conform the full scope to the language of the claims, wherein references to the singular element are not intended to mean “one and only one”, but rather “one or more”, unless specifically stated otherwise. Unless otherwise specifically stated, the term “some” means one or more. Male pronouns (e.g., his) include female and neutral pronouns (e.g., her and its), and vice versa. Titles and subtitles (if any) are for convenience only and do not limit the invention. The predicates “configured to,” “operable to,” and “programmed to” do not imply any particular tangible or intangible modification of the object, but are intended to be used interchangeably. For example, a processor configured to monitor and control an operation or component may also mean that the processor is programmed to monitor and control the operation, or that the processor is operable to monitor and control the operation. Similarly, a processor configured to execute code can be interpreted as a processor that is programmed to execute code or operable to execute code.

[0051] For example, phrases such as "aspect" do not imply that such an aspect is essential to the present subject matter, or that such an aspect applies to all configurations of the present subject matter. Disclosures relating to an aspect may apply to all configurations, or one or more configurations. An aspect may provide one or more examples. For example, phrases such as "aspect" may refer to one or more aspects, or vice versa. For example, phrases such as "embodiment" do not imply that such an embodiment is essential to the present subject matter, or that such an embodiment applies to all configurations of the present subject matter. Disclosures relating to embodiments may apply to all embodiments, or one or more embodiments. An embodiment may provide one or more examples. For example, phrases such as "embodiment" may refer to one or more embodiments, or vice versa. For example, phrases such as "configuration" do not imply that such a configuration is essential to the present subject matter, or that such a configuration applies to all configurations of the present subject matter. Disclosures relating to configurations may apply to all configurations, or one or more configurations. A configuration may provide one or more examples. For example, phrases such as "configuration" may refer to one or more configurations, or vice versa.

[0052] The word “example” as used in this document means “used as an example or illustration.” Any aspect or design described as an “example” in this document is not necessarily to be construed as being superior or more advantageous than other aspects or designs.

[0053] Equivalent structures and functions of elements throughout the various aspects described herein, as known or will be known by one of ordinary skill in the art, are expressly incorporated herein by reference and are intended to be covered by the claims. Furthermore, regardless of whether such disclosure is expressly recited in the claims, the contents disclosed herein are not intended as a donation to the public. No element of any claim shall be construed under 35 U.S.SC §112, paragraph 6, unless it is expressly stated using the phrase “means for…” or, in the case of a method claim, using the phrase “steps for…”. Furthermore, with regard to the terms “comprising,” “having,” etc., as used in the specification or claims, such terms are intended to be included in the manner of the term “comprising,” similar to the interpretation of “comprising” when it is used as a conjunction in a claim.

[0054] References to "one embodiment," "embodiment," "some embodiments," "various embodiments," etc., indicate that a particular element or feature is included in at least one embodiment of the invention. Although the phrase may appear in various places, the phrase does not necessarily refer to the same embodiment. Together with this disclosure, those skilled in the art may be able to design and combine any of the various mechanisms suitable for achieving the above-described functions.

[0055] It should be understood that this disclosure teaches only one example of an illustrative embodiment, and many variations of the invention can be readily devised by those skilled in the art upon reading this disclosure, the scope of which will be determined by the following claims.

Claims

1. A depth-of-field simulation method, comprising: Receive the first image; Receive the second image; Predict the interest mask layer of the first image; The cross-correlation operation of quantizing the first image includes: determining the first average brightness anchor value of the first window array of the interest mask layer of the first image, and setting the first set of binary codes of the first window array based on the first average brightness anchor value. Predict the interest mask layer of the second image; The cross-correlation operation of quantizing the second image includes: determining the second average brightness anchor value of the second window array of the interest mask layer of the second image, and setting the second set of binary codes of the second window array based on the second average brightness anchor value; Determine the Hamming distance between the first set of binary codes and the second set of binary codes; The cost capacity is determined by determining the Hamming distance; Based on the vertical ordinal direction of the cost capacity, resample the vertical cost; Based on the horizontal ordinal direction of the cost capacity, resample the horizontal cost; The full-focus layer is determined by performing semantic partitioning on the vertical cost and the horizontal cost; The defocus layer is determined by performing semantic partitioning on the vertical cost and the horizontal cost; The full-focus layer is filtered using an edge-preserving filter; The defocus layer is filtered using a smoothing filter; The filtered fully focused layer and the filtered out-of-focus layer are recombined and depth inference is performed; and The depths of the full-focus layer and the defocus layer are determined by rendering the depths of the full-focus layer and the defocus layer.

2. The method according to claim 1, characterized in that, It also includes decoupling the spatial and depth domains of the first and second images.

3. The method according to claim 2, characterized in that, It also includes estimating the full-focus depth of the full-focus layer.

4. The method according to claim 3, characterized in that, It also includes separating the full-focus layer and the defocus layer.

Citation Information

Patent Citations

  • Binocular deep learning method based on adaptive single-peak stereo matching cost filtering

    CN111709977A

  • Depth of field image refocusing

    WO2020187425A1