A depth-of-field rendering method and device, computer equipment and readable storage medium

By acquiring depth maps and calculating the diffusion circle matrix on embedded devices, generating and merging blur layers using convolution kernels, the problem of depth rendering latency on devices without GPUs is solved, achieving efficient and real-time depth rendering effects.

CN120747333BActive Publication Date: 2025-11-04MALANSHAN AUDIO & VIDEO LABORATORY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511247899.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-03
Publication Date
2025-11-04
Estimated Expiration
2045-09-03

AI Technical Summary

Technical Problem

In embedded device environments without GPUs, existing technologies struggle to achieve efficient depth-of-field rendering, especially when processing video streams where significant latency exists, making it difficult to meet real-time rendering requirements.

Method used

By acquiring the depth map of the original image, the blur circle matrix is ​​calculated using a preset calculation formula to determine the focal plane and the binary masks of the foreground and background. The layers are then convolved using preset convolution kernels to generate a blurred layer. Finally, the layers are blended using the binary mask extended from the focal plane to achieve depth-of-field rendering.

Benefits of technology

It achieves efficient depth-of-field rendering on embedded devices with limited computing power, enabling high-speed real-time rendering and generating bokeh effects with rich layers and realism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747333B_ABST
    Figure CN120747333B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of computer vision, and discloses a depth-of-field rendering method and device, computer equipment and a readable storage medium. The method comprises the following steps: acquiring an original image and a depth map thereof; extracting a focusing distance in the depth map, obtaining a circle of confusion matrix according to the depth map and the focusing distance, the circle of confusion matrix comprising a focal plane, a foreground and a background; determining a binary mask of the focal plane in the circle of confusion matrix, and calculating a binary mask of focal plane expansion according to the binary mask of the focal plane and the circle of confusion matrix; determining a binary mask of the foreground and a binary mask of the background in the depth map, extracting a foreground layer and a background layer in the depth map, and adopting a preset convolution kernel to convolve the foreground layer, the background layer, the binary mask of the foreground and the binary mask of the background, so as to obtain a blurring layer; and fusing the depth map and the blurring layer by using the binary mask of the focal plane expansion, so as to obtain a depth-of-field rendering result. The application solves the problem of depth-of-field rendering in environments such as GPU-embedded devices, and completes high-speed rendering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a depth rendering method, apparatus, computer device, and readable storage medium. Background Technology

[0002] In images captured by a mobile phone, each pixel has a depth value relative to the camera lens. The area the lens focuses on corresponds to a specific depth, which in turn corresponds to a focal plane. Pixels closer to the focal plane at a given depth are sharper, while those further away from the focal plane (whether near or far from the lens) will exhibit a blurring effect. This blurring effect is a physical optical phenomenon known as the diffusion effect, which presents a depth-related bokeh effect (DoF, depth of field).

[0003] In existing technologies, if you want to capture images with a bokeh effect, you need to rely on the parallel computing power of the GPU (Graphics Processing Unit). In embedded device environments without GPUs and only with NPU (Neural Network Processing Unit), it is difficult to achieve efficient inference, and there is a significant delay when processing video streams, which makes it difficult to meet the requirements of real-time rendering. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to overcome the shortcomings of the prior art and provide a depth-of-field rendering method, apparatus, computer device and readable storage medium.

[0005] This invention provides the following technical solution:

[0006] In a first aspect, this disclosure provides a depth-of-field rendering method, the method comprising:

[0007] Obtain the original image and obtain the depth map of the original image;

[0008] The depth at preset coordinates in the depth map is extracted to obtain the focus distance. Then, based on the depth map and the focus distance, each pixel in the depth map is interpolated and truncated using a preset calculation formula to obtain a circle of confusion matrix. The circle of confusion matrix includes the focal plane, foreground, and background.

[0009] Determine the binary mask of the focal plane in the circle of confusion matrix, and calculate the binary mask of the focal plane extension based on the binary mask of the focal plane and the circle of confusion matrix;

[0010] The depth map is reduced in size, and the binary mask of the foreground and the binary mask of the background in the reduced depth map are determined respectively. The foreground layer and the background layer in the reduced depth map are extracted. The foreground layer, the background layer, the binary mask of the foreground and the binary mask of the background are convolved with a preset convolution kernel respectively to obtain the convolution result. The blur layer is obtained based on the convolution result.

[0011] The depth map and the bokeh layer are fused using the binary mask extended by the focal plane to obtain the depth rendering result.

[0012] In an optional implementation, the preset calculation formula is:

[0013]

[0014] In the formula, Let be the diffusion circle matrix. Let be the depth of the i-th pixel in the depth map. The focusing distance is... The preset nearest distance threshold for the foreground. The preset maximum distance threshold for the background.

[0015] In an optional implementation, determining the binary mask of the focal plane in the circle of confusion matrix, and calculating the binary mask of the focal plane extension based on the binary mask of the focal plane and the circle of confusion matrix, includes:

[0016] The region within the preset value range of the diffusion circle matrix is ​​used as a binary mask for the focal plane;

[0017] Calculate the absolute value of the blur circle matrix, multiply the binary mask of the focal plane by the absolute value of the blur circle matrix, and obtain the binary mask of the extended focal plane.

[0018] In an optional implementation, the convolution result includes a foreground blur layer and a background blur layer. The step of convolving the foreground layer, the background layer, the binary mask of the foreground, and the binary mask of the background using a preset convolution kernel to obtain the convolution result includes:

[0019] The foreground layer and the background layer are convolved using the preset convolution kernels to obtain foreground blur and background blur results, respectively.

[0020] The binary mask of the foreground and the binary mask of the background are convolved using the preset convolution kernel to obtain the foreground mask and the background mask respectively;

[0021] Divide the foreground blur result by the foreground mask to obtain the foreground blur layer, and divide the background blur result by the background mask to obtain the background blur layer.

[0022] In an optional implementation, obtaining the blurred layer based on the convolution result includes:

[0023] Multiply the foreground blur layer with the binary mask of the foreground to obtain the filtered foreground blur layer;

[0024] Multiply the background blur layer with the binary mask of the background to obtain the filtered background blur layer;

[0025] The filtered foreground blur layer and the filtered background blur layer are superimposed to obtain the blurred layer.

[0026] In an optional implementation, the step of fusing the depth map and the bokeh layer using the binary mask extended by the focal plane to obtain a depth-of-field rendering result includes:

[0027] The size of the blurred layer is enlarged to the same size as the depth map to obtain an enlarged blurred layer;

[0028] The binary mask of the focal plane extension is nonlinearly distorted to obtain the distorted binary mask of the focal plane extension.

[0029] The depth map and the bokeh layer are weighted and fused using a binary mask that extends the distorted focal plane to obtain the depth rendering result.

[0030] In an optional implementation, obtaining the depth map of the original image includes:

[0031] The resolution of the original image is reduced to a preset resolution to obtain the reduced original image;

[0032] The scaled-down original image is input into a depth estimation network to obtain a depth map of the original image.

[0033] Secondly, this disclosure provides a depth-of-field rendering apparatus, the apparatus comprising:

[0034] An acquisition module is used to acquire the original image and acquire the depth map of the original image;

[0035] The calculation module is used to extract the depth of the preset coordinates in the depth map, obtain the focus distance, and use a preset calculation formula to interpolate and truncate each pixel in the depth map according to the depth map and the focus distance to obtain a circle of confusion matrix, wherein the circle of confusion matrix includes a focal plane, foreground and background.

[0036] The determination module is used to determine the binary mask of the focal plane in the circle of confusion matrix, and to calculate the binary mask of the focal plane extension based on the binary mask of the focal plane and the circle of confusion matrix.

[0037] The convolution module is used to reduce the depth map, determine the binary mask of the foreground and the binary mask of the background in the reduced depth map, extract the foreground layer and the background layer in the reduced depth map, and convolve the foreground layer, the background layer, the binary mask of the foreground and the binary mask of the background using a preset convolution kernel to obtain the convolution result, and obtain the blur layer based on the convolution result;

[0038] The fusion module is used to fuse the depth map and the bokeh layer using the binary mask extended by the focal plane to obtain the depth rendering result.

[0039] Thirdly, this disclosure provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the depth-of-field rendering method described in the first aspect.

[0040] Fourthly, this disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the depth-of-field rendering method described in the first aspect.

[0041] The beneficial effects of this application are:

[0042] The depth-of-field rendering method provided in this application embodiment acquires an original image and obtains a depth map of the original image; extracts the depth at preset coordinates in the depth map to obtain the focus distance, and uses a preset calculation formula to interpolate and truncate each pixel in the depth map according to the depth map and the focus distance to obtain a circle of confusion matrix, the circle of confusion matrix including a focal plane, foreground and background; determines the binary mask of the focal plane in the circle of confusion matrix, and calculates the extended binary mask of the focal plane according to the binary mask of the focal plane and the circle of confusion matrix; shrinks the depth map, determines the binary mask of the foreground and the binary mask of the background in the shrunken depth map, and extracts the foreground layer and the background layer in the shrunken depth map, and convolves the foreground layer, the background layer, the binary mask of the foreground and the binary mask of the background respectively using a preset convolution kernel to obtain the convolution result, and obtains a blur layer according to the convolution result; and merges the depth map and the blur layer using the extended binary mask of the focal plane to obtain the depth-of-field rendering result. This application effectively solves the problem of depth-of-field rendering in extreme environments with limited computing power, such as those without GPU embedded devices, and can achieve high-speed real-time rendering.

[0043] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0044] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort. In the various drawings, similar components are numbered similarly.

[0045] Figure 1 This illustration shows a schematic diagram of light being projected onto an imaging plane through a convex lens, according to an embodiment of this application.

[0046] Figure 2 A schematic diagram of an imaging model provided in an embodiment of this application is shown;

[0047] Figure 3 This illustration shows a schematic diagram illustrating the relationship between depth and focal plane diameter according to an embodiment of this application;

[0048] Figure 4 This illustration shows a schematic diagram of the effect of the superposition of surrounding blur circles on a pixel at the center point position, according to an embodiment of this application.

[0049] Figure 5A flowchart of a depth-of-field rendering method provided in an embodiment of this application is shown;

[0050] Figure 6 This illustration shows a schematic diagram of a foreground layer and a background layer provided in an embodiment of this application;

[0051] Figure 7 The illustration shows a schematic diagram of the actual process of pixel projection overlay and the convolution process of neighborhood pixel aggregation provided in the embodiments of this application;

[0052] Figure 8 This illustration shows a schematic diagram of the structure of a depth-of-field rendering device provided in an embodiment of this application;

[0053] Figure 9 A schematic diagram of the structure of a computer device provided in an embodiment of this application is shown. Detailed Implementation

[0054] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0055] It should be noted that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein in the template description is for the purpose of describing particular embodiments only and is not intended to limit the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0057] First, let's introduce the physical principles of DOF and the diffusion effect:

[0058] When a camera captures an image of a scene, the distance (depth value) between different points of various objects in the scene and the convex lens is different, and the result of the beam of light reflected from each point being converged by the convex lens is also different.

[0059] Due to the converging effect of a convex lens, only objects within a specific distance (the focal plane) of the lens can have their light beams converged into a single point to form a sharp pixel. Other points deviating from the focal plane will still be projected as circular spots of light (circle of confusion, or coc), creating a blurring effect. The further an object is from the focal plane, the blurrier the image; the degree of blur is related to the size of the coc. Figure 1 As shown, when light is projected onto the imaging plane on the right through a convex lens, it can be perfectly focused, underfocused, or overfocused, thus forming the light spots and patches in the right image.

[0060] To quantify the blurring effect, we first introduce the convex lens imaging formula. Where u represents the object distance, v represents the image distance, and F represents the focal length. For an object that is exactly on the focal plane, the image distance v = I, and the object distance u = P, which is also the distance from the lens to the focal plane. Its pixels converge at a point on the pixel plane to form a sharp image, such as... Figure 2 As shown by the dashed lines, where I is the preset reference image distance, representing the distance from the image of an object on the focal plane to the optical center of the convex lens after imaging by the lens; for objects at other distances, the object distance u=D, and their light rays converge at points before or after the pixel plane, projecting a circular spot onto the pixel plane. Figure 2 As shown by the solid line, its image distance is greater than the distance from the pixel plane to the optical center, where D is the object distance of the non-focal plane object, representing the distance from the object at other distances to the optical center of the convex lens.

[0061] According to the convex lens imaging formula, substituting into... Figure 2 In the imaging model, geometry can deduce Figure 3 The conclusion is as follows: the relationship between aperture diameter A, focal length F, depth z, focusing distance P, and coc diameter:

[0062] When depth z = P, the corresponding focal plane coc = 0, resulting in sharp pixels. The coc of foreground and background objects increases non-linearly with their distance from the focal plane, and the trends are slightly different: coc < 0 indicates foreground objects between the focal plane and the lens, which change drastically; coc > 0 indicates background objects far from the focal plane, which converge smoothly to Maxcoc. The coc diameter is determined by the aperture diameter A, focal length F, depth z, and focusing distance P, and its upper limit is determined by the aperture, focal length, and focusing distance, following a mathematical formula.

[0063] The process of calculating the depth-of-field rendering effect from the COC matrix is ​​similar to a color projection superposition process. That is, each pixel will form a point or circular area of ​​light spot (COC) on the imaging plane according to its COC diameter. The larger the area of ​​the blur circle, the larger its influence area, but the influence intensity is averaged by the area. Conversely, the smaller the circle area, the smaller the influence range, but the higher the intensity. The final imaging effect is the weighted sum of the blur circle projections of all pixels.

[0064] like Figure 4 As shown, considering only the five pixel sampling points around the central black dot, the effect of the pixel at the center point (central crosshair) on the superposition of surrounding blur circles is demonstrated. It can be seen that whether the RGB value of the current pixel is affected by surrounding pixels depends entirely on whether the blur circles of surrounding pixels encompass it in distance. This also reflects that only pixels with similar depth values ​​and close distances will influence each other. The final image is formed by superimposing the blur circles (cocs) of each ray onto the imaging plane. The accumulation of different blur circle radii creates the final effect of depth-of-field rendering.

[0065] Example 1

[0066] like Figure 5 The diagram shown is a flowchart of a depth-of-field rendering method according to an embodiment of this application. The depth-of-field rendering method provided in this embodiment includes the following steps:

[0067] Step S110: Obtain the original image and obtain the depth map of the original image.

[0068] In this embodiment, the original image is first acquired. The original image is a vertical image with a 2K resolution (1920x1080) as an example. Of course, other resolutions are also possible. The resolution is related to the computing power of the device itself. For example, on mobile devices with low computing power, a 720p resolution can be used; on high-end devices, 4K or even higher resolutions can be used to meet the image quality requirements and performance balance in different scenarios. This application embodiment does not limit this.

[0069] Before performing depth rendering, a depth estimation network is needed to obtain the depth map of the original image. To accelerate depth estimation, the original image is usually scaled down to a preset resolution (e.g., 1k) to fit the depth estimation network. The depth estimation network is a conventional neural network configuration, such as the U-Net architecture based on convolutional neural networks (CNN), the Hourglass network, etc. This application will not elaborate on this further.

[0070] The steps described above, by acquiring the original image and its corresponding depth map, provide a fundamental and accurate data source for subsequent depth-of-field rendering operations. Appropriately scaling down the image during the depth estimation stage effectively reduces the computational load on the depth estimation network, improving the speed of depth estimation and enabling a more efficient depth-of-field rendering process. This is particularly suitable for running on devices with limited computing power, ensuring the system's efficiency and accuracy in the initial stage of acquiring image and depth information, and laying a solid foundation for subsequent depth-of-field effect generation based on depth information.

[0071] Step S120: Extract the depth at preset coordinates in the depth map to obtain the focus distance, and use a preset calculation formula to interpolate and truncate each pixel in the depth map according to the depth map and the focus distance to obtain a circle of confusion matrix, which includes the focal plane, foreground and background.

[0072] Understandably, in the actual implementation process, a virtual coordinate (x, y) needs to be preset. The selection of the preset coordinate can be determined according to the actual application scenario and shooting intention, directly from the depth map. The depth of this preset coordinate is extracted as the focus depth P. The focus depth P refers to the desired depth at which the camera should focus under the current camera parameters, where the object pixels at the focus depth P are the sharpest.

[0073] The focusing distance corresponds to a focal plane. This embodiment uses a preset calculation formula, which is a truncated linear function, to approximate the original nonlinear relationship. Using this preset formula, based on the depth map and the focusing distance, each pixel in the depth map is interpolated and truncated to obtain the offset of each pixel from the focal plane. This results in a two-dimensional circle of confusion (COC) matrix with the same size as the depth map, divided into foreground (negative values) and background (positive values). Pixels located precisely on the focal plane have a value of 0. The preset calculation formula is as follows:

[0074]

[0075] In the formula, The diffusion circle matrix, Let be the depth of the i-th pixel in the depth map. Focusing distance The preset nearest distance threshold for the foreground. The preset maximum distance threshold for the background.

[0076] From the above formula, we can see that The offset of all pixels relative to the focal plane can be obtained, with negative values ​​for foreground, 0 for the focal plane, and positive values ​​for the background. A simplified linear formula for calculating the circle of confusion matrix based on the offsets is used to simplify and accelerate the calculation.

[0077] The final result is a blur circle matrix normalized to the range [-1, 1], which can represent the blur level of each pixel. The calculation process of the blur circle matrix... The preset maximum distance threshold represents the background. Pixels whose distance from the focal plane falls within this threshold have a coc value between 0 and 1. However, pixels whose distance from the focal plane exceeds this threshold have a coc value exceeding 1. In this embodiment, for calculation simplification, all coc values ​​exceeding 1 are truncated, forcing the coc value to be 1. This means the distance from the focal plane is set to 8 meters, i.e., a distance exceeding the focal plane is considered a maximum distance. pixel blur level and Treat it the same way. The foreground cutoff distance is... This means that if the focusing distance is very close, then that distance is taken as the nearest cutoff distance. Beyond the cutoff distance, the blurriness will not increase further.

[0078] It should be noted that the background cutoff distance Usually more promising A larger background cutoff distance is preferable because the background distance typically varies more during shooting, allowing for greater flexibility. A larger background cutoff distance helps maintain the depth of the blur effect. The foreground cutoff distance is the smaller of the focus distance and the preset value. This is because the area closest to the lens is always the most blurred, and this maintains this characteristic when the focus distance is relatively close.

[0079] The above steps, by extracting the depth from preset coordinates in the depth map as the focus distance, accurately determine the key areas in the image that need to be focused on. This allows subsequent blurring to be applied around these areas, highlighting the subject and enhancing the image's three-dimensionality and depth. Using a preset calculation formula to interpolate and truncate the pixels in the depth map to obtain the circle of confusion matrix, effectively represents the offset relationship between each pixel and the focal plane, providing a quantitative basis for determining the degree of blur based on different offsets. By appropriately setting the cutoff distances for the foreground and background, the blur effect becomes more natural and consistent with actual visual experience. This ensures that foreground objects close to the lens have sufficient blur while also giving the background blur variations a sense of depth, avoiding over-blurring or under-blurring and improving the overall quality of depth-of-field rendering.

[0080] Step S130: Determine the binary mask of the focal plane in the circle of confusion matrix, and calculate the binary mask of the focal plane extension based on the binary mask of the focal plane and the circle of confusion matrix.

[0081] This step begins by inputting the circle of confusion matrix into the depth-of-field rendering network for subsequent operations until the final depth-of-field rendering result is obtained.

[0082] Understandably, within the preset value range of the circle of confusion matrix, the value of the binary mask of the focal plane is 0. Therefore, this region is used as the binary mask of the focal plane. The preset value range can be -0.1≤coc≤0.1, leaving ±0.1 (which can also be adjusted as needed) as a distance interval to protect the sharp area, in order to increase the sharp pixel area in focus.

[0083] Next, the absolute value of the blur circle matrix is ​​calculated. Multiplying the binary mask of the focal plane by the absolute value of the blur circle matrix yields the binary mask of the extended focal plane, which serves as the alpha channel for subsequent fusion of the sharp and blurred layers.

[0084] The above steps, by determining the binary mask of the focal plane, clearly delineate the range of sharp areas in the image, providing a clear boundary between sharp and blurred regions for subsequent bokeh processing. By generating a binary mask that extends the focal plane and using it as the alpha channel, a smooth transition and fusion of the sharp and blurred layers is achieved, avoiding harsh boundary segmentation. This results in a more natural and realistic depth-of-field effect, enhancing the visual depth and three-dimensionality of the image, and allowing the viewer's attention to focus more on the subject corresponding to the focal plane. Furthermore, by reasonably setting the preset value range, the fusion quality is further improved, ensuring the aesthetics and professionalism of the final depth-of-field rendered image.

[0085] Step S140: Reduce the depth map, determine the binary mask of the foreground and the binary mask of the background in the reduced depth map, extract the foreground layer and the background layer in the reduced depth map, and use a preset convolution kernel to convolve the foreground layer, the background layer, the binary mask of the foreground and the binary mask of the background to obtain the convolution result, and obtain the blur layer based on the convolution result.

[0086] Furthermore, the depth map and blur circle matrix are reduced to a smaller size to accelerate the calculation and simulation of different degrees of blur. Binary masking matrices for the foreground (coc<0) and background (coc>0) are generated in the reduced depth map based on whether coc is greater than 0 (the background can also be further divided into two layers with coc=0.5 as the boundary, where coc>0.5 represents the background region, which has a higher degree of blur and therefore needs to be reduced further). The foreground and background layers are then extracted from the reduced depth map, such as... Figure 6 As shown.

[0087] Current technologies directly implement physical color projection overlay, which involves too much computation and cannot be done using image convolution operations (convolution operations are usually weighted summations). Therefore, this embodiment reverses the above process, transforming the process of projecting the center RGB value onto the plane into a process of weighted summation of the neighboring RGB values ​​toward the center. Figure 7The true process of pixel projection superposition is described in the left figure, Figure 7 and the convolution process of neighborhood pixel aggregation is described in the right figure. It can be seen that the two are equivalent. The RGB value of each pixel is obtained by weighted summation of the neighborhood pixel points (including itself) within the sampling circle radiating from the center of this pixel. The size of this sampling circle is represented by a constant we set to represent the maximum sampling radius R, restricting that all diffusion circle radii r in the coc matrix satisfy: r < R. Next, all pixel points in the image are traversed, and by calculating the straight-line distance of the pixel point to the center point in the pixel coordinate system, it is determined which pixel points within the sampling disk centered on them have their coc covering the center point.

[0088] For the disk convolution process described above, it is necessary to dynamically calculate the distance of the sampling points from the current center, while the conventional convolutional neural network only supports the static convolution process with determined parameters. Therefore, in this embodiment, it is approximately considered that all sampling pixel points in the disk convolution should directly participate in the convolution calculation. This will lead to some errors. However, if the layer is divided into several layers according to the coc values, and each layer is convolved with disk convolution kernels of different sizes, although the errors will not be completely eliminated, they will be reduced. Just like using a smaller disk convolution in the area close to the focal plane and a larger disk convolution in the area far from the focal plane, which intuitively also conforms to the hierarchical change of the blur effect in the spatial distance.

[0089] Therefore, in this embodiment, the foreground layer and the background layer are respectively convolved by using a preset convolution kernel in the depthwise (depthwise separable convolution) manner to adjust the size of the layer, which is equivalent to adjusting the size of the convolution kernel in disguise, thereby affecting the blur degree, and respectively obtaining the foreground blur result and the background blur result .

[0090] It should be noted that when using a convolutional neural network in this embodiment, in order to simulate disk convolution, it is found that a 7x7 disk convolution kernel is relatively close to a circle by setting the positions to be retained as 1 and the unnecessary ones as 0, and if the convolution kernel is larger, it will seriously affect the calculation speed. Therefore, this embodiment uses a 7x7 disk convolution kernel for convolution, and the specific size can be determined according to the actual situation, and this application embodiment does not limit this.

[0091] Then, the foreground binary mask and the background binary mask are respectively convolved with the same preset convolution kernel to obtain the foreground mask and the background mask . Then, the foreground blur result and the foreground mask are divided, that is , to obtain the foreground blur layer, and the background blur result and the background mask Divide, that is This results in a blurred background layer.

[0092] It should be noted that if a background layer with a coc value greater than 0.5 is separated... Then it and its background mask need to be further reduced in size, and the reduced background mask Then convolve the layer using the preset convolution kernel. have to Finally, calculate This results in a blurred background layer. The blurriness will be stronger, creating a more blurred background effect.

[0093] The foreground blur layer is further multiplied by the foreground binary mask to filter out pixels outside the foreground binary mask, resulting in a filtered foreground blur layer. Similarly, the background blur layer is multiplied by the background binary mask to filter out pixels outside the background binary mask, resulting in a filtered background blur layer. This ensures that pixels in each layer do not overlap. Finally, the filtered foreground blur layer and the filtered background blur layer are overlaid to create a blurred layer using direct addition.

[0094] The above method uses static disk convolution combined with multiple layers to simulate an approximate blur effect. Different layers use convolution kernels of different sizes to simulate different degrees of blur, thus simulating blur effects at different depths of field, and can be effectively run on the NPU. When creating blur effects in convolutional layers, the process of convolving different convolution kernels into the same layer is changed to using a fixed preset (7×7) disk convolution kernel to convolve and scale down to different layers to simulate blur effects at different depths of field. This can further reduce the amount of computation, because large kernel convolution is very computationally intensive. In comparison, scaling operations consume much less computation, allowing the entire depth-of-field rendering process to generate a blur effect with rich layers and realism while maintaining high efficiency.

[0095] Step S150: The depth map and the bokeh layer are fused using the binary mask of the extended focal plane to obtain the depth rendering result.

[0096] Understandably, the size of the blurred layer is enlarged to the same size as the depth map to obtain an enlarged blurred layer. At the same time, in order to expand the range of sharp pixels on the focal plane, the binary mask of the expanded focal plane is non-linearly distorted. For example, the regions in the absolute value matrix of coc that are less than δ are forcibly set to 0, where δ is an adjustable range, such as 0.1, or the values ​​of coc above 0.5 are uniformly increased, but not exceeding 1. The purpose is to increase the required degree of background, and then the distorted binary mask of the expanded focal plane is obtained.

[0097] Finally, the sharp depth map and the blurred layer are weighted and blended using a binary mask extended by the distorted focal plane to obtain the final depth rendering result. At this point, the core depth rendering network with approximate diffusion circle is completed, and the output effect has built-in depth rendering.

[0098] The depth rendering network, through trade-offs and approximations, can directly model the mathematical formulas of matrix operations in the entire process using regular PyTorch operators. It does not rely on rendering engines such as OpenGL and can run directly on the NPU unit, solving the problem of depth rendering in extreme environments without GPU embedded devices.

[0099] The above steps, by enlarging the size of the blurred layer and applying non-linear distortion to the mask, allow the mask to better adapt to the actual depth-of-field distribution of the image, enhancing the naturalness of the transition between blurred and sharp areas. Using the distorted focal plane to expand the binary mask for weighted fusion, the proportion of sharp and blurred components can be rationally allocated according to the depth offset of each pixel, thus generating an image with a natural and delicate depth-of-field rendering effect. This mask-based fusion method fully utilizes the depth and blur information generated in the preceding steps, effectively integrating the sharp and blurred parts to achieve high-quality depth-of-field rendering. It makes the image highlight the subject while creating a more realistic and natural background blur effect, greatly enhancing the artistic expression and visual appeal of the image, and meeting the need for efficient operation and high-quality depth-of-field rendering results on different hardware platforms such as embedded devices.

[0100] The depth-of-field rendering method provided in this application embodiment acquires an original image and obtains a depth map of the original image; extracts the depth at preset coordinates in the depth map to obtain the focus distance, and uses a preset calculation formula to interpolate and truncate each pixel in the depth map according to the depth map and the focus distance to obtain a circle of confusion matrix, the circle of confusion matrix including a focal plane, foreground and background; determines the binary mask of the focal plane in the circle of confusion matrix, and calculates the extended binary mask of the focal plane according to the binary mask of the focal plane and the circle of confusion matrix; shrinks the depth map, determines the binary mask of the foreground and the binary mask of the background in the shrunken depth map, and extracts the foreground layer and the background layer in the shrunken depth map, and convolves the foreground layer, the background layer, the binary mask of the foreground and the binary mask of the background respectively using a preset convolution kernel to obtain the convolution result, and obtains a blur layer according to the convolution result; and merges the depth map and the blur layer using the extended binary mask of the focal plane to obtain the depth-of-field rendering result. This application effectively solves the problem of depth-of-field rendering in extreme environments with limited computing power, such as those without GPU embedded devices, and can achieve high-speed real-time rendering.

[0101] Example 2

[0102] like Figure 8 The diagram shown is a structural schematic of a depth-of-field rendering device 800 according to an embodiment of this application. The device includes:

[0103] The acquisition module 810 is used to acquire the original image and acquire the depth map of the original image;

[0104] The calculation module 820 is used to extract the depth of the preset coordinates in the depth map, obtain the focus distance, and use a preset calculation formula to interpolate and truncate each pixel in the depth map according to the depth map and the focus distance to obtain a circle of confusion matrix, wherein the circle of confusion matrix includes a focal plane, foreground and background.

[0105] The determination module 830 is used to determine the binary mask of the focal plane in the circle of confusion matrix, and to calculate the binary mask of the focal plane extension based on the binary mask of the focal plane and the circle of confusion matrix.

[0106] The convolution module 840 is used to reduce the depth map, determine the binary mask of the foreground and the binary mask of the background in the reduced depth map, extract the foreground layer and the background layer in the reduced depth map, and convolve the foreground layer, the background layer, the binary mask of the foreground and the binary mask of the background respectively using a preset convolution kernel to obtain the convolution result, and obtain the blur layer based on the convolution result;

[0107] The fusion module 850 is used to fuse the depth map and the bokeh layer using the binary mask extended by the focal plane to obtain a depth rendering result.

[0108] The depth-of-field rendering apparatus provided in this application embodiment can implement each process of the depth-of-field rendering method corresponding to Embodiment 1 and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0109] The depth rendering device provided in this application embodiment effectively solves the problem of depth rendering in extreme environments with limited computing power, such as those without GPU embedded devices, and can complete high-speed real-time rendering.

[0110] Example 3

[0111] This application also provides a computer device. Please refer to the following for details. Figure 9 , Figure 9 This is a basic structural block diagram of the computer device in this embodiment.

[0112] The computer device 9 includes a memory 91, a processor 92, and a network interface 93 that are interconnected via a system bus. It should be noted that only a computer device 9 with a memory 91, a processor 92, and a network interface 93 is shown in the figure; however, it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0113] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.

[0114] The memory 91 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or D slot compatibility test memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 91 may be an internal storage unit of the computer device 9, such as the hard disk or memory of the computer device 9. In other embodiments, the memory 91 may also be an external storage device of the computer device 9, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the computer device 9. Of course, the memory 91 may include both the internal storage unit and its external storage device of the computer device 9. In this embodiment, the memory 91 is typically used to store the operating system and various application software installed on the computer device 9, such as computer-readable instructions for slot compatibility testing methods. In addition, the memory 91 can also be used to temporarily store various types of data that have been output or will be output.

[0115] In some embodiments, the processor 92 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other depth-of-field rendering chip. The processor 92 is typically used to control the overall operation of the computer device 9. In this embodiment, the processor 92 is used to execute computer-readable instructions stored in the memory 91 or to process data, such as executing computer-readable instructions for the slot compatibility testing method.

[0116] The network interface 93 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 9 and other electronic devices.

[0117] The computer device provided in this embodiment can execute the above-described depth-of-field rendering method. The depth-of-field rendering method here can be any of the depth-of-field rendering methods described in the various embodiments above.

[0118] Example 4

[0119] This embodiment also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the depth-of-field rendering method in this embodiment.

[0120] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the computer-readable storage medium can be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device. Of course, the computer-readable storage medium can also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer-readable storage medium is typically used to store the operating system and various application software installed on the computer device. In addition, the computer-readable storage medium can also be used to temporarily store various types of data that have been output or will be output.

[0121] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative; for example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that, as an alternative implementation, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0122] In addition, the functional modules or units in the various embodiments of the present invention can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0123] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a smartphone, personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium can be a non-volatile storage medium or a volatile storage medium. For example, the storage medium can be a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or any other medium capable of storing program code.

[0124] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A depth-of-field rendering method, characterized in that, The method includes: Obtain the original image and obtain the depth map of the original image; The depth at preset coordinates in the depth map is extracted to obtain the focus distance. Then, based on the depth map and the focus distance, each pixel in the depth map is interpolated and truncated using a preset calculation formula to obtain a circle of confusion matrix. The circle of confusion matrix includes the focal plane, foreground, and background. Determine the binary mask of the focal plane in the circle of confusion matrix, and calculate the binary mask of the focal plane extension based on the binary mask of the focal plane and the circle of confusion matrix; The depth map is reduced in size, and the binary mask of the foreground and the binary mask of the background in the reduced depth map are determined respectively. The foreground layer and the background layer in the reduced depth map are extracted. The foreground layer, the background layer, the binary mask of the foreground and the binary mask of the background are convolved with a preset convolution kernel respectively to obtain the convolution result. The blur layer is obtained based on the convolution result. The depth map and the bokeh layer are fused using the binary mask extended by the focal plane to obtain the depth rendering result.

2. The depth-of-field rendering method according to claim 1, characterized in that, The preset calculation formula is: In the formula, Let be the diffusion circle matrix. Let be the depth of the i-th pixel in the depth map. The focusing distance is... The preset nearest distance threshold for the foreground. The preset maximum distance threshold for the background.

3. The depth-of-field rendering method according to claim 1, characterized in that, The step of determining the binary mask of the focal plane in the circle of confusion matrix, and calculating the binary mask of the focal plane extension based on the binary mask of the focal plane and the circle of confusion matrix, includes: The region within the preset value range of the diffusion circle matrix is ​​used as a binary mask for the focal plane; Calculate the absolute value of the blur circle matrix, multiply the binary mask of the focal plane by the absolute value of the blur circle matrix, and obtain the binary mask of the extended focal plane.

4. The depth-of-field rendering method according to claim 1, characterized in that, The convolution result includes a foreground blur layer and a background blur layer. The foreground layer, the background layer, the binary mask of the foreground, and the binary mask of the background are convolved using a preset convolution kernel to obtain the convolution result, including: The foreground layer and the background layer are convolved using the preset convolution kernels to obtain foreground blur and background blur results, respectively. The binary mask of the foreground and the binary mask of the background are convolved using the preset convolution kernel to obtain the foreground mask and the background mask respectively; Divide the foreground blur result by the foreground mask to obtain the foreground blur layer, and divide the background blur result by the background mask to obtain the background blur layer.

5. The depth-of-field rendering method according to claim 4, characterized in that, The step of obtaining the blurred layer based on the convolution result includes: Multiply the foreground blur layer with the binary mask of the foreground to obtain the filtered foreground blur layer; Multiply the background blur layer with the binary mask of the background to obtain the filtered background blur layer; The filtered foreground blur layer and the filtered background blur layer are superimposed to obtain the blurred layer.

6. The depth-of-field rendering method according to claim 1, characterized in that, The process of fusing the depth map and the bokeh layer using a binary mask extended by the focal plane to obtain a depth-of-field rendering result includes: The size of the blurred layer is enlarged to the same size as the depth map to obtain an enlarged blurred layer; The binary mask of the focal plane extension is nonlinearly distorted to obtain the distorted binary mask of the focal plane extension. The depth map and the bokeh layer are weighted and fused using a binary mask that extends the distorted focal plane to obtain the depth rendering result.

7. The depth-of-field rendering method according to claim 1, characterized in that, The process of obtaining the depth map of the original image includes: The resolution of the original image is reduced to a preset resolution to obtain the reduced original image; The scaled-down original image is input into a depth estimation network to obtain a depth map of the original image.

8. A depth-of-field rendering device, characterized in that, The device includes: An acquisition module is used to acquire the original image and acquire the depth map of the original image; The calculation module is used to extract the depth of the preset coordinates in the depth map, obtain the focus distance, and use a preset calculation formula to interpolate and truncate each pixel in the depth map according to the depth map and the focus distance to obtain a circle of confusion matrix, wherein the circle of confusion matrix includes a focal plane, foreground and background. The determination module is used to determine the binary mask of the focal plane in the circle of confusion matrix, and to calculate the binary mask of the focal plane extension based on the binary mask of the focal plane and the circle of confusion matrix. The convolution module is used to reduce the depth map, determine the binary mask of the foreground and the binary mask of the background in the reduced depth map, extract the foreground layer and the background layer in the reduced depth map, and convolve the foreground layer, the background layer, the binary mask of the foreground and the binary mask of the background using a preset convolution kernel to obtain the convolution result, and obtain the blur layer based on the convolution result; The fusion module is used to fuse the depth map and the bokeh layer using the binary mask extended by the focal plane to obtain the depth rendering result.

9. A computer device, characterized in that, It includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the depth-of-field rendering method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the depth-of-field rendering method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Depth of field real-time rendering method based on fast guided filtering

    CN108665494A

  • Generating non-destructive synthetic lens blur with in-focus edge rendering

    US20250104196A1