ToF depth image denoising method based on confidence perception diffusion model

The confidence-aware diffusion model for ToF depth images addresses the challenge of balancing quality and detail preservation by using confidence-guided networks and dynamic range normalization, enhancing denoising performance and robustness in complex scenes.

CN120318104APending Publication Date: 2025-07-15TONGJI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510216025.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Existing ToF depth image denoising methods are difficult to balance global generation quality and local detail fidelity while removing noise, and relying on a large amount of real data to train, resulting in high cost and poor adaptability.

Method used

The confidence-aware diffusion model is adopted, and dynamic range normalization and confidence graph generation is generated, combining the confidence-guided network and the diffusion denoising network to denoise the ToF depth image, and train it using Gaussian noise simulation data to adapt to different noise environments.

Benefits of technology

Effectively remove noise, balance global generation quality and local detail fidelity, reduce dependence on real data, and improve denoising robustness and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318104A_ABST
    Figure CN120318104A_ABST
Patent Text Reader

Abstract

The invention relates to a ToF depth image denoising method and device based on a confidence perception diffusion model, and the method employs the confidence perception diffusion model to carry out the denoising of a single-frequency ToF depth image or a multi-frequency ToF depth image, and comprises the steps: converting raw data into IQ data, carrying out the dynamic range normalization, and generating a confidence map based on a ToF depth image; and de-noising based on the IQ data after the dynamic range normalization and the confidence map to obtain a de-noised ToF depth map. Compared with the prior art, the method not only enhances the recovery capability of the abnormal high-noise region, but also can accurately retain the structural details in the high-confidence region, avoids the problems of excessive smoothness and detail loss commonly existing in the prior art, and ensures the generation of the high-quality de-noised image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image denoising, and in particular to a ToF depth image denoising method based on a confidence-aware diffusion model. Background Art

[0002] In today's fields of computer vision and artificial intelligence, Time-of-Flight (ToF) depth imaging technology has been widely used in many fields such as autonomous driving, robot navigation, augmented reality and virtual reality (AR / VR), and industrial inspection because it can quickly obtain three-dimensional information of a scene. A ToF camera calculates the distance of a target object by measuring the flight time of a light pulse from emission to reception, thereby generating a depth image. However, due to the influence of various factors such as ambient light interference, electronic noise, and the limitations of the sensor itself, ToF depth images often contain a large amount of noise, which seriously reduces the quality of the depth image and the accuracy of subsequent applications. Currently, common ToF depth image denoising methods mainly include the following categories: mean filtering, median filtering, bilateral filtering, and model-based methods. Mean filtering, median filtering, and bilateral filtering are collectively referred to as filtering methods. Although this method can effectively remove the noise in the ToF depth image, it will easily lose the structural features of the depth map, only perform well in specific scenarios, or have high computational complexity and low efficiency when removing the noise; although using a model for ToF depth image denoising can better retain the details of the image. For example, Chinese Patent Application 《CN118096565A》 discloses a depth image denoising method, including: obtaining single-frequency raw data and converting it into IQ data; denoising the noisy IQ data based on a confidence-aware graph Laplacian regularization expansion network; converting the denoised IQ data into a depth map, and at the same time converting the denoised IQ data into phase data and calculating the final depth; removing low-confidence regions and setting them to invalid values. To a certain extent, it realizes the detail fidelity of depth map denoising, but it is easy to have the problem that detail fidelity and the quality of the denoised image cannot be achieved simultaneously.

[0003] Therefore, it is a technical problem to be solved to provide a ToF depth map denoising method that can achieve high quality and detail guarantee in a balanced manner. Summary of the Invention

[0004] The purpose of the present invention is to provide a ToF depth image denoising method based on a confidence-aware diffusion model to overcome the defects of the above-mentioned existing technologies. It can be used for high-quality denoising of single-frequency ToF depth maps and multi-frequency ToF depth maps. This method can effectively remove the noise in the ToF depth image and effectively balance the contradiction between global generation quality and local detail fidelity.

[0005] The purpose of the present invention can be achieved by the following technical solutions:

[0006] The present invention provides a ToF depth image denoising method based on a confidence-aware diffusion model, which uses the confidence-aware diffusion model to denoise a single-frequency ToF depth image or a multi-frequency ToF depth image.

[0007] Among them, the method for denoising a single-frequency ToF depth image includes:

[0008] Obtain single-frequency raw data and a single-frequency ToF depth map, convert the single-frequency raw data into first IQ data and perform dynamic range normalization, and at the same time calculate the gradient magnitude of each pixel in the single-frequency ToF depth map and generate a confidence map;

[0009] Based on the first IQ data after dynamic range normalization and the confidence map, use the confidence-aware diffusion model for denoising and output second IQ data; the confidence-aware diffusion model includes a confidence guidance network and a diffusion denoising network, both the confidence guidance network and the diffusion denoising network include an encoder, an intermediate layer and a decoder, the decoder includes a plurality of decoding blocks, and the decoding blocks in the confidence guidance network correspond one-to-one to the decoding blocks in the diffusion denoising network;

[0010] Perform post-processing based on the second IQ data to obtain a denoised single-frequency ToF depth map.

[0011] The method for denoising a multi-frequency ToF depth image includes:

[0012] Obtain multi-frequency raw data and a multi-frequency ToF depth map, split the multi-frequency raw data into high-frequency raw data and low-frequency raw data, and perform the following operations on the high-frequency raw data and the low-frequency raw data respectively to obtain corresponding depth maps:

[0013] Convert the raw data into first IQ data and perform dynamic range normalization, and at the same time calculate the gradient magnitude of each pixel in the multi-frequency ToF depth map and generate a confidence map;

[0014] Based on the first IQ data after dynamic range normalization and the confidence map, use the confidence-aware diffusion model for denoising and output second IQ data; the confidence-aware diffusion model includes a confidence guidance network and a diffusion denoising network, both the confidence guidance network and the diffusion denoising network include an encoder, an intermediate layer and a decoder, the decoder includes a plurality of decoding blocks, and the decoding blocks in the confidence guidance network correspond one-to-one to the decoding blocks in the diffusion denoising network;

[0015] Fuse the second IQ data obtained based on the high-frequency raw data and the second IQ data obtained based on the low-frequency raw data, and perform post-processing to obtain a denoised multi-frequency ToF depth map.

[0016] As a preferred technical solution, the method for dynamic range normalization is as follows:

[0017]

[0018] where x i represents the in-phase component in the IQ data with high dynamic range and noise; x q represents the quadrature component of the IQ data with high dynamic range and noise; represents the in-phase component of the low dynamic range after dynamic range normalization; represents the quadrature component of the low dynamic range after dynamic range normalization; k LDR represents the normalization parameter.

[0019] As a preferred technical solution, the method for generating a confidence map is as follows:

[0020] Calculate the gradient magnitude of each pixel in the ToF depth map:

[0021]

[0022] where d mag represents the gradient magnitude, and (u, v) represents the horizontal and vertical components of the ToF depth map;

[0023] Generate a confidence map based on the gradient magnitude:

[0024]

[0025] where min(d mag ) represents the minimum value of the gradient magnitude; max(d mag ) represents the maximum value of the gradient magnitude; C represents the confidence map.

[0026] As a preferred technical solution, the method for outputting the second IQ data is as follows:

[0027] Obtain the initial noise and sample the initial noise with a standard normal distribution to generate an initial noise feature;

[0028] Perform feature encoding on the first IQ data after dynamic range normalization and the confidence map to obtain a guiding feature;

[0029] Repeat the following steps until the second IQ data is generated:

[0030] Perform zero convolution processing on the described guiding feature, and use the guiding feature after zero convolution processing and the noise feature as the input of the confidence guiding network, while using the noise feature as the input of the diffusion denoising network; the noise feature includes an initial noise feature and an intermediate noise feature, and the inputs of the confidence guiding network and the diffusion denoising network during the first diffusion are the initial noise feature, and the inputs of the confidence guiding network and the diffusion denoising network during the Nth diffusion are the intermediate noise features output by the last decoding block of the diffusion denoising network during the (N - 1)th diffusion;

[0031] In the described confidence guiding network, except that the input of the first decoding block is the fused feature of the output of the intermediate layer and the guiding feature after zero convolution processing, the input of the remaining decoding blocks is the output of the previous decoding block;

[0032] In the described diffusion denoising network, except that the input of the first decoding block is the fused feature of the output of the intermediate layer and the output of the intermediate layer of the confidence guiding network, the input of the remaining decoding blocks is the fused feature of the output of the previous decoding block and the output of the corresponding previous decoding block in the confidence guiding network, and the output of the last decoding block is the intermediate noise feature.

[0033] As a preferred technical solution, the method for obtaining the fused feature of the output of the decoding block in the diffusion denoising network and the output of the corresponding decoding block in the confidence guiding network is:

[0034] The initial noise after the t-th diffusion, and t ∈ (0, T - 1); h (g) Represents the guiding feature; Represents the operations of the diffusion denoising network including the encoder, the intermediate layer, and all decoding blocks before and including the m-th decoding block; Represents the operations of the confidence guiding network including the encoder, the intermediate layer, and all decoding blocks before and including the m-th decoding block; Represents the zero convolution operation.

[0036] As a preferred technical solution, the method for data fusion is to perform data fusion on the second IQ data obtained based on the high-frequency raw data and the second IQ data obtained based on the low-frequency raw data according to phase unwrapping.

[0037] As a preferred technical solution, the method for post-processing is:

[0038] Perform feature decoding on the data to be processed to obtain its quadrature component and in-phase component; among them, when performing single-frequency ToF depth map denoising, the data to be processed is the second IQ data; when performing multi-frequency ToF depth map denoising, the data to be processed is the result of data fusion;

[0039] Calculate the phase data based on the quadrature component and the in-phase component of the data to be processed, and calculate the depth of each data point based on the phase data;

[0040] Calculate the confidence value of each data point based on the quadrature component and the in-phase component of the data to be processed, and perform data point screening based on the confidence value to remove data points with a confidence value lower than a preset value;

[0041] Generate a denoised ToF depth map based on the depth of the remaining data points.

[0042] As a preferred technical solution, the method for calculating the depth of each data point is:

[0043] Calculate the phase data, and its expression is:

[0044]

[0045] where x q * represents the quadrature component of the second IQ data; x i * represents the in-phase component of the second IQ data; represents the phase data;

[0046] Calculate the depth based on the phase data, and the expression for calculating the depth is:

[0047]

[0048] where f is the frequency of the signal emitted by the ToF sensor, c is the speed of light, and d is the depth.

[0049] As a preferred technical solution, the method for calculating the confidence value of each data point is:

[0050] r = abs(x i * ) + abs(x q * ),

[0051] where x q * represents the quadrature component of the second IQ data; x i * represents the in-phase component of the second IQ data.

[0052] Compared with the prior art, the present invention has the following beneficial effects:

[0053] 1) When performing diffusion denoising on the ToF depth map, the present invention adds a confidence map generated based on the pixel gradient of the original ToF depth image and noise features obtained by collecting noise from the original ToF depth map to guide the confidence of the diffusion denoising process. It can achieve dynamic adjustment of the denoising process and enhance the recovery ability for abnormally high-noise regions, and can accurately retain structural details in high-confidence regions. It can effectively balance the quality of the denoised image and the fidelity of image details caused by denoising only based on the diffusion model, avoiding the problems of over-smoothing and detail loss commonly existing in the prior art.

[0054] 2) The present invention provides a dynamic range normalization method for fitting IQ data to process high-dynamic-range IQ data, which can effectively reduce the difference between the input data and the pre-training domain of the diffusion model, enabling the rich prior knowledge of the diffusion model to better adapt to the ToF depth map denoising task. It solves the problem that the existing denoising methods based on the diffusion model do not consider the dynamic range of the input data, resulting in a large domain bias, making it difficult for the model to effectively generalize, and ultimately leading to distorted denoising results and even depth offset.

[0055] 3) The present invention adopts a confidence-aware diffusion model, which does not rely on a large amount of real data for training, but can directly use synthetic data simulating Gaussian noise for training, significantly reducing the dependence on expensive data collection and annotation, and having stronger robustness; and can effectively handle complex noise distributions and different noise environments, with higher adaptability and stability. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 is the flowchart of the method of the present invention;

[0057] Figure 2 is the flowchart of single-frequency ToF depth map denoising in Embodiment 2 of the present invention;

[0058] Figure 3 is the flowchart of multi-frequency ToF depth map denoising in Embodiment 2 of the present invention;

[0059] Figure 4 is the original ToF depth map of the simulation data processing experiment in Embodiment 3 of the present invention;

[0060] Figure 5 is the un-denoised ToF depth map of the simulation data processing experiment in Embodiment 3 of the present invention;

[0061] Figure 6 is the processing result map of ToFNet in the simulation data processing experiment in Embodiment 3 of the present invention; (6a) is the effect diagram; (6b) is the error schematic diagram;

[0062] Figure 7The UDA processing result graph of the simulation data processing experiment of Embodiment 3 of the present invention; (7a) is the effect graph; (7b) is the error schematic diagram;

[0063] Figure 8 The RADU processing result graph of the simulation data processing experiment of Embodiment 3 of the present invention; (8a) is the effect graph; (8b) is the error schematic diagram;

[0064] Figure 9 The Palette processing result graph of the simulation data processing experiment of Embodiment 3 of the present invention; (9a) is the effect graph; (9b) is the error schematic diagram;

[0065] Figure 10 The processing result graph of DepthGen of the simulation data processing experiment of Embodiment 3 of the present invention; (10a) is the effect graph; (10b) is the error schematic diagram;

[0066] Figure 11 The confidence graph of the method of the simulation data processing experiment of Embodiment 3 of the present invention;

[0067] Figure 12 The noise error schematic diagram of the method of the simulation data processing experiment of Embodiment 3 of the present invention;

[0068] Figure 13 The processing result graph of the method of the simulation data processing experiment of Embodiment 3 of the present invention; (13a) is the effect graph; (13b) is the error schematic diagram;

[0069] Figure 14 The result graph of the real scene data processing comparison experiment of Embodiment 4 of the present invention; (14a) is the comparison graph of the experimental results of some methods; (14b) is the comparison graph of the experimental results of another part. Detailed implementation manners

[0070] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0071] The details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more comprehensible.

[0072] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the ordinary meanings understood by those with ordinary skills in the technical field to which this application belongs. The words such as "a", "an", "one", "the" and the like involved in this application do not indicate a limitation in quantity and may represent singular or plural. The terms "comprising", "including", "having" and any variations thereof involved in this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or modules (units) is not limited to the listed steps or units, but may further include unlisted steps or units, or may further include other steps or units inherent to these processes, methods, products or devices. The similar words such as "connected", "coupled" and "joined" involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "plurality" involved in this application means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. The terms "first", "second", "third" and the like involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects.

[0073] Embodiment 1

[0074] In order to solve the problem in the prior art that it is impossible to achieve balanced high-quality and detail-guaranteed ToF depth map denoising, this embodiment provides a ToF depth image denoising method based on a confidence-aware diffusion model, aiming to improve the denoising performance and robustness of Time-of-Flight (ToF) depth images in complex scenes, ensure that the noise in ToF depth images can be effectively removed, and effectively balance the contradiction between global generation quality and local detail fidelity. The provided method can be used for denoising single-frequency ToF depth images or multi-frequency ToF depth images, and its process is as Figure 1 shown.

[0075] Taking the denoising of single-frequency ToF depth images as an example, its detailed steps are as Figure 2 shown and include:

[0076] S1. Conversion of raw data and generation of confidence maps.

[0077] S11. Obtain single-frequency raw data from a ToF sensor and obtain a single-frequency ToF depth map.

[0078] S12. Convert the single-frequency raw data into first IQ data, and its expression is:

[0079]

[0080] Among them, c θ is the original raw data collected by the ToF sensor, corresponding to the output signal at different phase shifts θ; x' i is the in-phase component of the first IQ data; x' q is the quadrature component of the first IQ data.

[0081] At this time, the first IQ data has a high dynamic range, that is, the amplitude of some pixels may reach a value of several thousand (such as a strong reflection area), while some other pixels are very close to 0 (such as a long-distance or dull area), and the numerical span of this high dynamic range can reach 1000 times or even higher.

[0082] S13. Normalize the dynamic range of the first IQ data, compress the values in the abnormal range to a controllable range, and keep the values in the normal range basically unchanged, which can reduce the domain difference between the IQ input and the training data of the pre-trained diffusion denoising network. Its expression is:

[0083]

[0084] Among them, x i represents the in-phase component in the IQ data with a high dynamic range with noise; x q represents the quadrature component in the IQ data with a high dynamic range with noise; represents the in-phase component with a low dynamic range after dynamic range normalization; represents the quadrature component with a low dynamic range after dynamic range normalization; k LDR represents the normalization parameter.

[0085] S14. Calculate the gradient magnitude of each pixel in the single-frequency ToF depth map and generate a confidence map.

[0086] S141. Calculate the gradient magnitude of each pixel in the ToF depth map:

[0087]

[0088] Among them, d mag represents the gradient magnitude, and (u, v) represents the horizontal and vertical components of the ToF depth map;

[0089] S142. Generate a confidence map based on the gradient magnitude:

[0090]

[0091] Among them, min(d mag ) represents the minimum value of the gradient magnitude; max(d mag) represents the maximum value of the gradient magnitude; C represents the confidence map.

[0092] S2. Output the second IQ data.

[0093] This process uses a confidence-aware diffusion model to obtain the second IQ data. Specifically, the confidence-aware diffusion model includes a confidence-guided network and a diffusion denoising network. Both the confidence-guided network and the diffusion denoising network include an encoder, an intermediate layer, and a decoder. The decoder includes multiple decoding blocks, and the decoding blocks in the confidence-guided network correspond one-to-one with the decoding blocks in the diffusion denoising network. Both the confidence-guided network and the diffusion denoising network use convolutional neural networks. During the training process, the parameters of the diffusion denoising network are locked, and only the confidence-guided network is optimized to retain the generation ability and prior knowledge of the diffusion denoising network, while enabling the confidence-guided network to adapt to the specific noise distribution and structural characteristics of the IQ data. And it is trained using a simulation dataset containing Gaussian noise. The loss function is as follows:

[0094]

[0095] Among them, represents the noise-free IQ feature, represents the real noise, represents the noise predicted by the diffusion denoising network at time step t, represents the loss, h (g) represents the guiding feature.

[0096] Specifically, this process includes:

[0097] S21. Obtain the initial noise and sample the initial noise with a standard normal distribution to generate the initial noise feature.

[0098] Sample the noise of the depth image from the standard normal distribution according to the user-defined time step T (where I represents the identity matrix) to generate the initial noise feature as the initial input of the confidence-aware diffusion model to participate in the reverse diffusion process of step-by-step denoising. The specific way of the reverse diffusion process is:

[0099]

[0100] Among them, represents the intermediate noise feature of the (T - t + 1)-th reverse diffusion; α t is the noise adjustment parameter, which usually increases with the time step t; represents the cumulative noise adjustment parameter from the initial state to the current time step t; ∈ θ is the output of the diffusion denoising network; σ tis the diffusion coefficient, which determines the intensity of the random noise added to the denoising result in each time step t; z ∼ N(0, 1) is a random noise term sampled from the standard normal distribution, used to maintain the randomness of the diffusion process.

[0101] S22. Feature-encode the first IQ data and the confidence map after normalizing the dynamic range to obtain the guiding feature h (g) .

[0102] Repeat steps S23 - S25 until the second IQ data is generated:

[0103] S23. Perform zero convolution on the guiding feature, use the guiding feature after zero convolution and the noise feature as the inputs of the confidence guiding network, and at the same time use the noise feature as the input of the diffusion denoising network; the noise feature includes the initial noise feature and the intermediate noise feature, and the inputs of the confidence guiding network and the diffusion denoising network at the first diffusion are the initial noise feature, and the inputs of the confidence guiding network and the diffusion denoising network at the Nth diffusion are the intermediate noise features output by the last decoding block of the diffusion denoising network at the (N - 1)th diffusion.

[0104] S24. In the confidence guiding network, except that the input of the first decoding block is the fusion feature of the output of the intermediate layer and the guiding feature after zero convolution, the inputs of the remaining decoding blocks are the outputs of the previous decoding blocks.

[0105] Among them, the method for outputting the fusion feature is:

[0106] the initial noise after the t-th diffusion, and t ∈ (0, T - 1); h (g) represents the guiding feature; represents the operations of the diffusion denoising network including the encoder, the intermediate layer, and all decoding blocks up to and including the m-th decoding block; represents the operations of the confidence guiding network including the encoder, the intermediate layer, and all decoding blocks up to and including the m-th decoding block; represents the zero convolution operation.

[0108] S25. In the diffusion denoising network, except that the input of the first decoding block is the fusion feature of the output of the intermediate layer and the output of the intermediate layer of the confidence guiding network, the inputs of the remaining decoding blocks are the fusion features of the output of the previous decoding block and the corresponding output of the previous decoding block in the confidence guiding network, and the output of the last decoding block is the intermediate noise feature.

[0109] S3. Perform post-processing based on the second IQ data to obtain the denoised single-frequency ToF depth map.

[0110] S31. Feature-decode the second IQ data to obtain its quadrature component and in-phase component.

[0111] S32. Calculate the phase data based on the quadrature component and in-phase component of the second IQ data, and calculate the depth of each data point based on the phase data.

[0112] S321. Calculate the phase data, and its expression is:

[0113]

[0114] where x q * represents the quadrature component of the second IQ data; x i * represents the in-phase component of the second IQ data; represents the phase data;

[0115] S322. Calculate the depth based on the phase data, and the expression for calculating the depth is:

[0116]

[0117] where f is the frequency of the signal emitted by the ToF sensor, c is the speed of light, and d is the depth.

[0118] S33. Calculate the confidence value r of each data point based on the quadrature component and in-phase component of the second IQ data. The lower the r value, the lower the confidence of the point. By setting a threshold r thres of r, some points with confidence lower than this value are set to invalid values, and the method for calculating the confidence value is:

[0119] r = abs(x i * ) + abs(x q * ),

[0120] where x q * represents the quadrature component of the second IQ data; x i * represents the in-phase component of the second IQ data.

[0121] S34. Generate a denoised ToF depth map based on the depths of the remaining data points.

[0122] The above is the denoising method for single-frequency ToF depth images. This application also provides a denoising method for multi-frequency ToF depth images, and its processing process is as Figure 3As shown, it is slightly different from single-frequency ToF depth image denoising. That is, the acquired multi-frequency raw data needs to be first split into high-frequency raw data and low-frequency raw data, and the operations in steps S1 - S2 are respectively performed on the high-frequency raw data and the low-frequency raw data to obtain the second IQ data based on the high-frequency raw data and the second IQ data based on the low-frequency raw data, and the operations performed on them are as follows: The second IQ data obtained based on the high-frequency raw data and the second IQ data obtained based on the low-frequency raw data are subjected to data fusion according to phase unwrapping and then post-processed to obtain the denoised multi-frequency ToF depth map. Among them, the steps of the mentioned post-processing are as shown in S31 - S34, but the data object of the post-processing is changed to the fusion data obtained after data fusion of the second IQ data obtained based on the high-frequency raw data and the second IQ data obtained based on the low-frequency raw data.

[0123] This embodiment also provides a ToF depth image denoising device, which is used to implement all the above methods. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process described can refer to the corresponding process in the foregoing method embodiment and will not be elaborated here.

[0124] Embodiment 2

[0125] To verify whether the method provided in Embodiment 1 is feasible, this embodiment uses simulation data containing Gaussian noise for verification experiments. Among them, the original ToF depth map constructed in this embodiment is as Figure 4 shown. Figure 5 is the non-denoised image of the original ToF depth map in this embodiment. This embodiment uses a total of two major types of methods for experiments, including deep learning and diffusion models. Deep learning includes three methods, namely: ToFNet, UDA, and RADU, and their results are as Figure 6 , Figure 7 and Figure 8 , where Fig. (6a), Fig. (7a), and Fig. (8a) are the effect diagrams, and Fig. (6b), Fig. (7b), and Fig. (8b) are the corresponding errors. It can be observed that the methods based on deep learning still have a large amount of noise remaining during the denoising process, especially the depth recovery effect at the edges is poor; the diffusion models are Palette and DepthGen, and their results are as Figure 9 and Figure 10 , where Fig. (9a) and Fig. (10a) are the effect diagrams, and Fig. (9b) and Fig. (10b) are the error schematic diagrams. It can be seen that although the methods based on the diffusion model achieve overall smoothing, there are obvious deviations in depth estimation.

[0126] Using the method provided in Embodiment 1 for denoising, the confidence map obtained is as Figure 11 shown, and the noise error is asFigure 12 As shown, the result can be obtained as Figure 13 , where Fig. (13a) is the effect diagram and Fig. (13b) is the error schematic diagram. It can be seen that this solution can significantly remove most of the noise and shows better performance in the reconstruction of details such as edges. This advantage is attributed to the guiding ability of the confidence-aware diffusion model. Figure 11 The confidence map shown can effectively guide the diffusion model to achieve an accurate balance between global generation quality and local detail fidelity.

[0127] In summary, the method provided in Embodiment 1 is feasible and performs excellently on simulation data.

[0128] Embodiment 3

[0129] In Embodiment 2, the feasibility and excellent performance of the method provided by the present invention on simulation data are verified. To prove that the method provided by this application can be applied to real-scene data, in this embodiment, the real-scene data obtained from a Kinect v2 camera is used as experimental data for testing. Due to sensor noise and environmental light effects, the sensor data contains a large amount of noise before denoising, as shown in Figure 14 the first figure in Fig. (14a).

[0130] Similarly, to verify the superiority of the method provided by the present invention, in this embodiment, ToFNet, UDA, RADU, Palette, and DepthGen are used as comparative experiments at the same time. The results are shown in Fig. (14a) and (14b). It can be seen that although the above methods remove a certain degree of noise, the noise residue is still significant, and the effect of retaining scene details is not good. However, the method provided by the present invention shows excellent generalization performance. It can not only remove noise more thoroughly but also perform better in detail retention and depth reconstruction. This result fully demonstrates the robustness and practicality of the method provided by the present invention in complex and real scenes.

[0131] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A ToF depth image denoising method based on a confidence-aware diffusion model, characterized in that The method uses a confidence-aware diffusion model to denoise a single-frequency ToF depth image, and the method includes: Obtain single-frequency raw data and a single-frequency ToF depth map, convert the single-frequency raw data into first IQ data and perform dynamic range normalization, and at the same time calculate the gradient magnitude of each pixel in the single-frequency ToF depth map and generate a confidence map; Based on the first IQ data and the confidence map after dynamic range normalization, use the confidence-aware diffusion model to denoise and output second IQ data; the confidence-aware diffusion model includes a confidence guidance network and a diffusion denoising network, and both the confidence guidance network and the diffusion denoising network include an encoder, an intermediate layer, and a decoder. The decoder includes multiple decoding blocks, and the decoding blocks in the confidence guidance network correspond one-to-one with the decoding blocks in the diffusion denoising network; Based on the second IQ data, perform post-processing to obtain a denoised single-frequency ToF depth map.

2. A ToF depth image denoising method based on a confidence-aware diffusion model, characterized in that, The method uses a confidence-aware diffusion model to denoise a multi-frequency ToF depth image, and the method includes: Obtain multi-frequency raw data and a multi-frequency ToF depth map, split the multi-frequency raw data into high-frequency raw data and low-frequency raw data, and perform the following operations on the high-frequency raw data and the low-frequency raw data respectively to obtain corresponding depth maps: Convert the raw data into first IQ data and perform dynamic range normalization, and at the same time calculate the gradient magnitude of each pixel in the multi-frequency ToF depth map and generate a confidence map; Based on the first IQ data and the confidence map after dynamic range normalization, use the confidence-aware diffusion model to denoise and output second IQ data; the confidence-aware diffusion model includes a confidence guidance network and a diffusion denoising network, and both the confidence guidance network and the diffusion denoising network include an encoder, an intermediate layer, and a decoder. The decoder includes multiple decoding blocks, and the decoding blocks in the confidence guidance network correspond one-to-one with the decoding blocks in the diffusion denoising network; Fuse the second IQ data obtained based on the high-frequency raw data and the second IQ data obtained based on the low-frequency raw data, and perform post-processing to obtain a denoised multi-frequency ToF depth map.

3. A ToF depth image denoising method based on a confidence-aware diffusion model according to claim 1 or claim 2, characterized in that The method of dynamic range normalization is: where x i represents the in-phase component in the IQ data with high dynamic range and noise; x q represents the quadrature component of the IQ data with high dynamic range and noise; represents the in-phase component of the low dynamic range after dynamic range normalization; represents the quadrature component of the low dynamic range after dynamic range normalization; k LDR represents the normalization parameter.

4. A ToF depth image denoising method based on a confidence-aware diffusion model according to claim 1 or claim 2, characterized in that, The method of generating the confidence map is: Calculate the gradient magnitude of each pixel in the ToF depth map: where d mag represents the gradient magnitude, and (u, v) represents the horizontal and vertical components of the ToF depth map; Based on the gradient magnitude, generate a confidence map: where, min(d mag ) represents the minimum value of the gradient magnitude; max(d mag ) represents the maximum value of the gradient magnitude; C represents the confidence map.

5. A ToF depth image denoising method based on a confidence-aware diffusion model according to claim 1 or claim 2, characterized in that, The method of outputting the second IQ data is: Obtain initial noise, sample the initial noise with a standard normal distribution to generate initial noise features; Perform feature encoding on the first IQ data and the confidence map after dynamic range normalization to obtain guidance features; Repeat the following steps until the second IQ data is generated: Perform zero convolution processing on the aforementioned guiding features, and use the guiding features after zero convolution processing and the noise features as the inputs of the confidence guiding network. At the same time, use the noise features as the inputs of the diffusion denoising network. The noise features include initial noise features and intermediate noise features. When initially diffusing, the inputs of the confidence guiding network and the diffusion denoising network are the initial noise features. When diffusing for the Nth time, the inputs of the confidence guiding network and the diffusion denoising network are the intermediate noise features output by the last decoding block of the diffusion denoising network during the (N - 1)th diffusion. In the aforementioned confidence guiding network, except that the input of the first decoding block is the fused feature of the output of the intermediate layer and the guiding features after zero convolution processing, the inputs of the remaining decoding blocks are the outputs of the previous decoding blocks. In the aforementioned diffusion denoising network, except that the input of the first decoding block is the fused feature of the output of the intermediate layer and the output of the intermediate layer of the confidence guiding network, the inputs of the remaining decoding blocks are the fused features of the output of the previous decoding block and the output of the corresponding previous decoding block in the confidence guiding network. The output of the last decoding block is the intermediate noise features.

6. A ToF depth image denoising method based on a confidence-aware diffusion model according to claim 5, characterized in that, The method for obtaining the fused feature of the output of the decoding block in the diffusion denoising network and the output of the corresponding decoding block in the confidence guiding network is as follows: Among them, h m (o) represents the fused feature output by the m-th decoding block in the diffusion denoising network, represents the initial noise after T - t times of diffusion, and t ∈ (0, T - 1); h (g) represents the guidance feature; represents the operations of the diffusion denoising network including the encoder, the intermediate layer, and all decoding blocks before and including the m-th decoding block; represents the operations of the confidence guidance network including the encoder, the intermediate layer, and all decoding blocks before and including the m-th decoding block; represents the zero convolution operation.

7. A ToF depth image denoising method based on a confidence-aware diffusion model according to claim 2, characterized in that, The method for data fusion is to perform data fusion on the second IQ data obtained based on the aforementioned high-frequency raw data and the second IQ data obtained based on the low-frequency raw data by phase unfolding.

8. A ToF depth image denoising method based on a confidence-aware diffusion model according to claim 1 or claim 2, characterized in that, The method for post-processing is as follows: Perform feature decoding on the data to be processed to obtain its quadrature component and in-phase component. Among them, when performing single-frequency ToF depth map denoising, the data to be processed is the second IQ data. When performing multi-frequency ToF depth map denoising, the data to be processed is the result of data fusion. Calculate the phase data based on the quadrature component and in-phase component of the data to be processed, and calculate the depth of each data point based on the phase data. Calculate the confidence value of each data point based on the quadrature component and in-phase component of the data to be processed, and perform data point screening based on the confidence value to remove data points with a confidence value lower than the preset value. Generate the denoised ToF depth map based on the depth of the remaining data points.

9. A ToF depth image denoising method based on a confidence-aware diffusion model according to claim 8, characterized in that, The method for calculating the depth of each data point is as follows: Calculate the phase data, and its expression is: where x q * represents the quadrature component of the second IQ data; x i * represents the in-phase component of the second IQ data; represents the phase data; Calculate the depth based on the phase data, and the expression for calculating the depth is: Where f is the frequency of the signal emitted by the ToF sensor, c is the speed of light, and d is the depth.

10. A ToF depth image denoising method based on a confidence-aware diffusion model according to claim 8, characterized in that, The method for calculating the confidence value of each data point is as follows: r = abs(x i * ) + abs(x q * ), where x q * represents the quadrature component of the second IQ data; x i * represents the in-phase component of the second IQ data.

Citation Information

Patent Citations

  • Depth image denoising method and device

    CN118096565A