Texture enhancement and denoising method applied to depth image
By combining adaptive threshold noise detection and an improved median filter with Canny edge detection and texture enhancement operators, the problems of texture corruption and computational complexity in depth image processing are solved, achieving fast and effective noise removal and edge enhancement.
Patent Information
- Application Number
- CN202511046553.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-28
AI Technical Summary
Existing depth image processing methods struggle to effectively preserve object boundary texture information during denoising, and are computationally complex, resource-intensive, and unable to respond quickly.
We employ adaptive threshold noise detection and an improved median filter to remove noise, and combine Canny edge detection and an eight-directional texture enhancement operator to enhance edge texture and reduce computation.
It effectively preserves image edge texture details, reduces computational load, enables fast processing and response, and is suitable for more complex scenarios, avoiding the tedious operations of depth map super-resolution and image registration.
Smart Images

Figure CN121032840A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image technology applied to CMOS depth image sensors, and in particular to a method for fast edge texture enhancement and noise reduction of depth images. Background Technology
[0002] For widely used Time-of-Flight (ToF) cameras, the system emits infrared light of a specific frequency. This light reaches the object and returns to the receiver. The distance to the object is determined by measuring the time or phase relationship between the emitted and returned light. Depth measurement systems have significant advantages in a wide range of applications. However, when objects in the scene obstruct the view, or when the surface of a projected object has excessively high or no reflectivity, the depth sensor may fail to correctly receive the returned signal, resulting in abnormal impulse noise in the depth image. Additionally, at the boundaries of object edges, the difference in distance between the foreground and background causes unstable light reception, introducing edge noise into the depth map.
[0003] To improve the accuracy of depth camera measurements, many types of depth map denoising methods have been proposed. Among them, the more mainstream methods are: (1) single-frame or multi-frame continuous depth map denoising. (2) color image-guided depth map filtering. (3) deep learning-based depth map denoising methods, etc. For single-frame image denoising, the traditional method is bilateral filtering. This method has made great progress in depth image processing technology, but it cannot well preserve the boundary texture information of objects during the denoising process, which is a disadvantage of this type of method. Later, considering the disadvantage of single-frame image denoising, color image-guided filtering was proposed. For the classic joint bilateral filtering, this method alleviates the disadvantage of insufficient boundary texture preservation of traditional bilateral filtering and has a better smoothing effect. The important problem faced by this method is that the result will produce texture replication and gradient inversion problems, and not all depths can generate color information. Moreover, considering that the depth map has a lower resolution than the color map, the depth map needs to be upsampled before registering the two to keep the depth map and the color map at the same resolution. For two heterogeneous color maps and depth maps, there are transformations such as rotation, scaling and translation. Registration between images from different sources is difficult, prone to mismatches, and the registration process is time-consuming. While camera-based denoising methods like deep learning have achieved good results, their effectiveness is highly dependent on the training dataset. For scenarios with incomplete datasets, they do not perform well. Furthermore, computation consumes significant time and resources, making it difficult to achieve rapid processing and response. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings and defects of existing technologies by providing a texture enhancement and denoising method for depth images that can both enhance edge textures and simultaneously remove residual noise during processing. This invention solves the problem of traditional depth map denoising methods damaging the texture details of object boundaries, and eliminates the cumbersome and problematic operations of depth map super-resolution and image registration. Furthermore, this method does not require large datasets for time-consuming computation, saving computational resources and making it more applicable to more complex and diverse scenarios.
[0005] One object of the present invention is to provide a method for texture enhancement and denoising applied to depth images, comprising the following steps:
[0006] For the input depth map f(x,y), noise detection is performed using a preset noise detection window W to obtain the depth values w(i,j) of eight non-zero pixels in a 3×3 neighborhood next to the center point depth value w(i0,j0).
[0007] Calculate the absolute value r(i,j) of the depth difference between the non-zero pixel point w(i,j) and the center point w(i0,j0);
[0008] Sort the absolute values r(i,j) in ascending order to r i Calculate the depth inconsistency R(x,y) between the center point w(i0,j0) and the non-zero pixel point w(i,j):
[0009] Set up a 3×3 calculation window and calculate the mean Ave of R(x,y) for all points within the window. Count the number of R(x,y) values less than Ave within the calculation window, denoted as Q. Multiply Q by Ave to obtain the comparison threshold T. i ,
[0010] Within the 3×3 calculation window, R(x,y) is greater than the comparison threshold T. i The point is marked as 1 and is used as a noise point;
[0011] In the 3×3 neighborhood Ω i,j Remove the marked noise points; sort the remaining unmarked non-zero pixels, assign the median value of the non-zero pixels to the center pixel depth value w(i0,j0) for denoising, and output the image point g(x,y);
[0012] The Canny edge detection algorithm is used to detect edge texture in the denoised image, and the detected edges are marked.
[0013] For the point I(x,y) marked with an edge texture, set a 5×5 neighborhood window Ω x,y It employs an eight-directional 5×5 texture enhancement operator O v As a mask, for the neighboring window Ω x,yPerform convolution to obtain the convolution output I(x,y). θ ;
[0014] The convolution output is I(x,y). θ The output F(x,y) is obtained by linear weighting and is used as the final output image. Preferably, the size of the noise detection window W is (2n+1)×(2n+1), with n initially set to 1.
[0015] When using a preset noise detection window W for noise detection, it checks whether the depth values w(i,j) of the eight pixels surrounding the center point depth value w(i0,j0) under a preset n value are zero. If so, it takes n = n + 1 (n < 7) and expands the detection window to 5×5, and so on, according to (i0,j0) → (i x ,j y The direction of the search is to find a non-zero value and replace the depth value of the zero point with the non-zero value.
[0016] Preferably, if no alternative non-zero value is found after the value of n reaches the upper limit of the preset range, then w(i) is selected. x ,j y The depth value of the zero point is replaced by the mean of the two nearest non-zero values in the outer neighborhood (2n+1)×(2n+1).
[0017] Preferably, if the two nearest non-zero values of the zero point cannot be found, it indicates that the center point depth value w(i0,j0) is within (i x ,j y The region in the direction of ) is a hollow area, so the depth value w(i) x ,j y The value is 0.
[0018] Preferably, the absolute value r(i,j) of the depth difference between a non-zero pixel w(i,j) and the center point w(i0,j0) is calculated as follows:
[0019] r(i,j)=|w(i,j)-w(i0,j0)|.
[0020] Preferably, the inconsistency R(x,y) between the depth values of the center point w(i0,j0) and the non-zero pixel points w(i,j) is calculated as follows:
[0021] m = (2n + 1) 2 -1.
[0022] Preferably, the comparison threshold T i The calculation process is as follows:
[0023]
[0024] Q = ∑logic(R(x,y) <Ave);
[0025] T i =Ave × Q.
[0026] Preferably, the output image point g(x,y) is represented as follows:
[0027] g(x,y)=med (i,j)∈Ω {f(i,j)},f ij ≠0, f(i,j) is the non-zero depth value in the 3x3 neighborhood, and i,j is the image index where the non-zero value is located;
[0028] Preferably, the 5×5 texture enhancement operator O in eight directions is used. v As a mask, for the neighboring window Ω x,y Perform convolution to obtain the convolution output I(x,y). θ The process is as follows:
[0029] I(x,y) θ =I(x,y)*O v ,x,y∈Ω x,y .
[0030] Preferably, the output F(x,y) is obtained by linearly weighting the results of the convolution, and the expression is as follows: θ = 0°, 45°…315°.
[0031] The method of this invention performs adaptive threshold noise detection and localization on the noise depth map, and then uses an improved median filter for noise removal, which alleviates the damage to image details and texture caused by uniformly using denoising methods; the improved filter can reduce the influence of hole pixels, and has a good denoising effect on impulse noise introduced under natural conditions.
[0032] The method of this invention uses the Canny operator to locate edge detail textures and a texture enhancement operator to enhance texture details. The method is relatively simple and can significantly reduce the amount of computation while effectively enhancing the texture details and denoising the depth image, thus achieving the goal of fast processing and fast response.
[0033] The method of this invention effectively improves the damage to image texture details caused by traditional algorithms, achieving the goal of simultaneous denoising and edge texture enhancement. It does not require color image guidance or training with a large dataset, thus increasing the application scenarios and scope of the method. At the same time, it reduces the complexity and computational load of the method, achieving the goal of fast processing and fast response. Attached Figure Description
[0034] Figure 1This is a flowchart illustrating a texture enhancement and denoising method for depth images according to an embodiment of the present invention.
[0035] Figures 2A-2B These are schematic diagrams of the eight-directional 5×5 texture enhancement operator and its different directions used in the edge texture enhancement technology of this invention.
[0036] Figure 3 This is the raw depth map obtained from the image to be processed.
[0037] Figure 4 This invention utilizes the texture enhancement and denoising method for depth images proposed in this embodiment. Figure 3 The depth map shown is the result of processing the original depth map. Detailed Implementation
[0038] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0039] The input raw depth map contains significant noise and black holes due to factors such as object surface material and reflectivity. To effectively remove this noise while preserving image edge texture to the maximum extent, an exemplary embodiment of this application proposes a fast edge texture enhancement and denoising method for depth images.
[0040] In an exemplary embodiment of this application, the texture enhancement and denoising method applied to a depth image first performs noise detection and localization on the input depth map using a preset noise detection window, removes the detected and localized noise points, and outputs image points without noise points; then, the Canny edge detection algorithm is used to perform edge texture detection on the denoised image, and the detected edges are marked; for the edge texture marked points, a neighborhood window is set, an eight-directional texture enhancement operator is used as a mask, and the neighborhood window is convolved to obtain the convolution output result; the convolution output result is linearly weighted to obtain the output, which is the final output image.
[0041] This invention overcomes the drawback of traditional depth map denoising methods that damage the texture details of object boundaries, and eliminates the cumbersome and error-prone operations of depth map super-resolution and image registration. Furthermore, this method does not require large datasets for time-consuming computation, saving computational resources and making it more applicable to a wider range of complex scenarios.
[0042] See Figure 1 As shown in the exemplary embodiment of this application, the texture enhancement and denoising method applied to a depth image includes the following steps:
[0043] For the input depth map f(x,y), noise detection is performed using a preset noise detection window W to obtain the depth values w(i,j) of eight non-zero pixels in a 3×3 neighborhood next to the center point depth value w(i0,j0).
[0044] Calculate the absolute value r(i,j) of the depth difference between the non-zero pixel w(i,j) and the center point w(i0,j0);
[0045] Sort the absolute values r(i,j) in ascending order to r i Calculate the depth inconsistency R(x,y) between the center point w(i0,j0) and the non-zero pixel point w(i,j):
[0046] Set up a 3×3 calculation window and calculate the mean Ave of R(x,y) for all points within the window. Count the number of R(x,y) values less than Ave within the calculation window, denoted as Q. Multiply Q by Ave to obtain the comparison threshold T. i ,
[0047] Within the 3×3 calculation window, R(x,y) is greater than the comparison threshold T. i The point is marked as 1 and is used as a noise point;
[0048] In the 3×3 neighborhood Ω i,j Remove the marked noise points; sort the remaining non-zero pixels, assign the median value of the non-zero pixels to the center pixel for noise reduction, and output the image point g(x,y);
[0049] The Canny edge detection algorithm is used to detect edge texture in the denoised image, and the detected edges are marked.
[0050] For the point I(x,y) marked with an edge texture, set a 5×5 neighborhood window Ω x,y It employs an eight-directional 5×5 texture enhancement operator O v As a mask, for the neighboring window Ω x,y Perform convolution to obtain the convolution output I(x,y). θ ;
[0051] The convolution output is I(x,y). θ The output F(x,y) is obtained by performing linear weighting, which is used as the final output image.
[0052] In this embodiment of the application, the size of the noise detection window W is (2n+1)×(2n+1);
[0053] When using a preset noise detection window W for noise detection, it checks whether the depth values w(i,j) of the eight pixels surrounding the center point depth value w(i0,j0) under a preset n value are zero. If so, it takes n = n + 1 (n < 7) and expands the detection window to 5×5, and so on, according to (i0,j0) → (i x ,j y The direction of the search is to find a non-zero value and replace the depth value of the zero point with the non-zero value.
[0054] In this embodiment of the application, after the preset value of n reaches the upper limit (n≧7), if no alternative non-zero value has been found, then w(i) is selected. x ,j y The depth value of the zero point is replaced by the mean of the two nearest non-zero values in the outer neighborhood (2n+1)×(2n+1).
[0055] In this embodiment of the application, if the two nearest non-zero values of the zero point cannot be found, it means that the center point depth value w(i0,j0) is within (i x ,j y The region in the direction of ) is a hollow area, so the depth value w(i) x ,j y The value is 0.
[0056] In this embodiment, the absolute value r(i,j) of the depth difference between the non-zero pixel point w(i,j) and the center point w(i0,j0) is calculated as follows:
[0057] r(i,j)=|w(i,j)-w(i0,j0)|.
[0058] In this embodiment of the application, the inconsistency R(x,y) between the depth values of the center point w(i0,j0) and the non-zero pixel point w(i,j) is calculated as follows:
[0059] m = (2n + 1) 2 -1.
[0060] In this embodiment of the application, the comparison threshold T i The calculation process is as follows:
[0061]
[0062] Q = ∑logic(R(x,y) <Ave);
[0063] T i =Ave × Q.
[0064] In this embodiment of the application, the output image point g(x,y) is represented as follows:
[0065] g(x,y)=med (i,j)∈Ω {f(i,j)},f ij ≠0, f(i,j) is the non-zero depth value in the 3x3 neighborhood, and i,j is the image index where the non-zero value is located;
[0066] In this embodiment of the application, the 5×5 texture enhancement operator O in eight directions is used. v As a mask, for the neighboring window Ω x,y Perform convolution to obtain the convolution output I(x,y). θ The process is as follows:
[0067] I(x,y) θ =I(x,y)*O v ,x,y∈Ω x,y .
[0068] In this embodiment of the application, the output F(x,y) is obtained by linearly weighting the convolution output, and the expression is as follows:
[0069] θ = 0°, 45°…315°.
[0070] In this embodiment of the application, the texture enhancement and denoising method applied to depth images includes the following overall processing procedure and steps:
[0071] Step 1: Let the input depth map be f(x,y). First, set a noise detection window W with a size of (2n+1)×(2n+1). Initially, n=1 (n<7), selecting a 3×3 window, with w(i0,j0) as the depth value of the center point. Detect the depth values w(i,j) of the eight pixels surrounding w(i0,j0). If the pixel w(i0,j0)... x ,j y The point is located inside the black hole, i.e., w(i) x ,j y ) = 0; To avoid the influence of hole values, it is necessary to find the point w(i) = 0. x ,j y () can replace depth values.
[0072] Step 2, if w(i) x ,j y If ) = 0, then n = n + 1 (n < 7), meaning the detection window expands to 5×5. This process continues, following the formula (i0, j0) → (i x ,j y Continue searching for non-zero values in the direction of ), and replace the zero point w(i) with the found value. x ,j y ).
[0073] Step 3: If no alternative non-zero value is found, repeat step 2. If no alternative non-zero value is found when n = 6, then choose w(i). x ,j y The depth value w(i) is replaced by the mean of the two nearest non-zero values in the outer neighborhood (2n+1)×(2n+1). x ,j y If w(i) still cannot be found x ,j y The two nearest non-zero values indicate that the center point depth value w(i0,j0) is within (i... x ,j y The direction is a hollow region, so the depth value w(i) is... x ,j y The value is 0.
[0074] Step 4: After obtaining the depth values w(i,j) of the 8 non-zero pixels in the 3×3 neighborhood of the center point depth value w(i0,j0), calculate the absolute value r(i,j) of the difference between the depth values of w(i,j) and the center point w(i0,j0), that is:
[0075] r(i,j)=|w(i,j)-w(i0,j0)|
[0076] Step 5: Sort r(i,j) in ascending order to get r i The degree of inconsistency between the depth of the center point w(i0,j0) and the non-zero pixel point w(i,j) is calculated and denoted as R(x,y):
[0077] m = (2n + 1) 2 -1
[0078] Step 6: Set up a 3×3 calculation window, calculate the R(x,y) values of all points within the calculation window, and denote the mean as Ave. Count the number of R(x,y) values less than Ave within the calculation window and denote it as Q. Calculate a comparison threshold T. i To determine whether the center point depth value w(i0,j0) is an abnormal noise point:
[0079]
[0080] Q = ∑logic(R(x,y) <Ave);
[0081] T i =Ave × Q.
[0082] Step 7: Determine whether R(x,y) within the 3×3 calculation window is greater than T. iIf so, mark the center point (i0,j0) as 1, that is, the center point depth value w(i0,j0) is a noise point; otherwise, it is a normal pixel point. Repeat the above steps until all noise points are marked.
[0083] Step 8: In the selected calculation window Ω i,j In a 3×3 neighborhood, marked noise points are directly removed. The remaining non-zero pixel depth values that were not marked as noise are sorted, and the median depth value of the non-zero pixels is then selected. (i,j)∈Ω The value is assigned to the center pixel w(i0,j0) for noise reduction, and the output image point is g(x,y).
[0084] g(x,y)=med (i,j)∈Ω {f(i,j)},f ij ≠0, f(i,j) is the non-zero depth value in the 3x3 neighborhood, and i,j is the image index where the non-zero value is located;
[0085] Step 9: Perform edge texture detection on the denoised image that has better preserved edge texture. Canny edge detection, which has good edge extraction performance, is used to extract edges, and the detected edges are marked.
[0086] Step 10: For the point I(x,y) marked with the edge texture, set a 5×5 neighborhood window Ω. x,y It employs an eight-directional 5×5 texture enhancement operator O v (Using counterclockwise as the direction, θ = 0°, 45°, 90°…315°, where the most appropriate values for a and b are a = 1.6 and b = 2.08) is the convolution kernel, for a 5×5 Ω... x,y Convolution is performed on regions, as follows:
[0087] I(x,y) θ =I(x,y)*O v ,x,y∈Ω x,y
[0088] Step 11: Process the convolution output I(x,y) θ The output F(x,y) is obtained by performing linear weighting, which is the final output image; the process is as follows:
[0089] θ = 0°, 45°…315°
[0090] F(x,y) is the depth image output after processing, which has edge texture enhancement and noise reduction effects.
[0091] test:
[0092] The proposed method was tested using depth maps with unclear boundary contours and high noise levels.
[0093] The experimental environment was a PC with a main frequency of 2.5GHz and 8GB of memory, and the software development tool was MATLAB 2018a.
[0094] The specific steps are as follows:
[0095] First of all Figure 3 The input depth map is subjected to noise detection and localization marking. Then, the proposed method is used to remove most of the noise. Next, the Canny method is used to extract image texture details from the noise-removed depth map. Finally, the proposed edge texture enhancement method is used to enhance edge texture details and remove a small amount of residual noise.
[0096] In the experiment, the depth map was captured using a self-made ToF camera, with a size of 1080×1920. The results are attached. Figure 3 As shown, the results obtained using the proposed method can be seen. Figure 4 Compared to the original depth map Figure 3 Most of the noise has been removed, and edge details have been enhanced, resulting in good performance. The small gray and black squares in the lower right corner are from the test pixel area, which differs from the other pixels and is normal.
[0097] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. It will be apparent to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or basic features of the present invention.
[0098] Therefore, the embodiments should be regarded as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, it is intended that all variations falling within the meaning and scope of the equivalents of the claims be included within the invention.
[0099] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A texture enhancement and denoising method applied to depth images, characterized in that, Includes the following steps: For the input depth map f(x,y), noise detection is performed using a preset noise detection window W to obtain the depth values w(i,j) of eight non-zero pixels in a 3×3 neighborhood next to the center point depth value w(i0,j0). Calculate the absolute value r(i,j) of the depth difference between the non-zero pixel w(i,j) and the center point w(i0,j0); Sort the absolute values r(i,j) in ascending order to r i Calculate the depth inconsistency R(x,y) between the center point w(i0,j0) and the non-zero pixel point w(i,j): Set up a 3×3 calculation window and calculate the mean Ave of R(x,y) for all points within the window. Count the number of R(x,y) values less than Ave within the calculation window, denoted as Q. Multiply Q by Ave to obtain the comparison threshold T. i , Within the 3×3 calculation window, R(x,y) is greater than the comparison threshold T. i The point is marked as 1 and is used as a noise point; In the 3×3 neighborhood Ω i,j Remove marked noise points; Sort the remaining unlabeled non-zero pixels, assign the median value of the non-zero pixels to the depth value w(i0,j0) of the center pixel for denoising, and output the image point g(x,y); The Canny edge detection algorithm is used to detect edge texture in the denoised image, and the detected edges are marked. For the point I(x,y) marked with an edge texture, set a 5×5 neighborhood window Ω x,y It employs an eight-directional 5×5 texture enhancement operator O v As a mask, for the neighboring window Ω x,y Perform convolution to obtain the convolution output I(x,y). θ ; The convolution output is I(x,y). θ The output F(x,y) is obtained by performing linear weighting, which is used as the final output image.
2. The texture enhancement and denoising method for depth images according to claim 1, characterized in that, The size of the noise detection window W is (2n+1)×(2n+1), with n initially set to 1. When using a preset noise detection window W for noise detection, the depth values w(i,j) of the eight pixels surrounding the center point depth value w(i0,j0) under a preset n value are checked to see if they are zero. If so, n = n + 1 (n < 7), and the detection window is expanded to 5×5. This process is repeated sequentially, following the order (i0,j0) → (i x ,j y The direction of the search is to find a non-zero value and replace the depth value of the zero point with the non-zero value.
3. The texture enhancement and denoising method for depth images according to claim 2, characterized in that, If, after the value of n reaches the upper limit of the preset range, no alternative non-zero value has been found, then w(i) is selected. x ,j y The depth value of the zero point is replaced by the mean of the two nearest non-zero values in the outer neighborhood (2n+1)×(2n+1).
4. The texture enhancement and denoising method for depth images according to claim 3, characterized in that, If the two nearest non-zero values of the zero point cannot be found, it means that the center point depth value w(i0,j0) is within (i x ,j y The region in the direction of ) is a hollow area, so the depth value w(i) x ,j y The value is 0.
5. The texture enhancement and denoising method for depth images according to claim 4, characterized in that, The absolute value r(i,j) of the depth difference between a non-zero pixel w(i,j) and the center point w(i0,j0) is calculated as follows: r(i,j)=|w(i,j)-w(i0,j0)|.
6. The texture enhancement and denoising method for depth images according to claim 5, characterized in that, The inconsistency R(x,y) between the depth values of the center point w(i0,j0) and the non-zero pixel points w(i,j) is calculated as follows:
7. The texture enhancement and denoising method for depth images according to claim 6, characterized in that, The comparison threshold T i The calculation process is as follows: Q = ∑logic(R(x,y) <Ave); T i =Ave×Q。 8. The texture enhancement and denoising method for depth images according to claim 7, characterized in that, The output image point g(x,y) is represented as follows: g(x,y)=med (i,j)∈Ω {f(i,j)},f ij ≠0, f(i,j) is the non-zero depth value in the 3x3 neighborhood, and i,j is the image index where the non-zero value is located.
9. The texture enhancement and denoising method for depth images according to claim 8, characterized in that, The 5×5 texture enhancement operator O used in eight directions v As a mask, for the neighboring window Ω x,y Perform convolution to obtain the convolution output I(x,y). θ The process is as follows: I(x,y) θ =I(x,y)*O v ,x,y∈Ω x,y 。 10. The texture enhancement and denoising method for depth images according to claim 9, characterized in that, The output F(x,y) is obtained by linearly weighting the results of the convolution, and the expression is as follows: θ = 0°, 45°…315°.