An image optimization method, device and medium

CN120852172BActive Publication Date: 2026-09-29QINGDAO UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510968003.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2026-09-29
Estimated Expiration
2045-07-14

AI Technical Summary

Technical Problem

然而,该类方法大多采用固定参数或静态核尺度,缺乏对局部深度密度、边缘结构的动态适应能力,导致补全结果往往出现边缘模糊、结构扭曲或误补现象,尤其在处理复杂结构或高梯度变化区域时表现不佳

Benefits of technology

[0093]本发明通过构建多级自适应深度补全机制,在无需依赖外部RGB引导或深度学习模型训练的前提下,实现了对深度图空洞的高质量修复。具体而言,该方法首先通过深度反转与局部区域占比分析建立深度分布感知能力,使膨胀核尺寸能动态适应不同区域的稀疏特性;随后通过闭合操作与选择性膨胀策略的协同作用,在保留原始几何结构的同时分阶段填补不同尺度的空洞区域;顶部延展策略专门针对场景上方的系统性缺失进行物理约束补偿,而联合滤波最终消除补全过程中的局部不一致性。这种基于深度图自身特征逐级优化的技术路径,既克服了传统方法固定参数导致的边缘模糊问题,又规避了深度学习方案对数据与算力的强依赖性,最终输出具有结构连贯性、边缘清晰度与场景合理性的稠密深度图,为各类实时性要求高或资源受限的智能感知系统提供了轻量化且可靠的深度修复解决方案。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852172B_ABST
    Figure CN120852172B_ABST
Patent Text Reader

Abstract

The application discloses an image optimization method and device and a medium, comprising: performing inversion processing on pixels of a depth map to obtain an inverted depth map; adjusting the inflation core size of the inverted depth map, performing inflation operation on each preset local area based on the inflation core size to obtain an inflation depth map; performing closed operation on the inflation depth map to obtain a closed depth map; performing first preset size filling on pixels in a first hollow area based on a pre-set selective inflation strategy to obtain a first filled depth map; performing upward extrapolation of depth values on the first filled depth map by using a top extension strategy to obtain a top extension depth map; performing second preset size filling on pixels in a second hollow area to obtain a second filled depth map, the second preset size being greater than the first preset size; performing processing on the second filled depth map by using a pre-set joint filtering strategy to obtain a filtered depth map; and restoring the filtered depth map to an original depth map conforming to a real scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to an image optimization method, device, and medium. Background Technology

[0002] A depth map is a method of encoding the distance information between an object in three-dimensional space and a viewpoint into a two-dimensional image. Its core feature is that each pixel value corresponds to the depth information of that point in space, rather than the brightness or color information found in traditional RGB images. Depth maps are widely used in various scenarios requiring spatial structure perception, including but not limited to autonomous driving environmental perception, 3D reconstruction modeling, augmented reality system localization and fusion, industrial robot path planning, drone obstacle avoidance and navigation, and virtual reality interaction. They are a key data carrier supporting intelligent systems in achieving high-precision spatial cognition and operation.

[0003] Typical depth map acquisition methods rely on various active or passive sensing technologies, such as LiDAR, structured light, time-of-flight (TOF), stereo vision, and phase-shift ranging. LiDAR, with its long-range measurement and high precision, is widely used in autonomous driving and high-precision map building. Structured light and TOF sensors, due to their high integration and real-time performance, are frequently used in consumer AR / VR devices and robot vision. However, these devices still face several limitations in practical applications. Factors such as viewpoint occlusion, ambient light interference, differences in object material reflectivity, sampling sparsity, edge blurring, and sensor accuracy limitations can all cause a large number of invalid pixels (i.e., holes) in the depth map, especially in transparent objects, distant targets, or areas with drastic lighting changes. The presence of these invalid areas severely weakens the usability and stability of the depth map in subsequent tasks, becoming a significant bottleneck in the current visual perception chain.

[0004] To address these issues, depth map completion technology has emerged. Its goal is to infer and complete missing regions using only partially observed sparse depth values, resulting in a dense depth map with coherent structure, clear edges, geometric consistency, and good noise suppression. This task is essentially a sparse data completion and interpolation problem. However, due to the often non-uniform distribution of holes in depth maps, complex local edge structures, and weak global geometric information, the completion process is far more complex than simple interpolation, placing higher demands on structural reconstruction capabilities, edge detail preservation, local coherence, and global consistency.

[0005] Currently, depth map completion techniques can be mainly divided into two categories: guided depth completion methods and unguided depth completion methods.

[0006] (1) Guided methods use RGB images acquired from the same viewpoint as the depth map as structural guidance information, and assist in depth completion through constraints such as image texture features and edge distribution. These methods, such as GuideNet, FusionNet, and MSG-CHN, demonstrate strong performance in edge restoration and geometric consistency by fusing the complementary characteristics of images and depth. However, their performance is highly dependent on the quality and accurate registration of the RGB images, and can introduce additional errors when there are drastic changes in lighting or mismatches between the image and depth.

[0007] (2) Unguided methods rely solely on the original sparse depth map for completion, eliminating the dependence on image guidance and emphasizing geometric reasoning capabilities under single-modal information. These methods are more adaptable to complex lighting conditions or scenarios where RGB images are unavailable. Representative methods include traditional interpolation, propagation, filtering algorithms (such as Joint Bilateral Filter, Total Variation Inpainting, etc.) and unsupervised learning methods.

[0008] From a technical perspective, current depth map completion methods can be further divided into two major schools of thought: deep learning-based methods and traditional image processing-based methods.

[0009] (1) Deep Learning-Based Methods: In recent years, with the maturation of deep structures such as Convolutional Neural Networks (CNN), Graph Neural Networks (GNN), and Transformers, many researchers have attempted to treat the deep map completion process as a supervised learning problem, utilizing end-to-end network structures to perform completion predictions on input sparse deep maps. These methods, such as MSG-CHN, DeepLiDAR, and NLSPN (Non-Local Spatial Propagation Network), are typically trained on large-scale public datasets (such as the KITTIDepth Completion Benchmark) and achieve leading performance on evaluation metrics. However, these methods often rely on large amounts of high-quality labeled data, have long training cycles, poor generalization ability, and often require retraining or fine-tuning when transferred to real-world applications. Furthermore, their complex network structures, large model parameters, and high runtime resource consumption make them difficult to meet the real-time and deployment requirements of resource-constrained environments such as embedded devices and edge computing platforms.

[0010] (2) Methods based on traditional image processing: These methods often employ classic image processing operators and prior modeling strategies, such as morphological operations (dilation, erosion), bilateral filtering, guided filtering, Poisson interpolation, and gradient diffusion. Their advantages include clear algorithm structure, strong interpretability, low computational complexity, and no need for training, making them particularly suitable for industrial environments where algorithm stability and resource consumption are critical. For example, the hole completion schemes built into industrial vision libraries such as OpenCV are largely based on these methods and are widely used in embedded sensing systems and automotive systems. However, these methods mostly use fixed parameters or static kernel scales, lacking the ability to dynamically adapt to local depth density and edge structures, leading to blurred edges, structural distortion, or mis-completion in the completion results, especially when dealing with complex structures or regions with high gradient changes.

[0011] In summary, although the field of depth completion has developed a relatively mature technical system, existing methods still face certain technical bottlenecks when dealing with depth map completion problems in practical applications. Therefore, there is an urgent need to propose an improved completion scheme that fully leverages the inherent features of the depth map without requiring deep learning training. This would enhance completion quality and practical deployability, providing more reliable dense depth map data support for applications such as autonomous driving, robot navigation, and AR / VR. Summary of the Invention

[0012] This specification provides one or more embodiments of an image optimization method, device, and medium to solve the technical problems mentioned in the background art.

[0013] One or more embodiments of this specification employ the following technical solutions:

[0014] This specification provides an image optimization method according to one or more embodiments, the method comprising:

[0015] The pixels of the depth map are inverted to obtain an inverted depth map;

[0016] Determine the proportion of effective depth pixels occupied by each preset local region in the inverted depth map;

[0017] Based on the occupancy ratio, the dilation kernel size of the depth inversion depth map is adjusted, and dilation operations are performed on each of the preset local regions based on the dilation kernel size to obtain the dilated depth map;

[0018] The expansion depth map is closed to connect the effective depth regions within the local area, fill the structural gaps, and obtain a closed depth map.

[0019] A first hole region is determined in the closed depth map, and the pixels in the first hole region are filled with a first preset size based on a preset selective dilation strategy to obtain a first filled depth map.

[0020] The first filled depth map is pushed upwards and outwards using a top extension strategy to fill in the missing area at the top of the image, thus obtaining a top extended depth map.

[0021] A second hole region is determined in the top extended depth map, and the pixels in the second hole region are filled with a second preset size to obtain a second filling depth map, wherein the second preset size is larger than the first preset size;

[0022] The second fill depth map is processed using a preset joint filtering strategy to obtain a filtered depth map;

[0023] The filtered depth map is then restored to the original depth map that matches the real scene.

[0024] It should be noted that this invention achieves high-quality repair of holes in depth maps by constructing a multi-level adaptive depth completion mechanism, without relying on external RGB guidance or deep learning model training. Specifically, the method first establishes depth distribution perception capability through depth inversion and local region proportion analysis, enabling the dilation kernel size to dynamically adapt to the sparsity characteristics of different regions. Subsequently, through the synergistic effect of closure operation and selective dilation strategy, it fills in hole regions of different scales in stages while preserving the original geometric structure. The top extension strategy specifically compensates for the physical constraints of systematic deficiencies above the scene, and joint filtering ultimately eliminates local inconsistencies in the completion process. This technical path, based on the step-by-step optimization of the depth map's own features, overcomes the edge blurring problem caused by fixed parameters in traditional methods and avoids the strong dependence of deep learning schemes on data and computing power. The final output is a dense depth map with structural coherence, edge clarity, and scene rationality, providing a lightweight and reliable depth repair solution for various intelligent perception systems with high real-time requirements or limited resources.

[0025] Further, determining the occupancy ratio of effective depth pixels in each preset local region of the inverted depth map includes:

[0026] The neighborhood of each hole pixel in the inverted depth map is set as a corresponding preset local region;

[0027] Extract fixed-size windows within each preset local area and calculate the percentage of effective depth pixels occupied.

[0028] It should be noted that this invention constructs a fine-grained spatial distribution perception mechanism by dynamically dividing the neighborhood of hole pixels into preset local regions and statistically analyzing the proportion of effective depth pixels within the window. This method abandons the traditional coarse processing approach with a fixed kernel size, and captures the non-uniform distribution characteristics of the depth map in real time through local window statistics, enabling subsequent processing to adaptively adjust the completion strategy based on the actual data density of different regions. This quantitative analysis based on neighborhood statistics avoids over-completion of sparse regions or under-completion of dense regions by global uniform processing, and can accurately identify the structural features of edge transition areas, providing reliable regional feature support for subsequent adaptive expansion kernel adjustment, thereby achieving reasonable filling of holes in the depth map while maintaining the original geometric structure.

[0029] Furthermore, determining the occupancy ratio of effective depth pixels in each preset local region of the inverted depth map includes:

[0030] Calculation function based on local density The proportion of effective depth pixels in each preset local region of the inverted depth map is determined, where ρ(x,y) is the local density value, representing the proportion of effective depth pixels in the window centered at point (x,y); W and H represent the width and height of the local window in pixels; (i,j) represents the pixel coordinates traversed within the window; and N(i,j) is an indicator function that determines whether the depth value of a pixel within the local window is an effective depth value. τ′ is the preset density threshold, and depth(i,j) is the original depth value of point (i,j).

[0031] It should be noted that this invention achieves intelligent perception of the spatial distribution characteristics of depth maps by introducing a precise quantitative evaluation mechanism based on a local density calculation function. This technical solution constructs a mathematical model that includes window size parameters, a coordinate traversal mechanism, and a validity judgment function. This model not only accurately calculates the effective depth ratio within any window location but also ensures that the local density evaluation results reflect both the statistical characteristics of the data distribution and maintain physical consistency with the real scene through the dual constraints of a preset density threshold and the original depth value. This structured density calculation method overcomes the shortcomings of traditional regional statistical methods, such as sensitivity to window size and significant boundary effects. It provides a regional feature description that is both mathematically rigorous and physically reasonable for subsequent adaptive processing, thereby significantly improving the adaptability of the depth completion process to complex scenes while maintaining the algorithm's lightweight nature.

[0032] Furthermore, adjusting the dilation kernel size of the depth inversion depth map based on the occupancy ratio includes:

[0033] Input the occupancy ratio into formula K size (ρ)=K max-(K max -K min )×ρ(x,y), adjust the dilation kernel size of the depth inversion depth map, where K size (ρ) represents the expansion core size calculated based on the occupancy ratio; ρ(x,y) is the local density value, ranging from [0,1]; K max Indicates the preset maximum core size; K min This indicates the preset minimum core size.

[0034] It should be noted that this invention achieves intelligent control of parameter configuration in the depth map completion process by establishing an adaptive adjustment mechanism for the dilation kernel size based on mathematical formulas. This method utilizes a functional mapping relationship between local density values ​​and a preset kernel size range, enabling the dilation operation to automatically adjust the processing intensity according to the regional data distribution characteristics: a smaller kernel size is used in areas with dense effective pixels to preserve fine structural features, while the kernel size is increased in sparse areas to ensure effective filling. This nonlinear adjustment method avoids the over-dilation or under-dilation problems caused by a fixed kernel size, and ensures algorithm stability through preset boundary constraints. Ultimately, while maintaining computational efficiency, it achieves an optimal balance between structural continuity and detail preservation in the depth map completion result.

[0035] Furthermore, the method also includes:

[0036] The expansion kernel size is set to satisfy an odd number constraint to ensure central symmetry, resulting in the expansion kernel size of the depth-inverted depth map as follows:

[0037]

[0038] It should be noted that this invention ensures geometric symmetry and structural stability during depth map processing by introducing an odd-number constraint mechanism for the dilation kernel size. This technical solution mandates that the dilation kernel size maintain an odd number, ensuring that each processing window has a clearly defined center pixel, thereby maintaining the symmetry and predictability of spatial transformations in morphological operations. This centrally symmetric dilation process not only avoids pixel shifts and structural distortions that may occur with even-number kernel sizes, but also guarantees the uniform propagation of depth values ​​during region dilation. This allows subsequent depth completion operations to better maintain the geometric consistency of the original scene, ultimately outputting a depth map with higher structural integrity and spatial accuracy.

[0039] Furthermore, before obtaining the filtered depth map, the method further includes:

[0040] The gradient is calculated using the Scharr operator, and its kernel function formula is as follows:

[0041] in,

[0042] G x G represents the gradient response in the horizontal direction, used for detecting vertical edges; y This represents the gradient response in the vertical direction, used for detecting horizontal edges; D is the input depth map.

[0043] The formula for calculating gradient magnitude is:

[0044] in,

[0045] G(x,y) represents the gradient magnitude at pixel (x,y), indicating the edge strength of that pixel.

[0046] Normalization is performed as follows:

[0047] in,

[0048] G norm (x,y) is the normalized gradient magnitude, used for subsequent calculation of edge detection threshold and hybrid optimization factor.

[0049] It should be noted that this invention constructs a precise extraction system for depth map edge features by introducing a multi-directional gradient calculation and normalization mechanism based on the Scharr operator. This method utilizes the orthogonality of horizontal and vertical kernel functions to achieve highly sensitive detection of anisotropic edge structures in the depth map. Through gradient magnitude calculation and normalization, the original depth variations are transformed into standardized edge intensity representations, providing physically meaningful edge constraint information for subsequent filtering optimization. This differential operator-based edge extraction scheme not only overcomes the shortcomings of traditional gradient detection methods in terms of noise sensitivity and directional bias, but its normalization process also ensures the comparability of edge intensity under different scenarios. This allows the subsequent hybrid filtering process to dynamically adjust the optimization intensity based on real geometric features, ultimately preserving key structural information in the depth map while suppressing noise.

[0050] Furthermore, before obtaining the filtered depth map, the method further includes:

[0051] Median filtering is used to preprocess the global pixel to remove salt-and-pepper noise, so that the statistical median of neighboring pixels can be used to replace the current pixel value. The formula is as follows:

[0052] D med (x,y)=median{D(x+i,y+i)|(i,j)∈W}; where,

[0053] D(x,y) represents the pixel value of the input depth map at coordinates (x,y), W is the preset window size, and D med(x, y) is the depth map output after median filtering; (i, j) represents the coordinates of the pixel traversed in the window, that is, i represents the abscissa and j represents the ordinate.

[0054] The threshold T is set to distinguish the edge region E and the flat region F, and Gaussian filtering with different intensities is performed for different regions. For the edge region E(G norm >T), sharpening processing is adopted to compensate the blur caused by filtering and enhance edge details. The formula is:

[0055] D E (x,y)=(1+λ′)·D med (x,y)-λ′·G σ (D med (x,y)); wherein,

[0056] D E (x,y) represents the pixel value after sharpening processing, λ′ is the sharpening coefficient, which is used to control the edge sharpening intensity, G σ is a Gaussian kernel with a standard deviation of σ;

[0057] For the flat region F(G norm <T), strong smoothing processing is adopted to suppress residual noise and avoid introducing artifacts at the same time. The formula is:

[0058] D F (x,y)=G σ′ (D med (x,y)); wherein,

[0059] D F (x,y) represents the pixel value after smoothing processing, σ′ is the standard deviation of the Gaussian filter for the flat region, and it is necessary to ensure σ′>σ to guarantee a stronger smoothing effect.

[0060] It should be noted that the present invention realizes the collaborative optimization of depth map noise suppression and detail preservation by constructing an adaptive filtering system based on regional characteristics. In the method of the present invention, median filtering is firstly adopted for global preprocessing, which effectively eliminates the interference of discrete noise points on subsequent processing; then the image is divided into edge regions and flat regions with greatly different structural characteristics through an edge detection threshold, and differentiated processing strategies are respectively applied: Gaussian filtering with sharpening compensation is implemented on edge regions to strengthen geometric features, and strong smoothing processing is adopted on flat regions to completely suppress noise. This sub-regional processing mechanism dynamically adjusts the filtering intensity and sharpening coefficient, which not only overcomes the problems of edge blurring or noise residue caused by the traditional single filtering method, but also ensures the accurate matching between processing intensity and regional characteristics through the adaptive configuration of Gaussian kernel standard deviation. Finally, while maintaining the computational efficiency of the algorithm, the depth map achieves an optimal balance between overall smoothness and structural integrity.

[0061] Furthermore, before obtaining the filtered depth map, the method further includes:

[0062] Define an adaptive blending factor α, which dynamically adjusts the filter strength based on the gradient magnitude. The formula is as follows:

[0063] in,

[0064] G norm (x,y) represents the normalized gradient magnitude, indicating the edge strength at that pixel; α represents the adaptive blending factor, used to control the fusion weights of the original depth map and the filtered result, and to determine the smoothness of each pixel. The smaller α is, the stronger the edge region features and the more dependent it is on the original depth value; the larger α is, the stronger the flat region features and the more dependent it is on the filtered depth value; ε represents a minimal positive constant to avoid the denominator being zero and to control the gradient sensitivity, preventing extreme fusion situations caused by over-adjustment of the blending factor; n represents the shape parameter, which adjusts the nonlinearity of the blending factor as it changes with the gradient.

[0065] The fusion ratio of the original depth map and the filtered depth map is dynamically adjusted based on the blending factor α. The formula is as follows:

[0066] D final (x,y)=α·D raw (x,y)+(1-α)·D Gauss (x,y); where,

[0067] D final This represents the final optimized pixel value.

[0068] It should be noted that this invention achieves an intelligent balance between edge preservation and noise suppression during depth map filtering by constructing a gradient-sensitive adaptive fusion mechanism. This method innovatively introduces a nonlinear fusion factor function, transforming the normalized gradient magnitude into continuously adjustable fusion weights. This allows the system to automatically adjust the contribution ratio of the original data and the filtered result based on the structural characteristics of local regions—prioritizing the retention of original depth values ​​in areas with significant edge intensity to maintain geometric sharpness, while emphasizing the filtered result in flat areas to ensure noise suppression. This gradient-driven dynamic fusion strategy, through fine-tuning of shape parameters and minimal constants, avoids the edge blurring or noise residue problems caused by traditional fixed-weight fusion, and prevents drastic fluctuations in the fusion ratio caused by gradient abrupt changes. Ultimately, the output depth map achieves optimal smoothness and consistency while maintaining the integrity of structural features, providing a depth data foundation with both geometric accuracy and noise robustness for subsequent 3D vision tasks.

[0069] This specification provides an image optimization device according to one or more embodiments, comprising:

[0070] At least one processor; and,

[0071] A memory communicatively connected to the at least one processor; wherein,

[0072] The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to:

[0073] The pixels of the depth map are inverted to obtain an inverted depth map;

[0074] Determine the proportion of effective depth pixels occupied by each preset local region in the inverted depth map;

[0075] Based on the occupancy ratio, the dilation kernel size of the depth inversion depth map is adjusted, and dilation operations are performed on each of the preset local regions based on the dilation kernel size to obtain the dilated depth map;

[0076] The expansion depth map is closed to connect the effective depth regions within the local area, fill the structural gaps, and obtain a closed depth map.

[0077] A first hole region is determined in the closed depth map, and the pixels in the first hole region are filled with a first preset size based on a preset selective dilation strategy to obtain a first filled depth map.

[0078] The first filled depth map is pushed upwards and outwards using a top extension strategy to fill in the missing area at the top of the image, thus obtaining a top extended depth map.

[0079] A second hole region is determined in the top extended depth map, and the pixels in the second hole region are filled with a second preset size to obtain a second filling depth map, wherein the second preset size is larger than the first preset size;

[0080] The second fill depth map is processed using a preset joint filtering strategy to obtain a filtered depth map;

[0081] The filtered depth map is then restored to the original depth map that matches the real scene.

[0082] This specification provides one or more embodiments of a non-volatile computer storage medium storing computer-executable instructions, which, when executed by a computer, can perform the following:

[0083] The pixels of the depth map are inverted to obtain an inverted depth map;

[0084] Determine the proportion of effective depth pixels occupied by each preset local region in the inverted depth map;

[0085] Based on the occupancy ratio, the dilation kernel size of the depth inversion depth map is adjusted, and dilation operations are performed on each of the preset local regions based on the dilation kernel size to obtain the dilated depth map;

[0086] The expansion depth map is closed to connect the effective depth regions within the local area, fill the structural gaps, and obtain a closed depth map.

[0087] A first hole region is determined in the closed depth map, and the pixels in the first hole region are filled with a first preset size based on a preset selective dilation strategy to obtain a first filled depth map.

[0088] The first filled depth map is pushed upwards and outwards using a top extension strategy to fill in the missing area at the top of the image, thus obtaining a top extended depth map.

[0089] A second hole region is determined in the top extended depth map, and the pixels in the second hole region are filled with a second preset size to obtain a second filling depth map, wherein the second preset size is larger than the first preset size;

[0090] The second fill depth map is processed using a preset joint filtering strategy to obtain a filtered depth map;

[0091] The filtered depth map is then restored to the original depth map that matches the real scene.

[0092] The above-described at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects:

[0093] This invention achieves high-quality repair of holes in depth maps by constructing a multi-level adaptive depth completion mechanism, without relying on external RGB guidance or deep learning model training. Specifically, the method first establishes depth distribution perception capabilities through depth inversion and local region proportion analysis, enabling the expansion kernel size to dynamically adapt to the sparsity characteristics of different regions. Subsequently, through the synergistic effect of closure operations and selective expansion strategies, it fills in hole regions of different scales in stages while preserving the original geometric structure. The top extension strategy specifically compensates for the physical constraints of systematic deficiencies above the scene, and joint filtering ultimately eliminates local inconsistencies during the completion process. This technical path, based on the step-by-step optimization of the depth map's own features, overcomes the edge blurring problem caused by fixed parameters in traditional methods and avoids the strong dependence of deep learning schemes on data and computing power. The final output is a dense depth map with structural coherence, edge clarity, and scene rationality, providing a lightweight and reliable depth repair solution for various intelligent perception systems with high real-time requirements or limited resources. Attached Figure Description

[0094] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:

[0095] Figure 1 A flowchart illustrating an image optimization method provided in one or more embodiments of this specification;

[0096] Figure 2 This is a schematic diagram of adaptive expansion optimization provided for one or more embodiments of this specification;

[0097] Figure 3 A flowchart of the jointly optimized algorithm provided for one or more embodiments of this specification;

[0098] Figure 4 This is a schematic diagram of the structure of an image optimization device provided for one or more embodiments of this specification. Detailed Implementation

[0099] This specification provides an image optimization method, device, and medium through its embodiments.

[0100] In depth map completion tasks, while traditional image processing methods do not have the overwhelming advantage in accuracy as deep learning-based methods, they are still widely used in practical engineering scenarios such as autonomous driving, robot vision, and industrial 3D measurement due to their lower computational resource requirements, good interpretability, and ease of embedded deployment. They are particularly suitable for edge computing platforms or resource-constrained environments. Existing traditional methods mainly rely on morphological operations, image filtering techniques, and local interpolation strategies from the field of image processing to expand depth information and fill holes through methods such as constructing structural rules, spatial propagation, and edge guidance.

[0101] In the existing technological system, the IP-Basic (Image Processing Basic) method is one of the most representative and practical traditional depth completion schemes. It was first proposed by Jason Ku et al. in their early research on autonomous driving vision systems and has been widely disseminated. Based on the OpenCV image processing library, this method employs a series of classic morphological operations and filtering processes to gradually realize a complete processing path from local hole filling to full image densification. It has outstanding advantages such as simple structure, stable algorithm, high computational efficiency, and ease of implementation, and is especially suitable for embedded real-time processing systems in CPU environments.

[0102] Specifically, the IP-Basic method mainly includes the following eight consecutive processing steps, which strive to preserve local structural features and achieve edge smoothing while gradually diffusing depth information: ① Perform depth inversion to protect the edge information of nearby objects; ② Fill adjacent holes with custom kernel dilation; ③ Connect local structures using pinhole closure operations; ④ Use selective pinhole filling to fill missing areas and avoid erroneous modification of valid pixels; ⑤ Use a top extension strategy to push the depth value upwards to fill the missing areas at the top of the image; ⑥ For large remaining holes, further densification is achieved through large-hole filling; ⑦ Use median + Gaussian joint filtering to remove noise introduced by dilation and smooth edges; ⑧ Perform depth inversion to restore to the true scale. This method completes the process from shallow to deep and from local to global, taking efficiency into account, and achieves excellent results on the KITTI dataset. However, this method does not consider the impact of density changes on depth completion and has poor ability to preserve the edge information of objects. This paper optimizes the IP-Basic method to address two aspects: the error caused by ignoring local density changes in the depth completion process based on traditional image processing, and the balance between preserving object edge information and denoising.

[0103] The current method has the following limitations:

[0104] (1) The dilation operation does not consider the differences in local structures, and the fixed kernel size can easily lead to excessive expansion of the structural boundaries. It ignores the spatial variations in structural complexity, hole density and edge gradient in different regions of the depth map. Near the structural boundaries, if the dilation operation expands the adjacent effective pixels indiscriminately, it is very easy to cause edge "diffusion" or "leakage", resulting in blurred object contours, inaccurate boundaries, and even misconnection of structures at the junction, which seriously affects the downstream performance of 3D reconstruction and target recognition tasks.

[0105] (2) Most filtering methods are based on global strategies and cannot distinguish between edge regions and flat regions, which can easily lead to edge blurring or error propagation. During the completion process, filtering methods are often used to denoise and smooth the completion results. However, these methods are often based on uniform window scale and standard deviation parameters, making it difficult to flexibly adjust the processing strategy according to the local characteristics of the pixel's region (such as edge strength and texture complexity). This results in "false smoothing" in edge regions, where edge details are erased; while in flat regions, the filtering strength may be insufficient, leaving residual noise or interpolation artifacts. Furthermore, global parameter settings often require trade-offs and cannot simultaneously meet the dual requirements of edge preservation and denoising for different regions of the image.

[0106] (3) The lack of a dynamic adjustment mechanism makes it difficult to cope with complex scenarios where high-density and low-density regions coexist in depth maps. In practical applications, sparse depth maps often exhibit uneven spatial density. For example, LiDAR points are denser in near-field regions and sparser in far-field regions, or local holes may be significantly enlarged due to occlusion. Current methods generally use uniform parameter settings (such as fixed dilation times and filtering radius), failing to dynamically adjust the completion behavior according to the local effective pixel density. Low-density regions are prone to "hole residue" due to insufficient completion scale; high-density regions may be overprocessed, producing pseudo-structures or information redundancy, reducing completion accuracy and efficiency, and making it difficult to adapt to complex scenarios with extremely high requirements for environmental understanding, such as autonomous driving and unmanned navigation.

[0107] (4) Lack of structural prior aids under unguided conditions, especially poor performance under extremely sparse conditions. Traditional unguided methods rely solely on the depth map's own information for completion, lacking structural prior support from RGB images or other multimodal information. In extreme scenarios such as extremely sparse depth maps, missing edge information, or strong occlusion between objects, traditional methods struggle to infer reasonable geometric structures, easily leading to structural connection errors, object breaks, or over-interpolation, affecting the coherence and reliability of overall 3D understanding. Especially in autonomous driving perception systems, mis-completed structures may cause misjudgments, thus posing safety hazards to path planning and obstacle detection.

[0108] This invention aims to systematically address several key issues in current depth map completion techniques based on traditional image processing methods, including but not limited to poor structural adaptability, limited hole filling effects, weak ability to preserve edge details, and insufficient response to local features of the depth map. In practical applications, especially in the environmental perception modules of autonomous driving systems, the spatial mapping and path planning stages of robot navigation, and the strong reliance on high-fidelity depth information in 3D modeling and reconstruction tasks, the accuracy and robustness of depth map completion become core factors directly affecting system performance. However, traditional methods based on morphological operations and static filtering strategies often employ fixed kernel sizes and uniform filtering rules, neglecting the objective characteristics of uneven spatial density distribution and drastic changes in edge features within the depth map. This leads to structural distortion, blurred edges, loss of details, and even misfilling during the completion process, making it difficult to meet the practical needs of the aforementioned high-precision scenarios.

[0109] To address this, this invention proposes an optimized depth completion method that integrates "local structure awareness" and "regional adaptive response" mechanisms. Specifically, it includes two key innovative modules: First, an adaptive dilation method based on local density distribution dynamically adjusts the size of the dilation kernel by statistically analyzing the proportion of effective depth values ​​(i.e., local depth density) within the neighborhood of each pixel. This enhances the expansion capability of the dilation operation in sparse depth regions and preserves structural boundaries in dense regions, effectively improving the overall detail and continuity of the completion. Second, a gradient-aware regional adaptive filtering method introduces image gradient information to guide the structure of the depth map, dividing the entire image into edge and non-edge regions. Differentiated filtering strategies (such as sharpening or Gaussian smoothing) are then applied to each region, combined with a hybrid weighting mechanism, to achieve precise preservation of edge features and effective suppression of noisy regions.

[0110] The technical solution of this invention is not a simple superposition of single optimization modules, but rather a synergistic collaboration between dilation operations and filtering strategies within a unified algorithm framework. By decoupling information structures and reconstructing local features, it enhances the adaptability and robustness of traditional image processing methods in depth completion. The overall design not only retains the advantages of traditional methods such as low resource consumption, no training required, and strong interpretability, but also verifies its versatility and effectiveness in complex scenarios on multiple typical datasets (such as KITTI), demonstrating broad engineering applicability and theoretical significance.

[0111] This invention aims to address the prominent problems in current depth map completion methods based on traditional image processing techniques, such as poor local structure perception, insufficient hole filling accuracy, weak edge preservation, and rigid processing strategies. It proposes an optimized completion method with adaptive capabilities and structure perception characteristics to meet the comprehensive requirements of accuracy, efficiency, and stability in complex application scenarios. The specific objectives include the following four aspects:

[0112] (1) An adaptive dilation method based on local density distribution is proposed to achieve dynamic selection of kernel size, finely fill holes, and maintain edge structure. Traditional dilation operations mostly use structural elements of fixed size, which cannot adapt to the non-uniformity of hole distribution in depth maps. This invention innovatively introduces a local density modeling mechanism, which calculates the effective depth pixel density in the neighborhood of each hole to accurately quantify the sparsity and dynamically adjust the size and shape of the dilation kernel accordingly. Small kernels are used for fine filling in high-density areas to avoid excessive structural expansion; while in low-density or large-area hole areas, larger-scale dilation kernels are selected to improve filling speed and coverage. In addition, this method introduces an edge directionality protection strategy, which restricts the growth direction of the kernel to avoid crossing object boundaries or introducing pseudo-structures, effectively improving the geometric consistency and local coherence of edge regions.

[0113] (2) A region-adaptive filtering mechanism is introduced, which combines image gradient to accurately divide edge and flat regions, realizing a differentiated filtering strategy that balances denoising and structure preservation. Existing methods often use uniform filtering parameters (such as window size, standard deviation, etc.) to perform noise suppression processing on the entire image, ignoring the differences in structural complexity and texture features between different regions of the image, which can easily lead to edge blurring or error propagation. To this end, this invention proposes a local structure recognition mechanism based on image gradient response. By analyzing the gradient magnitude and direction information of the completed image, adaptive division of edge and flat regions is achieved. Based on this, differentiated filtering strategies are designed: a strong smoothing operation is applied to flat regions to remove isolated noise, while a guided filter with strong edge preservation or a light median filter is used to preserve structural information in edge regions, maximizing the balance between geometric clarity and denoising performance of the completed result.

[0114] (3) Two optimization methods are integrated within the same framework to achieve collaborative processing and overall performance improvement. The two optimization strategies mentioned above address key issues in the expansion and filtering stages, respectively, and have independent implementation value. To further improve the overall system performance and practical application efficiency, this invention integrates "adaptive expansion" and "region filtering" into a unified processing framework, realizing a continuous closed-loop processing mechanism from density-aware completion to structure-aware filtering. The specific process includes: hole region identification → local density modeling → dynamic adjustment of expansion kernel → preliminary completion → image gradient extraction → region classification → adaptive filtering → structure refinement. This collaborative process has good stage coupling and module compatibility, ensuring the controllability of processing logic while significantly improving the robustness and structural consistency of the completion process, making it particularly suitable for perception systems that balance real-time performance and accuracy requirements.

[0115] (4) A novel deep graph completion scheme that requires no training, consumes few resources, and is structurally adaptable is provided. Current research on deep graph completion largely focuses on deep learning methods. While these methods have achieved significant performance improvements under large-scale data-driven conditions, their versatility, interpretability, and resource dependence remain serious problems, making it difficult to meet the application needs of embedded devices, low-power platforms, or environments without network support. This invention explicitly avoids neural network dependence, implementing the entire process based on traditional image processing algorithms without any model training process, thus offering high deployment flexibility. Furthermore, the scheme design fully considers operational efficiency and memory overhead; all core operations can be smoothly executed on conventional CPUs or low-power edge computing units, greatly expanding the practical application potential of this technology in fields such as unmanned vehicles, robots, AR devices, and 3D scanners.

[0116] Furthermore, this method possesses excellent scalability and adjustability, allowing for dynamic adjustment of key parameters such as density threshold, gradient classification criteria, and structural kernel function to meet the needs of different scenarios. This enables flexible application across multiple scenarios and accuracy levels, making it a general-purpose depth map completion solution that combines engineering practicality with technological innovation.

[0117] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0118] Figure 1 This diagram illustrates a flowchart of an image optimization method provided in one or more embodiments of this specification. This process can be executed by an image optimization system. Certain input parameters or intermediate results in the process can be manually adjusted to help improve accuracy.

[0119] The method flow steps of the embodiments in this specification are as follows:

[0120] S101, invert the pixels of the depth map to obtain an inverted depth map.

[0121] In embodiment S101 of this specification, invalid pixels in the sparse depth map are typically 0, while the valid pixels of nearby objects are numerically very close to 0. This can lead to the situation where, when directly using OpenCV morphological operations on these objects, distant pixels may cover nearby pixels, resulting in the loss of edge information of nearby objects. To avoid this problem, all valid pixels are first inverted, and a preset offset (e.g., 20 meters) is introduced to separate valid and invalid pixels, ensuring that the edges of nearby objects are not destroyed during morphological operations. The specific inversion formula is as follows:

[0122] D inverted =100-D input , where D input This represents the original depth map as input; D inverted This represents the depth map after flipping.

[0123] S102, determine the occupancy ratio of effective depth pixels in each preset local region of the inverted depth map.

[0124] S103, based on the occupancy ratio, adjust the dilation kernel size of the depth inversion depth map, and perform dilation operation on each of the preset local regions based on the dilation kernel size to obtain the dilated depth map.

[0125] In embodiments S102-S103 of this specification, the optimization method takes each hole pixel as the processing object. First, it performs local density analysis: extracting a fixed-size window in the neighborhood of each hole pixel and statistically analyzing the proportion of effective depth pixels within it, which serves as a density index to measure the integrity of regional information and is the core guiding basis for subsequent operations. Second, based on the results of the local density analysis, the size of the dilation kernel is adaptively adjusted, following the principle of "the lower the density, the larger the dilation kernel; the higher the density, the smaller the dilation kernel." Finally, to ensure the symmetry and stability of morphological operations, the kernel size must be an odd number to meet the structural symmetry requirements. A schematic diagram of the adaptive dilation optimization is shown below. Figure 2 As shown. Black dots represent valid pixels, and the surrounding gray area represents the extent of the dilation kernel constructed centered on the black dots. In high-density areas, the dilation kernel size is 3×3, and in low-density areas, the dilation kernel size is 5×5.

[0126] S104, perform a closure operation on the expansion depth map to connect the effective depth regions within the local area, fill the structural gaps, and obtain a closed depth map.

[0127] In embodiment S104 of this specification, after the initial expansion described above, some pixel-level holes still exist in the depth map. To connect these local hole regions, a 5×5 full-kernel morphological closing operation can be used to connect the effective depth regions within the local neighborhood and fill the structural gaps.

[0128] S105, determine the first hole region in the closed depth map, and fill the pixels in the first hole region with a first preset size based on a preset selective dilation strategy to obtain a first filling depth map.

[0129] In embodiment S105 of this specification, some void areas in the depth map cannot be filled by the first two dilation operations. For these voids, a selective dilation strategy can be used, employing a 7×7 dilation kernel for dilation, but only updating the depth value of the void pixels, while leaving other already assigned pixels unchanged. This strategy avoids interfering with correctly estimated areas, improving the filling rate while maintaining structural realism.

[0130] S106, the first filling depth map is applied with a top extension strategy to push the depth value upwards and outwards to fill in the missing area above the image, thus obtaining a top extension depth map.

[0131] In embodiment S106 of this specification, in order to handle tall objects (such as trees and pillars) that exceed the vertical scanning range of the lidar (usually the lower two-thirds of the image), a "column extension" strategy can be introduced: the first effective depth value in each column is pushed upwards and outwards to the top of the image, reasonably inferring and supplementing the missing depth information at the top, so that the final generated depth map is suitable for application scenarios including the detection of tall objects.

[0132] S107, determine the second hole region in the top extended depth map, fill the pixels in the second hole region with a second preset size to obtain a second filling depth map, wherein the second preset size is greater than the first preset size.

[0133] In embodiment S107 of this specification, after the previous few steps of local filling, some large-scale void areas still remain. A 31×31 large-size kernel can be used for dilation operation to perform final filling. By copying the adjacent depth values ​​over a large area, the entire image is densified, while minimizing the damage to the overall consistency caused by local abnormal diffusion.

[0134] S108, the second filling depth map is processed using a preset joint filtering strategy to obtain a filtered depth map.

[0135] In embodiment S108 of this specification, the dilation operation introduces outliers while increasing coverage. To remove these outliers and enhance the realism of the output depth map, a joint filtering strategy can be adopted: first, a 5×5 median filter is used to remove salt-and-pepper noise, and then a 5×5 Gaussian filter is used to smooth local surfaces and soften object edges, so that the output conforms to the piecewise planar characteristics of the real scene.

[0136] Traditional methods for optimizing depth maps typically employ various filter combinations, such as median filtering + Gaussian filtering, median filtering + guided filtering, NLM filtering + Gaussian filtering, wavelet transform filtering + Gaussian filtering, and lightweight BM3D filtering. All of these combinations offer noise reduction, edge preservation, and overall smoothing capabilities. Building upon these traditional filtering combinations, three further optimization strategies can be introduced: precise edge detection, region differentiation processing, and dynamic hybrid optimization.

[0137] S109, restore the filtered depth map to the original depth map that conforms to the real scene.

[0138] In embodiment S109 of this specification, the processed depth map is restored to the original scale, through D... output =100-D blurred The depth map is then inverted to output a final, complete, dense depth map that matches the real-world scene. output This represents the final output dense depth map; D blurredThis represents the filtered depth map;

[0139] It's worth noting that the entire process follows a strategy from near to far and from local to global, effectively balancing the trade-off between preserving details and filling large areas of holes. This algorithm can run at up to 90Hz on an Intel Core i7-7700K CPU, making it suitable for deployment in real-world embedded systems. Furthermore, it outperforms several deep learning-based methods, such as NN+CNN, on the KITTI deep completion benchmark.

[0140] It should be noted that this invention achieves high-quality repair of holes in depth maps by constructing a multi-level adaptive depth completion mechanism, without relying on external RGB guidance or deep learning model training. Specifically, the method first establishes depth distribution perception capability through depth inversion and local region proportion analysis, enabling the dilation kernel size to dynamically adapt to the sparsity characteristics of different regions. Subsequently, through the synergistic effect of closure operation and selective dilation strategy, it fills in hole regions of different scales in stages while preserving the original geometric structure. The top extension strategy specifically compensates for the physical constraints of systematic deficiencies above the scene, and joint filtering ultimately eliminates local inconsistencies in the completion process. This technical path, based on the step-by-step optimization of the depth map's own features, overcomes the edge blurring problem caused by fixed parameters in traditional methods and avoids the strong dependence of deep learning schemes on data and computing power. The final output is a dense depth map with structural coherence, edge clarity, and scene rationality, providing a lightweight and reliable depth repair solution for various intelligent perception systems with high real-time requirements or limited resources.

[0141] Furthermore, when determining the occupancy ratio of effective depth pixels in each preset local region of the inverted depth map, the neighborhood of each hole pixel in the inverted depth map can be set as the corresponding preset local region; a fixed-size window is extracted in each preset local region, and the occupancy ratio of effective depth pixels is statistically analyzed.

[0142] It should be noted that this invention constructs a fine-grained spatial distribution perception mechanism by dynamically dividing the neighborhood of hole pixels into preset local regions and statistically analyzing the proportion of effective depth pixels within the window. This method abandons the traditional coarse processing approach with a fixed kernel size, and captures the non-uniform distribution characteristics of the depth map in real time through local window statistics, enabling subsequent processing to adaptively adjust the completion strategy based on the actual data density of different regions. This quantitative analysis based on neighborhood statistics avoids over-completion of sparse regions or under-completion of dense regions by global uniform processing, and can accurately identify the structural features of edge transition areas, providing reliable regional feature support for subsequent adaptive expansion kernel adjustment, thereby achieving reasonable filling of holes in the depth map while maintaining the original geometric structure.

[0143] Furthermore, when determining the proportion of effective depth pixels in each preset local region of the inverted depth map, a local density calculation function can be used. The proportion of effective depth pixels in each preset local region of the inverted depth map is determined, where ρ(x,y) is the local density value, representing the proportion of effective depth pixels in the window centered at point (x,y); W and H represent the width and height of the local window in pixels; (i,j) represents the pixel coordinates traversed within the window; and N(i,j) is an indicator function that determines whether the depth value of a pixel within the local window is an effective depth value. τ′ is a preset density threshold, and depth(i,j) is the original depth value of point (i,j). The above formula can be used to calculate the proportion of effective depth pixels within a local window centered on any pixel, providing necessary local feature information for subsequent adaptive processing.

[0144] It should be noted that this invention achieves intelligent perception of the spatial distribution characteristics of depth maps by introducing a precise quantitative evaluation mechanism based on a local density calculation function. This technical solution constructs a mathematical model that includes window size parameters, a coordinate traversal mechanism, and a validity judgment function. This model not only accurately calculates the effective depth ratio within any window location but also ensures that the local density evaluation results reflect both the statistical characteristics of the data distribution and maintain physical consistency with the real scene through the dual constraints of a preset density threshold and the original depth value. This structured density calculation method overcomes the shortcomings of traditional regional statistical methods, such as sensitivity to window size and significant boundary effects. It provides a regional feature description that is both mathematically rigorous and physically reasonable for subsequent adaptive processing, thereby significantly improving the adaptability of the depth completion process to complex scenes while maintaining the algorithm's lightweight nature.

[0145] Furthermore, when adjusting the dilation kernel size of the depth inversion depth map based on the occupancy ratio, the occupancy ratio can be input into the following adaptive density formula:

[0146] K size (ρ)=K max -(K max -K min )×ρ(x,y); Adjust the dilation kernel size of the depth inversion depth map, where K size (ρ) represents the expansion core size calculated based on the occupancy ratio; ρ(x,y) is the local density value, ranging from [0,1]; K max Indicates the preset maximum core size; K min This represents the preset minimum kernel size. The core of the optimization method of this invention lies in the design of an adaptive kernel, which achieves dynamic changes in the expanding kernel through an adaptive density formula.

[0147] It should be noted that this invention achieves intelligent control of parameter configuration in the depth map completion process by establishing an adaptive adjustment mechanism for the dilation kernel size based on mathematical formulas. This method utilizes a functional mapping relationship between local density values ​​and a preset kernel size range, enabling the dilation operation to automatically adjust the processing intensity according to the regional data distribution characteristics: a smaller kernel size is used in areas with dense effective pixels to preserve fine structural features, while the kernel size is increased in sparse areas to ensure effective filling. This nonlinear adjustment method avoids the over-dilation or under-dilation problems caused by a fixed kernel size, and ensures algorithm stability through preset boundary constraints. Ultimately, while maintaining computational efficiency, it achieves an optimal balance between structural continuity and detail preservation in the depth map completion result.

[0148] Furthermore, in the process of adjusting the dilation kernel size of the depth-inverted depth map in this application, the dilation kernel size can be set to satisfy an odd number constraint to ensure central symmetry, resulting in the following dilation kernel size for the depth-inverted depth map:

[0149]

[0150] It should be noted that this invention ensures geometric symmetry and structural stability during depth map processing by introducing an odd-number constraint mechanism for the dilation kernel size. This technical solution mandates that the dilation kernel size maintain an odd number, ensuring that each processing window has a clearly defined center pixel, thereby maintaining the symmetry and predictability of spatial transformations in morphological operations. This centrally symmetric dilation process not only avoids pixel shifts and structural distortions that may occur with even-number kernel sizes, but also guarantees the uniform propagation of depth values ​​during region dilation. This allows subsequent depth completion operations to better maintain the geometric consistency of the original scene, ultimately outputting a depth map with higher structural integrity and spatial accuracy.

[0151] Furthermore, edge detection is a crucial step in depth map optimization, determining the accuracy of edge and flat region segmentation. Common edge detection methods include the Sobel, Scharr, and Laplacian operators. The Sobel operator is computationally simple but has weak detection capability in the diagonal direction; the Laplacian operator can detect changes in all directions but is overly sensitive to noise, easily causing false edges. In contrast, the Scharr operator, through optimized convolution kernel coefficient distribution, enhances edge response while better suppressing noise, resulting in higher gradient accuracy and directional consistency. Therefore, the Scharr operator is chosen for gradient calculation to achieve more accurate edge detection. Before obtaining the filtered depth map, the Scharr operator can be used for gradient calculation; its kernel function formula is:

[0152] in,

[0153] G x G represents the gradient response in the horizontal direction, used for detecting vertical edges; y This represents the gradient response in the vertical direction, used for detecting horizontal edges; D is the input depth map.

[0154] The formula for calculating gradient magnitude is:

[0155] in,

[0156] G(x,y) represents the gradient magnitude at pixel (x,y), indicating the edge strength of that pixel.

[0157] Normalization is performed as follows:

[0158] in,

[0159] G norm (x,y) is the normalized gradient magnitude, used for subsequent calculation of edge detection threshold and hybrid optimization factor.

[0160] It should be noted that this invention constructs a precise extraction system for depth map edge features by introducing a multi-directional gradient calculation and normalization mechanism based on the Scharr operator. This method utilizes the orthogonality of horizontal and vertical kernel functions to achieve highly sensitive detection of anisotropic edge structures in the depth map. Through gradient magnitude calculation and normalization, the original depth variations are transformed into standardized edge intensity representations, providing physically meaningful edge constraint information for subsequent filtering optimization. This differential operator-based edge extraction scheme not only overcomes the shortcomings of traditional gradient detection methods in terms of noise sensitivity and directional bias, but its normalization process also ensures the comparability of edge intensity under different scenarios. This allows the subsequent hybrid filtering process to dynamically adjust the optimization intensity based on real geometric features, ultimately preserving key structural information in the depth map while suppressing noise.

[0161] Furthermore, this invention can employ a gradient-based regional differentiation processing strategy. Before obtaining the filtered depth map, median filtering can be used to preprocess the entire image to remove salt-and-pepper noise, so that the statistical median of neighboring pixels can be used to replace the current pixel value. The formula is as follows:

[0162] D med (x,y)=median{D(x+i,y+i)|(i,j)∈W}; where (i,j) represents the pixel coordinates traversed within the window.

[0163] D(x,y) represents the pixel value of the input depth map at coordinates (x,y), W is the preset window size, and Dmed (x,y) is the depth map output after median filtering;

[0164] Second, the threshold T can be set to distinguish between the edge region E and the flat region F, and Gaussian filtering with different intensities is performed for different regions. For edge region E(G norm >T), sharpening processing is adopted to compensate for the blurring caused by filtering and enhance edge details. The formula is:

[0165] D E (x,y)=(1+λ′)·D med (x,y)-λ′·G σ (D med (x,y)); wherein,

[0166] D E (x,y) represents the pixel value after sharpening processing, λ′ is the sharpening coefficient for controlling the edge sharpening intensity, and in the present invention, λ′ can be 0.5, G σ is a Gaussian with standard deviation σ, and in the present invention, σ can be 0.5 during sharpening processing. This operation can highlight the gray level jump of edges by enhancing the difference between the original image and the Gaussian blurred image, and avoid excessive smoothing of edges by traditional filtering;

[0167] For flat region F(G norm <T), strong smoothing processing is adopted to suppress residual noise and avoid introducing artifacts at the same time. The formula is:

[0168] D F (x,y)=G σ′ (D med (x,y)); wherein,

[0169] D F (x,y) represents the pixel value after smoothing processing, σ′ is the standard deviation of the Gaussian filter for the flat region, and σ′>σ is ensured to achieve a stronger smoothing effect, and in the present invention, σ′ is 1.5.

[0170] It should be noted that this invention achieves synergistic optimization of depth map noise suppression and detail preservation by constructing an adaptive filtering system based on regional characteristics. The method first employs median filtering for global preprocessing, effectively eliminating the interference of discrete noise points on subsequent processing. Then, by using an edge detection threshold, the image is divided into edge regions and flat regions with distinct structural features, and differentiated processing strategies are applied to each: Gaussian filtering with sharpening compensation is applied to edge regions to enhance geometric features, while strong smoothing is used to thoroughly suppress noise in flat regions. This regional processing mechanism, by dynamically adjusting the filtering intensity and sharpening coefficient, overcomes the edge blurring or noise residue problems caused by traditional single filtering methods. Furthermore, the adaptive configuration of the Gaussian kernel standard deviation ensures precise matching between processing intensity and regional features, ultimately achieving an optimal balance between overall smoothness and structural integrity of the depth map while maintaining the algorithm's computational efficiency.

[0171] Furthermore, regarding dynamic fusion optimization, after Gaussian filtering the depth map, this paper argues that edge blurring and noise residue in flat regions still exist. Therefore, this paper proposes a dynamic fusion optimization strategy based on gradient information, which adaptively fuses the original depth map and the filtered depth map at the pixel level to achieve a dynamic balance between edge preservation and noise suppression.

[0172] The core idea of ​​this strategy is to adaptively adjust the original depth map D based on the characteristics of different regions. raw and filtered depth map D Gauss The fusion ratio. For edge regions: the original depth map is usually clearer, but filtering may cause the edges to become blurred, so more original depth information should be retained in edge regions; for flat regions: the original depth map may contain a lot of noise, but filtering will reduce the amount of noise, so more reliance should be placed on the filtering results in flat regions.

[0173] Before obtaining the filtered depth map, an adaptive mixing factor α can be defined to dynamically adjust the filtering intensity based on the gradient magnitude. The formula is as follows:

[0174] in,

[0175] G norm (x,y) represents the normalized gradient magnitude, indicating the edge strength at that pixel; α represents the adaptive blending factor, used to control the fusion weights of the original depth map and the filtered result, and to determine the smoothness of each pixel. The smaller α is, the stronger the edge region features and the more dependent it is on the original depth value; the larger α is, the stronger the flat region features and the more dependent it is on the filtered depth value; ε represents a minimal positive constant to avoid the denominator being zero and to control the gradient sensitivity, preventing extreme fusion situations caused by over-adjustment of the blending factor; n represents the shape parameter, which adjusts the nonlinearity of the blending factor as it changes with the gradient.

[0176] Finally, the fusion ratio of the original depth map and the filtered depth map can be dynamically adjusted based on the blending factor α, as shown in the formula:

[0177] D final (x,y)=α·D raw (x,y)+(1-α)·D Gauss (x,y); where,

[0178] D final This represents the final optimized pixel value.

[0179] It should be noted that this invention achieves an intelligent balance between edge preservation and noise suppression during depth map filtering by constructing a gradient-sensitive adaptive fusion mechanism. This method innovatively introduces a nonlinear fusion factor function, transforming the normalized gradient magnitude into continuously adjustable fusion weights. This allows the system to automatically adjust the contribution ratio of the original data and the filtered result based on the structural characteristics of local regions—prioritizing the retention of original depth values ​​in areas with significant edge intensity to maintain geometric sharpness, while emphasizing the filtered result in flat areas to ensure noise suppression. This gradient-driven dynamic fusion strategy, through fine-tuning of shape parameters and minimal constants, avoids the edge blurring or noise residue problems caused by traditional fixed-weight fusion, and prevents drastic fluctuations in the fusion ratio caused by gradient abrupt changes. Ultimately, the output depth map achieves optimal smoothness and consistency while maintaining the integrity of structural features, providing a depth data foundation with both geometric accuracy and noise robustness for subsequent 3D vision tasks.

[0180] Considering the complementary advantages of the two optimization methods at different stages, this paper decides to merge them into a single algorithm framework to construct a more robust and intelligent depth map completion system, while simultaneously ensuring both hole filling quality and image detail preservation. The flowchart of the jointly optimized algorithm is shown below. Figure 3 As shown.

[0181] This invention aims to optimize existing depth map completion methods based on traditional image processing by introducing structure-aware and density-response mechanisms, thereby achieving a high-precision, strongly structure-preserving, and computationally-efficient dense depth map reconstruction scheme. Its main technical innovations include:

[0182] (1) An adaptive adjustment mechanism for the expansion kernel based on local density estimation is used to improve the structural adaptability of the completion process.

[0183] This invention patent breaks through the limitations of traditional fixed structural elements and innovatively proposes a dynamic kernel scale adjustment method based on local depth density estimation. During the completion process, the number of effective pixels per unit area around the hole region is first counted, and the local depth density index (ρ value) is calculated. This index is then mapped to an appropriate expansion kernel size k, achieving dynamic control of the structural element scale. A smaller kernel is used in high-density regions to maintain fine structure; a larger kernel is used in low-density regions for rapid filling, avoiding excessive iteration and improving overall efficiency. This approach enhances the algorithm's ability to perceive spatial structure, effectively avoiding edge crossings and structural ambiguity, and significantly improving the geometric coherence and boundary consistency of depth map completion.

[0184] (2) Image region segmentation is performed using gradient information to implement differentiated filtering strategies for edge enhancement and smoothing noise reduction.

[0185] To enhance edge preservation during the completion process, this invention further introduces an image gradient sensing mechanism. Gradient calculation is performed on the preliminary completed image to extract amplitude and direction information, dividing the image into edge-sensitive regions and flat regions. Based on the region division results, different filtering strategies are applied: strong median or mean filtering is applied to flat regions to remove isolated noise; guided filtering, bilateral filtering, or lightweight median processing is used in edge regions to enhance contours and prevent boundary blurring. This technique demonstrates a responsive processing capability for local structural differences, achieving structural adaptability in the filtering function, and is one of the key features of this invention.

[0186] (3) The filtering results are intelligently weighted by integrating a dynamic weighting mechanism to retain more original effective structural information.

[0187] Considering that a single filtering strategy may lead to information loss or error propagation in complex image regions, this invention introduces a weighted fusion mechanism based on dynamic structure awareness. For multiple filtering outputs (such as a combination of median filtering and guided filtering), a weighted synthesis of local pixels is achieved by constructing a weighted coefficient map based on gradient magnitude, regional connectivity, or density changes. This weighted mechanism not only improves the fidelity of detailed regions but also effectively suppresses noise propagation and enhances the accuracy of depth map edge and texture region restoration. It is a core technical path to achieve a balance between "denoising" and "edge preservation."

[0188] (4) The expansion and filtering modules are jointly optimized within the same processing framework to improve the completion effect and system coupling.

[0189] This invention breaks away from the existing processing flow where completion and filtering are executed separately. It integrates the "density-aware dilation" and "gradient-aware filtering" modules into a single processing architecture, and designs a unified flow control logic and parameter mapping mechanism to achieve collaborative optimization between stages. By passing the density estimation results output from the dilation stage to the filtering stage, it guides region partitioning and parameter configuration, making the two stages mutually referential. This joint design not only improves processing efficiency but also enhances the system's adaptability and robustness in different scenarios, avoiding information fragmentation and error accumulation in traditional serial processing methods. It is a practically valuable technology integration solution.

[0190] (5) Provides a traditional image processing framework that does not require training, suitable for edge computing and low-resource scenarios.

[0191] Unlike deep learning methods that rely on large-scale training data and high-performance GPU inference, this invention is based entirely on traditional image processing algorithms. It requires no model training, weight storage, or parameter updates, and the entire process can run in real-time on conventional CPUs or embedded processing units, significantly reducing system deployment costs. This technical framework boasts excellent interpretability, modular customizability, and cross-platform portability, making it particularly suitable for resource-constrained embedded computing environments such as autonomous driving perception systems, robot navigation platforms, and mobile terminal 3D modeling. It represents a low-power, highly generalizable, and highly reliable depth graph completion technology solution geared towards industrial deployment.

[0192] (1) Compared with the fixed kernel method, this method has the ability to dynamically adjust the kernel size, thereby improving the flexibility of boundary response.

[0193] Traditional methods, such as IP-Basic, typically use a fixed kernel size when performing morphological dilation. This fails to adaptively adjust the kernel size based on the hole density and structural complexity of different regions in the image, leading to two main problems: over-dilation causing boundary disruption in high-density detail areas, and insufficient completion in low-density areas with large holes. This invention, however, introduces a "kernel size mapping mechanism based on local density estimation," which dynamically determines the kernel size according to the effective depth pixel ratio of each pixel's neighborhood. This enables locally sensitive, globally responsive completion, making the completion process more accurately fit the image structure and improving the flexibility of edge response and overall restoration quality.

[0194] (2) Compared with the uniform filtering strategy, region adaptive processing can significantly reduce edge blurring and error diffusion.

[0195] Many traditional filtering schemes (such as uniform Gaussian filtering and median filtering) apply uniform smoothing to the entire image, making it difficult to distinguish between edges and flat regions. This often leads to two consequences: excessive edge smoothing results in blurred structures and unclear hole boundaries; and noise is retained or even diffused in local high-frequency regions, affecting overall quality. This invention utilizes gradient information to segment image regions, constructing an "edge-flat" dual-domain filtering mechanism, and designs differentiated filtering parameters for each region, thereby processing different areas specifically. This approach can effectively enhance structural boundaries while maximizing noise suppression, achieving a high balance between filtering and edge preservation.

[0196] (3) The adaptive mechanism can automatically adjust parameters according to the input, and has good generalization ability and easy deployment.

[0197] In traditional image processing methods, parameters (such as kernel size and filter radius) typically require manual adjustment based on experience. Furthermore, fixed-parameter schemes struggle to adapt to the characteristics of depth maps at different resolutions and in different scenes (such as outdoor, indoor, low-light, and high-reflectivity environments). This invention, however, utilizes a built-in density estimation and gradient perception module to automatically perceive input data features and dynamically drive the internal processing flow and parameter selection. It eliminates the need to readjust parameters for each new scene, exhibiting excellent scene transferability and application generalization, significantly improving the algorithm's practicality and deployment efficiency.

[0198] (4) The algorithm does not require training data, has wide applicability, high running efficiency, and is suitable for running on embedded devices.

[0199] While deep learning methods have demonstrated superior accuracy in depth completion in recent years, they often face bottlenecks such as reliance on large-scale training data, long training cycles, high hardware requirements (GPU / TPU), and poor generalization after deployment. This invention, however, is entirely based on traditional image processing algorithms, requiring no external datasets or model training. It is ready to use out of the box, and its computation consists of a series of low-complexity density calculations, morphological transformations, and region filtering. It supports efficient operation on low-power processing platforms such as ARM architecture, FPGA, and DSP chips, making it particularly suitable for embedded applications such as automotive computing units, drone onboard systems, and edge terminals.

[0200] This invention patent primarily employs a statistical method based on the proportion of locally effective pixels to calculate density distribution, which drives the kernel scale adjustment logic. To improve response accuracy in regions with complex textures or deep discontinuities, the following alternative implementation methods can be considered:

[0201] (1) Sparsity determination algorithm based on gradient histogram clustering

[0202] Instead of directly relying on the number of effective pixels, this method analyzes the distribution characteristics of the gradient histogram of the depth map to perform sparse clustering modeling of different regions in the image. This method can identify potential depth faults or reconstruct anomalous regions in scenes without explicit hole markers, improving the robustness of adaptive completion.

[0203] (2) Continuous density feedback mechanism based on Gaussian kernel density estimation

[0204] A Gaussian kernel function can be used to nonparametrically model the pixel value distribution within the local neighborhood of the depth map, generating a continuous density response map. This method does not rely on discrete pixel counts in density feedback, allowing for a more detailed representation of local structural change trends, and is suitable for density modeling in higher resolution and more complex image scenes.

[0205] 2. Alternative Inflation Strategy

[0206] The original design uses a traditional morphological dilation method based on an adaptive structural kernel for depth filling. While preserving the local structural boundaries, the following strategies can be considered as alternatives or supplements:

[0207] (1) Structure-guided region growth algorithm

[0208] By utilizing neighborhood grayscale consistency or gradient consistency rules, seed points are selected from the hole boundaries for recursive region filling. Compared to traditional dilation, its region expansion has stronger directional control capabilities, making it particularly suitable for handling long, thin, and asymmetrical hole regions.

[0209] (2) Introducing a direction-sensitive expansion mechanism assisted by structural tensor

[0210] By using a structure tensor to estimate the local gradient direction and principal axis direction, the kernel shape or expansion direction is defined according to the principal axis direction when performing dilation operations, avoiding the mispropagation of structural features in the vertical direction. This is particularly suitable for detail preservation near the edges.

[0211] 3. Alternative Filtering Fusion Method Scheme

[0212] This invention patent employs a gradient region partitioning and multi-kernel joint filtering strategy to achieve a balance between edge preservation and noise reduction. To further expand application scenarios or improve detail processing capabilities, the following alternative solutions can be adopted:

[0213] (1) Guided filtering and bilateral filtering as alternative strategies

[0214] Guided filters are used to establish structural alignment between depth maps and auxiliary maps (such as gradient maps or guided RGB maps), or bilateral filters are used to combine pixel value differences and spatial distances to achieve edge-preserving smoothing. These filtering methods respond well to edge structures and are suitable for preserving strong texture details.

[0215] (2) Use edge detection results as fusion weight control mask

[0216] In the image preprocessing stage, edge detection algorithms such as Canny and Sobel can be executed to extract high-confidence structural regions and generate edge mask maps. Subsequently, in the filtering result fusion, weights are assigned to the original pixels and the filtered results based on the mask map, thereby finely controlling the preservation of edge regions and the smoothing intensity of non-edge regions.

[0217] Figure 4 A schematic diagram of an image optimization device provided for one or more embodiments of this specification, comprising:

[0218] At least one processor; and,

[0219] A memory communicatively connected to the at least one processor; wherein,

[0220] The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to:

[0221] The pixels of the depth map are inverted to obtain an inverted depth map;

[0222] Determine the proportion of effective depth pixels occupied by each preset local region in the inverted depth map;

[0223] Based on the occupancy ratio, the dilation kernel size of the depth inversion depth map is adjusted, and dilation operations are performed on each of the preset local regions based on the dilation kernel size to obtain the dilated depth map;

[0224] The expansion depth map is closed to connect the effective depth regions within the local area, fill the structural gaps, and obtain a closed depth map.

[0225] A first hole region is determined in the closed depth map, and the pixels in the first hole region are filled with a first preset size based on a preset selective dilation strategy to obtain a first filled depth map.

[0226] The first filled depth map is pushed upwards and outwards using a top extension strategy to fill in the missing area at the top of the image, thus obtaining a top extended depth map.

[0227] A second hole region is determined in the top extended depth map, and the pixels in the second hole region are filled with a second preset size to obtain a second filling depth map, wherein the second preset size is larger than the first preset size;

[0228] The second fill depth map is processed using a preset joint filtering strategy to obtain a filtered depth map;

[0229] The filtered depth map is then restored to the original depth map that matches the real scene.

[0230] This specification provides one or more embodiments of a non-volatile computer storage medium storing computer-executable instructions, which, when executed by a computer, can perform the following:

[0231] The pixels of the depth map are inverted to obtain an inverted depth map;

[0232] Determine the proportion of effective depth pixels occupied by each preset local region in the inverted depth map;

[0233] Based on the occupancy ratio, the dilation kernel size of the depth inversion depth map is adjusted, and dilation operations are performed on each of the preset local regions based on the dilation kernel size to obtain the dilated depth map;

[0234] The expansion depth map is closed to connect the effective depth regions within the local area, fill the structural gaps, and obtain a closed depth map.

[0235] A first hole region is determined in the closed depth map, and the pixels in the first hole region are filled with a first preset size based on a preset selective dilation strategy to obtain a first filled depth map.

[0236] The first filled depth map is pushed upwards and outwards using a top extension strategy to fill in the missing area at the top of the image, thus obtaining a top extended depth map.

[0237] A second hole region is determined in the top extended depth map, and the pixels in the second hole region are filled with a second preset size to obtain a second filling depth map, wherein the second preset size is larger than the first preset size;

[0238] The second fill depth map is processed using a preset joint filtering strategy to obtain a filtered depth map;

[0239] The filtered depth map is then restored to the original depth map that matches the real scene.

[0240] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and non-volatile computer storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0241] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0242] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0243] In the embodiments provided in this application, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0244] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0245] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The aforementioned units can be implemented in hardware or software.

[0246] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.

[0247] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. An image optimization method, characterized in that, The method includes: The pixels of the depth map are inverted to obtain an inverted depth map; Determine the proportion of effective depth pixels occupied by each preset local region in the inverted depth map; Based on the occupancy ratio, the dilation kernel size of the inverted depth map is adjusted, and dilation operations are performed on each of the preset local regions based on the dilation kernel size to obtain the dilated depth map; The expansion depth map is closed to connect the effective depth regions within the local area, fill the structural gaps, and obtain a closed depth map. A first hole region is determined in the closed depth map, and the pixels in the first hole region are filled with a first preset size based on a preset selective dilation strategy to obtain a first filled depth map. The first filled depth map is pushed upwards and outwards using a top extension strategy to fill in the missing area at the top of the image, thus obtaining a top extended depth map. A second hole region is determined in the top extended depth map, and the pixels in the second hole region are filled with a second preset size to obtain a second filling depth map, wherein the second preset size is larger than the first preset size; The second fill depth map is processed using a preset joint filtering strategy to obtain a filtered depth map; The filtered depth map is then restored to the original depth map that matches the real scene. Determining the occupancy ratio of effective depth pixels in each preset local region of the inverted depth map includes: The neighborhood of each hole pixel in the inverted depth map is set as a corresponding preset local region; Extract fixed-size windows within each preset local area and calculate the percentage of effective depth pixels occupied.

2. The method according to claim 1, characterized in that, Determining the proportion of effective depth pixels in each preset local region of the inverted depth map includes: Calculation function based on local density Determine the proportion of effective depth pixels occupied by each preset local region in the inverted depth map, wherein, The local density value is represented by a point. The proportion of effective depth pixels occupied within the center window; This represents the width and height of a local window, in pixels. Represents the pixel coordinates traversed within the window; This is an indicator function that determines whether the depth value of a pixel within a local window is a valid depth value. , For the preset density threshold, For point The original depth value.

3. The method according to claim 2, characterized in that, Adjusting the dilation kernel size of the inverted depth map based on the occupancy ratio includes: Input the occupancy ratio into the formula Adjust the dilation kernel size of the inverted depth map, wherein, This indicates the size of the expansion core calculated based on the stated occupancy ratio; These are local density values, ranging from... between; Indicates the preset maximum core size; This indicates the preset minimum core size.

4. The method according to claim 3, characterized in that, The method further includes: The expansion kernel size is set to satisfy an odd number constraint to ensure central symmetry, resulting in the expansion kernel size of the inverted depth map as follows: 。 5. The method according to claim 1, characterized in that, Before obtaining the filtered depth map, the method further includes: The gradient is calculated using the Scharr operator, and its kernel function formula is as follows: ;in, This represents the gradient response in the horizontal direction, used for detecting vertical edges; This represents the gradient response in the vertical direction, used for detecting horizontal edges; Input depth map; The formula for calculating gradient magnitude is: ;in, Represents pixels The gradient magnitude at a given point represents the edge strength of that pixel. Normalization is performed as follows: ;in, It is the normalized gradient magnitude, used for subsequent calculation of edge detection threshold and hybrid optimization factor.

6. The method according to claim 5, characterized in that, Before obtaining the filtered depth map, the method further includes: Median filtering is used to preprocess the global pixel to remove salt-and-pepper noise, so that the statistical median of neighboring pixels can be used to replace the current pixel value. The formula is as follows: ;in, Indicates the input depth map in coordinates Pixel value at that location, To preset window size, This is the depth map output after median filtering; Represents the pixel coordinates traversed within the window; By setting a threshold To distinguish edge areas and flat areas And different strengths of Gaussian filtering are applied to different regions, especially for edge regions. Sharpening is applied to compensate for the blur caused by filtering and enhance edge details. The formula is as follows: ;in, This represents the pixel value after sharpening. This is the sharpening factor, used to control the edge sharpening intensity. The standard deviation is Gaussian kernel; For flat areas Strong smoothing is employed to suppress residual noise while avoiding the introduction of artifacts. The formula is as follows: ;in, This represents the pixel value after smoothing. The standard deviation of the Gaussian filter in the flat region must be ensured. To ensure a smoother finish.

7. The method according to claim 6, characterized in that, Before obtaining the filtered depth map, the method further includes: Define adaptive mixing factor The filter strength is dynamically adjusted based on the gradient magnitude, using the following formula: ;in, This represents the normalized gradient magnitude and the edge strength at that pixel. This represents the adaptive blending factor, used to control the fusion weights of the original depth map and the filtered result, and to determine the smoothness of each pixel. A smaller value indicates stronger edge region features and a greater dependence on the original depth value. A larger value indicates a stronger flat region feature and a greater dependence on the filtered depth value. This represents extremely small positive numbers, avoids zero denominators, controls gradient sensitivity, and prevents extreme fusion cases caused by over-adjustment of the mixing factor; This represents the shape parameter, which adjusts the degree of nonlinearity of the mixing factor as it changes with the gradient; Based on mixing factors The formula for dynamically adjusting the fusion ratio of the original depth map and the filtered depth map is as follows: ;in, This represents the final optimized pixel value.

8. An image optimization device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to: The pixels of the depth map are inverted to obtain an inverted depth map; Determine the proportion of effective depth pixels occupied by each preset local region in the inverted depth map; Based on the occupancy ratio, the dilation kernel size of the inverted depth map is adjusted, and dilation operations are performed on each of the preset local regions based on the dilation kernel size to obtain the dilated depth map; The expansion depth map is closed to connect the effective depth regions within the local area, fill the structural gaps, and obtain a closed depth map. A first hole region is determined in the closed depth map, and the pixels in the first hole region are filled with a first preset size based on a preset selective dilation strategy to obtain a first filled depth map. The first filled depth map is pushed upwards and outwards using a top extension strategy to fill in the missing area at the top of the image, thus obtaining a top extended depth map. A second hole region is determined in the top extended depth map, and the pixels in the second hole region are filled with a second preset size to obtain a second filling depth map, wherein the second preset size is larger than the first preset size; The second fill depth map is processed using a preset joint filtering strategy to obtain a filtered depth map; The filtered depth map is then restored to the original depth map that matches the real scene. Determining the occupancy ratio of effective depth pixels in each preset local region of the inverted depth map includes: The neighborhood of each hole pixel in the inverted depth map is set as a corresponding preset local region; Extract fixed-size windows within each preset local area and calculate the percentage of effective depth pixels occupied.

9. A non-volatile computer storage medium, characterized in that, It stores computer-executable instructions, which, when executed by a computer, can achieve the following: The pixels of the depth map are inverted to obtain an inverted depth map; Determine the proportion of effective depth pixels occupied by each preset local region in the inverted depth map; Based on the occupancy ratio, the dilation kernel size of the inverted depth map is adjusted, and dilation operations are performed on each of the preset local regions based on the dilation kernel size to obtain the dilated depth map; The expansion depth map is closed to connect the effective depth regions within the local area, fill the structural gaps, and obtain a closed depth map. A first hole region is determined in the closed depth map, and the pixels in the first hole region are filled with a first preset size based on a preset selective dilation strategy to obtain a first filled depth map. The first filled depth map is pushed upwards and outwards using a top extension strategy to fill in the missing area at the top of the image, thus obtaining a top extended depth map. A second hole region is determined in the top extended depth map, and the pixels in the second hole region are filled with a second preset size to obtain a second filling depth map, wherein the second preset size is larger than the first preset size; The second fill depth map is processed using a preset joint filtering strategy to obtain a filtered depth map; The filtered depth map is then restored to the original depth map that matches the real scene. Determining the occupancy ratio of effective depth pixels in each preset local region of the inverted depth map includes: The neighborhood of each hole pixel in the inverted depth map is set as a corresponding preset local region; Extract fixed-size windows within each preset local area and calculate the percentage of effective depth pixels occupied.

Citation Information

Patent Citations

  • Real-time depth completion method based on pseudo depth map guidance

    CN112861729A

  • Sparse radar depth completion method based on Boolean mask constraint

    CN119850433A