A deep depth map super-resolution method and system based on entropy map guidance
By using entropy map guidance and wavelet decomposition techniques, the guidance intensity of the color image is adaptively adjusted, which solves the problems of texture migration and structural inconsistency in depth map super-resolution reconstruction and generates high-quality depth maps.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ROCKET FORCE UNIV OF ENG
- Filing Date
- 2026-03-18
- Publication Date
- 2026-05-29
AI Technical Summary
The lack of an adaptive guidance mechanism based on the complexity of local image information in existing technologies leads to texture migration errors and inconsistencies in global structure during depth map super-resolution reconstruction.
An entropy-map-guided depth map super-resolution method is adopted. The entropy-map-guided module quantifies the complexity of local information, and the wavelet decomposition and hierarchical orientation perception module are combined to perform adaptive feature fusion. A composite loss function is used for training to collaboratively optimize the global structure and local details.
It achieves adaptive and precise guidance, avoids texture migration errors, improves the overall quality and realism of the depth map, and ensures global structural consistency and accuracy of local detail reconstruction.
Smart Images

Figure CN122115219A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of computer vision and deep learning technology, and in particular to a depth map super-resolution reconstruction technique guided by high-resolution color images. Background Technology
[0002] Depth maps can represent the three-dimensional geometric information of a scene and have important application value in fields such as autonomous driving and 3D reconstruction. However, due to limitations in the hardware performance of acquisition devices and complex imaging environments, directly acquired depth maps often suffer from low resolution, lack of detail, and noise contamination, which seriously affects the accuracy of subsequent advanced vision tasks. Therefore, using high-resolution color images that are strongly correlated with depth information and are easily accessible to guide the super-resolution reconstruction of depth maps has become the mainstream technical approach in this field.
[0003] In existing technologies, some methods employ deep learning networks to fuse features from color and depth images. For example, some methods construct pyramid structures to form multi-scale feature representations and use feedback fusion strategies to combine information from different levels, thereby gradually improving the resolution and quality of the depth map.
[0004] However, these methods still have some shortcomings. First, their fusion strategies rely on a relatively implicit and singular basis for determining where and how heavily to use color images for guidance, failing to adaptively adjust based on the local complexity of the image content. This can easily lead to insufficient guidance in texture-rich areas or the introduction of unnecessary texture artifacts in smooth areas, resulting in texture transfer problems. Second, existing methods often focus more on restoring high-frequency details during the fusion process, potentially neglecting low-frequency information that plays a decisive role in the overall structure of the depth map. This can lead to inconsistencies or distortions in the global structure of the final restored depth map. Summary of the Invention
[0005] The purpose of this application is to provide a depth map super-resolution method and system based on entropy map guidance, which aims to solve the technical problems in the prior art, such as the lack of an adaptive guidance mechanism based on the complexity of local image information, and the failure to effectively coordinate the processing of global structural information and local detail information of the depth map, resulting in structural inconsistencies or texture migration errors in the recovered depth map.
[0006] To achieve the above objectives, this application provides a depth map super-resolution method based on entropy map guidance. The method is implemented using a deep neural network and includes: acquiring a low-resolution depth map, a corresponding high-resolution color map, and a ground-value high-resolution depth map and a ground-value entropy map for training; generating a global entropy guidance map based on the low-resolution depth map and the high-resolution color map using an entropy map guidance module; extracting depth features from the low-resolution depth map and color map features from the high-resolution color map, respectively; and performing the following operations on the depth features and color map features in one or more hierarchical orientation sensing modules: using wavelet decomposition to decompose the depth features and color map features into low-frequency components and high-frequency components, respectively; fusing the low-frequency components and the high-frequency components based on the global entropy guidance map to generate a fused depth feature; and generating a high-resolution depth map based on the fused depth feature; wherein the deep neural network is trained using a composite loss function, the composite loss function including a first loss term for constraining the difference between the high-resolution depth map and the ground-value high-resolution depth map, and a second loss term for constraining the difference between the global entropy guidance map and the ground-value entropy map.
[0007] Optionally, the hierarchical orientation sensing module is called iteratively multiple times; wherein the fused depth features generated in the previous iteration are used as the input depth features of the hierarchical orientation sensing module in the next iteration.
[0008] Optionally, the depth features and the color map features are extracted using a residual network structure.
[0009] Optionally, the step of fusing low-frequency components and high-frequency components includes: generating a first attention weight using the low-frequency components of the color image feature and applying it to the low-frequency components of the depth feature; and generating a second attention weight using the high-frequency components of the color image feature and / or the high-frequency components of the depth feature to selectively fuse the high-frequency components of the color image feature into the high-frequency components of the depth feature.
[0010] Optionally, the step of generating a high-resolution depth map includes: generating a residual map based on the fused depth features, and adding the residual map to an image obtained by performing bicubic interpolation on the low-resolution depth map.
[0011] Optionally, the step of fusing high-frequency and low-frequency components includes: concatenating the low-frequency components of the depth features with the low-frequency components of the color image features, and concatenating the high-frequency components of the depth features with the high-frequency components of the color image features, and then performing information fusion through a convolutional layer.
[0012] Optionally, the deep neural network employs an encoder-decoder architecture, which includes an encoder, wherein the hierarchical orientation sensing module is invoked at different scales of the encoder.
[0013] Optionally, the first loss term is the spatial domain L1 loss, which can be expressed as: Furthermore, the second loss term is the entropy domain L1 loss, which can be expressed as: The composite loss function is a weighted sum of the two losses mentioned above, and its form can be expressed as: ,in These are the preset weight hyperparameters.
[0014] Another aspect of this application provides a depth map super-resolution system, comprising: a memory storing computer program instructions; and a processor; wherein the processor is configured to implement the aforementioned entropy map-guided depth map super-resolution method when executing the computer program instructions.
[0015] Optionally, in the system, when the processor is configured to execute the computer program instructions, it implements the scheme in the aforementioned method where the hierarchical orientation perception module is iteratively invoked multiple times.
[0016] Compared with existing technologies, the technical solution provided in this application has the following beneficial effects: First, this application achieves adaptive and precise guidance. By introducing an entropy map guidance module, this application can explicitly quantify the local information complexity of the image and dynamically adjust the guidance intensity of the color image accordingly. This allows the network to fully utilize the texture information of the color image in areas rich in detail, while avoiding the introduction of irrelevant texture artifacts in smooth areas, effectively solving the problem of texture transfer errors. Second, this application synergistically optimizes global structure and local details. By employing wavelet decomposition in the hierarchical orientation perception module, features are decoupled into low-frequency components representing global structure and high-frequency components representing local details, and differentiated processing is performed. This allows the network to simultaneously focus on the global structural consistency of the depth map and the reconstruction accuracy of local details, solving the problem of unbalanced high- and low-frequency information recovery in existing technologies. Finally, this application ensures the structural authenticity of the reconstruction results. By employing a composite loss function that includes spatial domain loss and entropy domain loss for end-to-end training, this application not only constrains numerical accuracy at the pixel level, but also constrains the generated results to align with the structure of the real depth map at the level of information complexity distribution, thereby comprehensively improving the overall quality and realism of the super-resolution depth map. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A flowchart illustrating a depth map super-resolution method guided by entropy graphs, provided in an embodiment of this application;
[0019] Figure 2 This is a schematic diagram of the overall network architecture provided in Embodiment 1 of this application;
[0020] Figure 3 A schematic diagram of the internal structure of the layered orientation sensing module (HDA) provided in an embodiment of this application;
[0021] Figure 4 This is a schematic diagram of the structure of a depth map super-resolution system provided in an embodiment of this application.
[0022] The main reference numerals in the accompanying drawings are explained as follows: 100: System device; 110: Processor; 120: Memory; 121: Instruction; 122: Data; 130: Input interface; 140: Output interface; S101: Step of receiving low-resolution depth map and high-resolution color map; S102: Step of generating global entropy guide map; S103: Step of extracting and performing wavelet decomposition to obtain high and low frequency feature components; S104: Step of adaptively fusing high and low frequency components according to global entropy guide map; S105: Step of generating residual map through inverse wavelet transform and subsequent network; S106: Step of adding residual map with interpolated depth map to obtain final result; Low-resolution depth map; Super-resolution depth map; DAM: Direction Aware Module; EGM: Entropy Map Guidance Module; Global entropy guidance graph; : Deep features; : Residual characteristics after fusion; Color image features; HDA: Hierarchical orientation sensing module; HFDF: High-frequency feature fusion module; High-resolution color image; LFDF: Low-frequency feature fusion module; Wavelet Decomposition; Wavelet Reconstruction. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0024] Example 1
[0025] This embodiment details an entropy-map-guided depth map super-resolution method and the deep neural network upon which it relies. This method aims to address the problems in existing technologies where the lack of adaptive guidance for the complexity of local image information and the inability to effectively coordinate the processing of global structure and local details leads to structural distortion or texture transfer errors in the reconstructed depth map.
[0026] Reference Figure 1 This is a schematic diagram illustrating the overall process of a depth map super-resolution method guided by entropy maps, provided in one embodiment of this application. In this embodiment, the method is mainly implemented through a deep neural network, and its specific process includes the following steps:
[0027] Step S101: Receive a low-resolution depth map and a high-resolution color image. Specifically, in one application scenario, the system acquires a pair of spatially aligned image data to be processed via an input interface, namely, a low-resolution depth map D_lr and a high-resolution color image. It is understandable that the low-resolution depth map This can be derived from low-cost depth sensors or obtained through downsampling simulation of high-resolution depth maps, and it typically suffers from low resolution, blurred edges, and lack of detail. Correspondingly, the high-resolution color image... These images are typically captured by a standard camera, thus containing rich texture and structural information. It should be noted that during the network training phase, ground-value high-resolution depth maps also need to be acquired for supervised learning. And the truth entropy map calculated based on this truth depth map. .
[0028] Step S102: Generate a global entropy guidance map. This involves using the low-resolution depth map obtained in step S101. With high-resolution color images They are all fed into an entropy map guidance module (EGM). The core function of this module is to quantify the information complexity of the input image in different local regions. Specifically, the entropy map guidance module (EGM) can first calculate... and The local information entropy is used to generate respective entropy maps. The calculation of information entropy can be based on Shannon entropy theory. One implementation is to define a neighborhood window (such as 3x3 or 5x5) centered on each pixel in the image, then statistically analyze the probability distribution of pixel values within this window, and calculate the entropy value accordingly. Typically, high entropy regions correspond to areas with drastic content changes and rich details (such as object edges and complex textures), while low entropy regions correspond to areas with smooth content and simple structure (such as walls and floors). Subsequently, the Entropy Map Guidance Module (EGM) fuses these two entropy maps from different modalities. Fusion methods include, but are not limited to, pixel-by-pixel addition, convolution after concatenation, or the use of more complex network structures, ultimately generating a single global entropy guidance map that comprehensively reflects the structural complexity of the scene. The global entropy guidance graph This serves as a global and explicit guiding signal for the subsequent feature fusion process.
[0029] Step S103: Extract features and perform wavelet decomposition to obtain high- and low-frequency feature components. In parallel with step S102, the system processes the low-resolution depth map... With high-resolution color images The data are fed into two independent initial feature extraction networks. As an optional implementation, this feature extraction network can employ a residual network structure to facilitate training deeper networks and effectively extracting deep semantic features. Through these feature extraction networks, initial deep features are obtained. Features of color images Next, within the Hierarchical Direction Awareness (HDA) module, wavelet decomposition is performed on these two feature maps. It is understood that wavelet decomposition is an effective signal processing technique capable of decomposing a signal or image into different frequency sub-bands. In this embodiment, a two-dimensional discrete wavelet transform is used to decompose the depth features... Decompose it into a low-frequency component and multiple high-frequency components (e.g., details in the horizontal, vertical, and diagonal directions); similarly, color image features are also decomposed. It is decomposed into one low-frequency component and multiple high-frequency components. The low-frequency component mainly captures global structural information such as the overall contour and smooth areas of the image, while the high-frequency components mainly contain local detail information such as the edges and textures of the image.
[0030] Step S104: Adaptive fusion of high and low frequency components based on the global entropy guidance map. This step is also completed within the hierarchical orientation sensing module (HDA). This module employs differentiated fusion strategies for the low-frequency and high-frequency components obtained from step S103. Specifically, the fusion process is influenced by the global entropy guidance map generated in step S102. Modulation. For example, the global entropy guide graph can be used. This can be used as an additional input channel, or transformed into attention weights, applied to various stages of feature fusion. This mechanism enables the network to... The network adaptively adjusts the guidance strength from color image features based on the indicated local region complexity. In regions with high information complexity (such as object boundaries), the network enhances the utilization of high-frequency details in the color image; while in regions with low information complexity (such as flat surfaces), it suppresses the introduction of color image information to avoid unnecessary texture transfer. In this way, adaptive and precise fusion of depth features and color image features across different frequency bands is achieved, ultimately generating fused depth features.
[0031] Step S105: Generate a residual map through wavelet reconstruction and subsequent networks. The low-frequency and high-frequency components obtained through adaptive fusion in step S104 are recombined into a unified fused feature map through wavelet reconstruction. This feature map can then be processed by further convolutional layers or upsampling networks to adjust the number of channels and spatial resolution, ultimately generating a predicted residual map.
[0032] Step S106: Add the residual map to the interpolated depth map to obtain the final result. The system first processes the original low-resolution depth map. Upsampling is performed to match the target high resolution. A preferred implementation is bicubic interpolation. Then, the upsampled depth map is pixel-by-pixel added to the residual map generated in step S105. This residual learning-based strategy allows the network to focus solely on learning the subtle differences between the high-resolution depth map and the low-resolution interpolation result, thus reducing the learning difficulty and improving reconstruction accuracy. The result of the addition is the final output super-resolution depth map. .
[0033] For a more detailed explanation of the technical solution in this embodiment, please refer to [link / reference]. Figure 2 This diagram illustrates the overall architecture of the deep neural network (EHDA-Net) used in this embodiment. The network receives low-resolution depth maps. and high-resolution color images As input.
[0034] The network first processes the input through an entropy graph guidance module (EGM) to generate a global entropy guidance graph, which is crucial for adaptive guidance. .at the same time, Initial deep features are generated through a deep feature extraction branch consisting of multiple residual sets. ; Then, a color image feature extraction branch consisting of residual blocks is used to generate color image features. This method of feature extraction using residual network structures can build deeper network models to extract richer feature representations, while effectively avoiding the gradient vanishing problem.
[0035] The network backbone employs an iterative structure, consisting of three cascaded hierarchical orientation sensing (HDA) modules. In the first iteration, the first HDA module receives initial depth features. and color image features As input, and in the global entropy guiding graph Under the guidance of the first-level fusion, the layered fusion is performed, and the residual features after the first-level fusion are output. Subsequently, this residual characteristic Through a skip connection to the deep features of the input. Element-wise addition is performed to obtain the updated depth features. These updated depth features will then be compared with the original color image features. Together, these features are used as input to the second HDA module. This process is repeated, meaning that the fused depth features generated in the previous iteration are used as input depth features to the hierarchical orientation perception module in the next iteration, thus achieving progressive refinement and optimization of the depth map features. After three iterations, the last HDA module outputs the final fused residual features.
[0036] The final fused residual features are fed into a dense projection upsampling network, which upscales the feature map to the target size and generates the final residual map. This residual map is then compared with the feature map after bicubic interpolation. Adding them together yields the final super-resolution depth map. .
[0037] Please see now Figure 3 This is a schematic diagram of the internal structure of the Hierarchical Direction Awareness (HDA) module, revealing the specific implementation mechanism of adaptive fusion. The HDA module receives the depth features of the current iteration. and color image features .
[0038] First, through wavelet decomposition units, and They are decomposed into low-frequency components ( , ) and high-frequency components ( , ).
[0039] These components are then fed into two parallel sub-modules for processing: the low-frequency feature fusion module LFDF and the high-frequency feature fusion module HFDF.
[0040] In the Low-Frequency Feature Fusion (LFDF) module, the primary objective is to utilize the global structure of the color image to correct and enhance the global structure of the depth map. Specifically, the low-frequency components of the color image... The data is fed into a direction-aware module (DAM), which typically consists of convolutional layers, activation functions (such as ReLU), and a final softmax or sigmoid layer to generate an initial attention weight map. This attention weight map is then applied to the low-frequency components of the depth map via element-wise multiplication. This allows for the weighting and adjustment of the structural information in the depth map, generating fused low-frequency features.
[0041] In the High Frequency Feature Fusion (HFDF) module, the goal is to selectively borrow useful edge and detail information from the color image. Specifically, the high frequency components of the depth map... and high-frequency components of color images First, the data is concatenated along the channel dimension. Then, the concatenated features are fed into another orientation-aware module (DAM) to generate a second attention weight. This weight is also used to selectively enhance or suppress high-frequency components of the color image. The information in the text. Weighted. Subsequently, it was compared with the high-frequency components of the original depth map. By combining them, high-frequency details are effectively integrated.
[0042] The fused low-frequency and high-frequency features are then merged using wavelet reconstruction units. The merged feature map is compared with the global entropy guided map from the EGM. They are fed together into a reversible network for final feature adjustment and output. (Global entropy guidance graph) This module plays a crucial role here, performing final modulation on the fused features to ensure the fusion process adheres to the guidelines of global information complexity. The HDA module ultimately outputs the fused residual features from this iteration. .
[0043] During the network training phase, this embodiment employs a composite loss function for end-to-end optimization. This composite loss function consists of two weighted components: the first loss term is a spatial domain loss, preferably an L1 loss, used to constrain the output super-resolution depth map. Compared with true high-resolution depth maps The pixel-level differences between them. Its mathematical form is:
[0044]
[0045] Where N is the batch size and i represents the sample index. This loss term ensures the numerical accuracy of the reconstruction results.
[0046] The second loss term is the entropy domain loss, preferably L1 loss, used to constrain the global entropy guidance graph generated by the network. According to the truth depth map Calculated truth entropy map The difference between them. Its mathematical form is:
[0047]
[0048] This loss term serves as an auxiliary supervision signal, forcing the entropy graph guidance module (EGM) to learn and generate a guidance graph that accurately reflects the structural complexity of the real scene, thereby indirectly improving the performance of the main task and ensuring the structural authenticity of the reconstruction results.
[0049] The total loss function is the weighted sum of the two losses mentioned above:
[0050]
[0051] in, It is a preset weight hyperparameter used to balance the two optimization objectives of spatial domain accuracy and entropy domain structural consistency.
[0052] Through the above structure and training strategy, the method in this embodiment can generate high-quality super-resolution depth maps, which clearly show details such as object edges and fine structures, while maintaining the continuity and distortion-free nature of the overall three-dimensional structure of the scene. This is significantly better than traditional interpolation methods and some deep learning methods that fail to provide adaptive guidance.
[0053] Example 2
[0054] This embodiment provides a variation of the scheme described in Embodiment 1. This embodiment aims to illustrate that the core idea of this application, namely, entropy-map-guided frequency band fusion, is not limited to the complex fusion module based on the attention mechanism used in Embodiment 1, but can also be implemented in a simpler way, thereby reducing computational complexity while maintaining good performance.
[0055] The overall network architecture of this embodiment and Figure 2 The architecture shown is basically the same, also including the entropy map guidance module (EGM), the iterative hierarchical orientation sensing module (HDA), and the final reconstruction part. The main difference lies in the specific implementation of the low-frequency feature fusion module (LFDF) and the high-frequency feature fusion module (HFDF) within the hierarchical orientation sensing module (HDA).
[0056] In this embodiment, the fusion steps for high-frequency and low-frequency components have been simplified. Specifically, the Direction Awareness Module (DAM) is no longer used to generate explicit attention weights.
[0057] In the Low-Frequency Feature Fusion (LFDF) module, the low-frequency components of the depth features and the low-frequency components of the color image features are directly concatenated along the channel dimension. The concatenated feature tensor is then input into one or more convolutional layers (e.g., a 1x1 convolutional layer followed by a 3x3 convolutional layer). These convolutional layers are responsible for learning how to automatically fuse the low-frequency information from the two modalities and restoring the number of channels to a level consistent with the input depth features through dimensionality reduction.
[0058] Similarly, in the high-frequency feature fusion module HFDF, the high-frequency components of the depth features and the high-frequency components of the color image features are also concatenated along the channel dimension, and then information is fused through another set of independent convolutional layers.
[0059] In this simplified scheme, the global entropy guiding graph Its guiding role can still be retained. For example, the global entropy guiding graph can be used. (Size matching and channel duplication may be required) The concatenated features, along with the features to be fused, are concatenated again along the channel dimension and then fed into the fusion convolutional layer. In this way, when learning the fusion weights, the convolutional kernel can use the entropy information of the local region as one of the judgment criteria, thereby indirectly realizing entropy-guided adaptive fusion.
[0060] The working process of this embodiment is similar to that of Embodiment 1: The low-resolution depth map... and high-resolution color images The input is fed into the network, where iterative feature fusion is performed through an HDA module containing a simplified fusion strategy, and finally outputs a super-resolution depth map. .
[0061] In terms of expected results, this embodiment can achieve super-resolution reconstruction of depth maps with lower computational resources and faster inference speed. Although its ability to suppress texture artifacts may be slightly inferior to the refined fusion strategy based on explicit attention in Embodiment 1 when dealing with extremely complex texture regions, its overall performance is still significantly better than the baseline method without frequency division processing or adaptive guidance. This demonstrates the universality and efficiency of the two core ideas proposed in this application, namely "entropy guidance" and "frequency division fusion," which can be combined with fusion operators of different complexities to achieve different balances between performance and efficiency.
[0062] Example 3
[0063] This embodiment provides another variation of the scheme described in Embodiment 1, which aims to illustrate that the core functional modules of this application, namely the Entropy Graph Guidance Module (EGM) and the Hierarchical Direction Awareness Module (HDA), have good modularity and scalability, and can be flexibly embedded into different types of network backbone architectures, rather than being limited to the iterative structure in Embodiment 1.
[0064] In this embodiment, the deep neural network employs a non-iterative, U-Net-like encoder-decoder architecture. This architecture has proven effective in many image-to-image conversion tasks.
[0065] The specific structural flow is as follows: 1. Encoder section: The encoder consists of a series of downsampling stages. In each stage, the resolution of the feature map decreases while the number of channels increases. At the initial input of the network, the entropy map guidance module (EGM) also operates based on the input low-resolution depth map. and high-resolution color images Generate a global entropy bootstrap graph At each scale of the encoder (i.e., after each downsampling stage), a hierarchical orientation awareness (HDA) module is deployed. This HDA module receives depth and color map features at that scale (these two feature streams pass through the encoder in parallel) and performs adaptive fusion by frequency band. The entropy guidance map used to guide the HDA module at this scale is derived from the original global entropy guidance map. This was achieved through a corresponding number of downsampling operations. Therefore, at each level of the encoding, entropy-guided multimodal feature fusion was implemented.
[0066] 2. Decoder Section: The decoder is structurally symmetrical to the encoder and consists of a series of upsampling stages to progressively restore the resolution of the feature maps. In each upsampling stage, the decoder's features are concatenated with the encoder's features at the corresponding scale, after being fused by the HDA module. This skip-connection approach transfers the fused features, rich in spatial information, from the encoder to the decoder, helping the decoder to better recover image details.
[0067] 3. Output section: The decoder ultimately outputs a feature map with the same resolution as the target map. This feature map can be directly used as the final super-resolution depth map. Or as a residual map and the interpolated Adding them together gives .
[0068] The working process of this embodiment is as follows: after inputting the image pair into the network, it undergoes one forward propagation to complete entropy-guided layered fusion at each level of the encoding-decoding process, ultimately generating a high-resolution depth map. The training process also uses the composite loss function defined in Embodiment 1.
[0069] The intended effect of this embodiment is to demonstrate that the EGM and HDA modules can be seamlessly integrated into mainstream encoder-decoder architectures such as U-Net as plug-and-play functional units, effectively improving their performance in depth map super-resolution tasks. This shows that the technical concept proposed in this application has broad applicability, does not depend on a specific network backbone design, and provides those skilled in the art with the possibility of applying this concept to other related network architectures.
[0070] Example 4
[0071] This embodiment provides another variation of the scheme described in Embodiment 1, focusing primarily on the loss function used during the training phase. This embodiment aims to illustrate that the core idea of the proposed "spatial domain + entropy domain" dual-supervision framework is to simultaneously constrain the numerical accuracy and structural complexity of the output results, rather than being limited to using a specific form of loss function (such as L1 loss).
[0072] In this embodiment, the overall structure and operation of the network are exactly the same as in Embodiment 1. The only difference lies in the composite loss function used during the training phase. Its specific composition.
[0073] As an optional implementation, spatial domain loss can be used. Replace it with L2 loss. L2 loss penalizes the square of the error, so it is more sensitive to larger error values than L1 loss, tending to produce results with smaller overall error but potentially slightly blurry edges.
[0074] As an alternative implementation method, spatial domain loss Alternatively, a loss function that focuses more on image structural similarity can be used, such as the structural similarity index loss or its multi-scale version. The structural similarity index measures image similarity from three aspects: brightness, contrast, and structure, and its loss function is typically in the form of… Using structural similarity index loss for optimization helps the network generate results that are visually closer to the ground truth image, especially in preserving the structure of objects.
[0075] Similarly, for entropy domain loss Alternatively, L2 loss or structural similarity index loss can be used to measure the predicted global entropy guided graph. With the truth entropy diagram The differences between them.
[0076] For example, a new composite loss function can be defined as a weighted sum of the spatial domain structural similarity exponential loss and the entropy domain L2 loss:
[0077]
[0078] in This represents the square of the L2 norm.
[0079] The working process of this embodiment is to train the EHDA-Net described in Embodiment 1 using the modified composite loss function described above. During inference (i.e., actual use), since the network structure remains unchanged, its working process is exactly the same as in Embodiment 1.
[0080] The intended effect of this embodiment is that by changing the loss function, the network can be guided to optimize different objectives. For example, using L2 loss can train a model that is insensitive to outliers, while using structural similarity index loss can train a model that focuses more on maintaining the integrity of the image structure. This demonstrates that the dual-supervision framework proposed in this application has high flexibility, and can select or combine the most suitable loss function form according to specific application requirements and different emphases on the reconstruction results (e.g., whether to pursue the highest peak signal-to-noise ratio or the best visual effect), thereby achieving optimal training results.
[0081] Example 5
[0082] This application also provides a depth map super-resolution system capable of performing the methods described in any of the foregoing embodiments. Please refer to... Figure 4 This is a schematic diagram of the system device.
[0083] The system device 100 can be a general-purpose computing device, such as a personal computer, server, or workstation, or a dedicated embedded system, such as a domain controller for an autonomous vehicle or a robot control unit. The system device 100 mainly includes a processor 110, a memory 120, an input interface 130, and an output interface 140. These components can communicate with each other via a bus or other forms of connection mechanism.
[0084] The processor 110 is the computing core of the system device 100, and may be one or more central processing units, graphics processing units, digital signal processors, or specially designed neural network processing units, etc. The processor 110 is responsible for executing instructions stored in the memory 120.
[0085] The memory 120 can be any type of volatile or non-volatile storage medium, such as random access memory, read-only memory, hard disk drive, solid-state drive, etc. The memory 120 stores computer program instructions 121 executable by the processor 110, which constitute a software program implementing the method of this application. Furthermore, the memory 120 is also used to store various data 122 required during processing, such as low-resolution depth maps received through the input interface 130. and high-resolution color images The ground truth data required for network training and This includes various intermediate feature maps generated by the network model during the computation process and the final output results.
[0086] The input interface 130 is used to acquire data from the external world. For example, it can be an interface for connecting to a depth camera and a color camera, or it can be an interface for reading image files from a storage device.
[0087] Output interface 140 is used to present the processing results to the user or send them to other system modules. For example, it can be connected to a display to show the generated super-resolution depth map. Alternatively, the result data can be transmitted to subsequent 3D reconstruction or path planning modules.
[0088] When system device 100 is operational, processor 110 is configured to execute computer program instructions 121 stored in memory 120. When these instructions are executed, processor 110 controls the entire system device 100 to implement the depth map super-resolution method as described in any one of Embodiments 1 to 4. For example, processor 110 loads a neural network model, also stored as data 122 in memory 120, receives image data through input interface 130, performs calculations according to the network-defined process (such as entropy map generation, iterative layered fusion, residual reconstruction, etc.), and finally outputs a high-quality super-resolution depth map through output interface 140. In this way, system device 100 of this embodiment constitutes a physical implementation of the method of this application.
[0089] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A depth map super-resolution method guided by entropy graphs, characterized in that, The method is implemented using a deep neural network and includes: Obtain a low-resolution depth map, a corresponding high-resolution color map, and a ground-value high-resolution depth map and a ground-value entropy map for training; An entropy-guided module generates a global entropy-guided map based on the low-resolution depth map and the high-resolution color map. The depth features of the low-resolution depth map and the color features of the high-resolution color map are extracted respectively. In one or more hierarchical orientation sensing modules, the following operations are performed on the depth features and the color image features: using wavelet decomposition, the depth features and the color image features are decomposed into low-frequency components and high-frequency components, respectively; according to the global entropy guidance map, the low-frequency components and the high-frequency components are fused to generate a fused depth feature; Based on the fused depth features, a high-resolution depth map is generated; The deep neural network is trained using a composite loss function, which includes a first loss term for constraining the difference between the high-resolution depth map and the ground truth high-resolution depth map, and a second loss term for constraining the difference between the global entropy guidance map and the ground truth entropy map.
2. The method according to claim 1, characterized in that, The hierarchical orientation sensing module was called iteratively multiple times; The fused depth features generated in the previous iteration are used as the input depth features for the hierarchical orientation perception module in the next iteration.
3. The method according to claim 1 or claim 2, characterized in that, The depth features and the color map features are extracted using a residual network structure.
4. The method according to claim 1 or 2, characterized in that, The step of fusing low-frequency components and high-frequency components includes: A first attention weight is generated using the low-frequency components of the color image features and applied to the low-frequency components of the depth features; and A second attention weight is generated using the high-frequency components of the color image features and / or the high-frequency components of the depth features to selectively fuse the high-frequency components of the color image features into the high-frequency components of the depth features.
5. The method according to claim 1 or 2, characterized in that, The steps for generating a high-resolution depth map include: A residual map is generated based on the fused depth features, and the residual map is added to the image obtained by bicubic interpolation and magnification of the low-resolution depth map.
6. The method according to claim 1, characterized in that, The step of fusing high-frequency and low-frequency components includes: concatenating the low-frequency components of the depth features with the low-frequency components of the color image features, and concatenating the high-frequency components of the depth features with the high-frequency components of the color image features, and then performing information fusion through a convolutional layer.
7. The method according to claim 1, characterized in that, The deep neural network adopts an encoder-decoder architecture, which includes an encoder, wherein the hierarchical orientation perception module is invoked at different scales of the encoder.
8. The method according to claim 1, characterized in that, The first loss term is the spatial domain L1 loss, and the second loss term is the entropy domain L1 loss.
9. A depth map super-resolution system, comprising: A memory that stores computer program instructions; One processor; Wherein, the processor is configured to execute the computer program instructions to implement the method as described in claim 1.
10. The system according to claim 9, characterized in that, When the processor is configured to execute the computer program instructions, it implements the method as described in claim 2.