Image reconstruction method and device, electronic equipment and storage medium

By optimizing the low-frequency and high-frequency components of the depth map using dual-tree complex wavelet transform and cross-modal attention model, and combining finite element analysis and lightweight optical flow prediction, the semantic guidance and physical rule compatibility issues of the depth map in the digital twin scenario are solved, and the generation and real-time processing of high-quality depth maps are realized.

CN121120446APending Publication Date: 2025-12-12ZHEJIANG DAHUA SYST ENG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511190167.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing deep repair methods suffer from insufficient semantic guidance capabilities, poor physical rule compatibility, limited dynamic adaptability, and computational efficiency bottlenecks in digital twin scenarios, making it difficult to meet the requirements for high-quality depth maps.

Method used

The low-frequency and high-frequency components of the depth map are decomposed using dual-tree complex wavelet transform. Combined with the rigid body deformation rule library and multi-scale feature pyramid generated by finite element analysis, the semantic enhancement and physical constraint optimization of the depth map are achieved through cross-modal attention model and lightweight optical flow prediction. Real-time processing is performed through an end-to-end acceleration framework.

Benefits of technology

It improves the quality of depth maps, enhances semantic guidance capabilities and physical rule compatibility, improves adaptability to dynamic scenes and computational efficiency, and meets the real-time requirements of digital twin scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121120446A_ABST
    Figure CN121120446A_ABST
Patent Text Reader

Abstract

Embodiments of the invention disclose an image reconstruction method and apparatus, an electronic device and a storage medium. The method comprises the steps of determining a target depth map and a target RGB map of a to-be-processed image; the target depth map is a depth map after space alignment, and the target RGB map is an enhanced RGB map; processing the target depth map to obtain a low-frequency component, a high-frequency component, an edge feature and a feature pyramid; based on the target RGB image, the low-frequency component, the edge features and the feature pyramid, obtaining semantic enhanced depth features; applying the target RGB image, the high-frequency component, the feature pyramid and the semantic enhanced depth features to generate a repaired depth image meeting physical consistency; and processing the repaired depth map by using a multi-feature pyramid to generate a target depth map meeting the time sequence consistency. By applying the method provided by the embodiment of the invention, a high-quality depth map can be obtained, so that the requirements of a digital twinning scene can be met.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, and particularly relates to an image reconstruction method and device, electronic equipment and a storage medium. BACKGROUND

[0002] Under the background of rapid development of digital twinning technology, as the core data carrier for accurate mapping between the physical world and the virtual space, the quality of the depth map directly affects the reliability of key scenes such as consistency control, dynamic simulation modeling and real-time interaction.

[0003] The depth map obtained by the depth repair method in the related art has insufficient semantic guidance ability, poor physical rule compatibility and limited dynamic adaptability, and it is difficult to meet the requirements of the digital twinning scene.

[0004] Therefore, there is an urgent need for an image reconstruction method to obtain a high-quality depth map to meet the requirements of the digital twinning scene. SUMMARY

[0005] The embodiments of the present application provide an image reconstruction method, device, electronic equipment and storage medium, which can obtain a high-quality depth map, and thus can meet the requirements of the digital twinning scene.

[0006] In a first aspect, an embodiment of the present application provides an image reconstruction method, comprising:

[0007] determining a target depth map and a target RGB map of a to-be-processed image; wherein the target depth map is a spatially aligned depth map, and the target RGB map is an enhanced RGB map;

[0008] processing the target depth map to obtain a low-frequency component, a high-frequency component, an edge feature and a feature pyramid;

[0009] obtaining a semantically enhanced depth feature based on the target RGB map, the low-frequency component, the edge feature and the feature pyramid;

[0010] applying the target RGB map, the high-frequency component, the feature pyramid and the semantically enhanced depth feature to generate a repaired depth map meeting physical consistency;

[0011] applying a multi-feature pyramid to process the repaired depth map to generate a target depth map meeting temporal consistency.

[0012] In this embodiment, firstly, a target depth map and a target RGB image of the image to be processed are determined. The target depth map is a spatially aligned depth map, and the target RGB image is an enhanced RGB image. Secondly, the target depth map is processed to obtain low-frequency components, high-frequency components, edge features, and a feature pyramid. Thirdly, based on the target RGB image, low-frequency components, edge features, and feature pyramid, semantically enhanced depth features are obtained. Then, the target RGB image, high-frequency components, feature pyramid, and semantically enhanced depth features are applied to generate a repaired depth map that satisfies physical consistency. Finally, a multi-feature pyramid is applied to process the repaired depth map to generate a target depth map that satisfies temporal consistency. The target depth map obtained by applying this image reconstruction method considers low-frequency components, high-frequency components, and edge features, and constructs a feature pyramid and considers semantically enhanced depth features, resulting in a higher quality target depth map that can better support scenarios such as digital twins.

[0013] Secondly, one embodiment of this application provides an image reconstruction apparatus, comprising:

[0014] The determining unit is used to: determine the target depth map and the target RGB map of the image to be processed; wherein the target depth map is a spatially aligned depth map, and the target RGB map is an enhanced RGB map;

[0015] The data processing unit is used to process the target depth map to obtain low-frequency components, high-frequency components, edge features, and feature pyramids.

[0016] The data processing unit is also used to: obtain semantically enhanced deep features based on the target RGB image, low-frequency components, edge features, and feature pyramid;

[0017] The data processing unit is also used to: apply the target RGB image, high-frequency components, feature pyramid and semantically enhanced deep features to generate a repair depth map that satisfies physical consistency;

[0018] The data processing unit is also used to: apply a multi-feature pyramid to process the repair depth map and generate a target depth map that meets temporal consistency requirements.

[0019] Thirdly, one embodiment of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of any of the above methods.

[0020] Fourthly, one embodiment of this application provides a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, implement the steps of any of the above methods.

[0021] Fifthly, one embodiment of this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above methods. Attached Figure Description

[0022] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments of this application will be briefly introduced below. Obviously, the drawings introduced below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 A schematic flowchart of an image reconstruction method provided in an embodiment of this application;

[0024] Figure 2 This is a schematic diagram of the structure of an image reconstruction apparatus provided in one embodiment of this application;

[0025] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0027] For ease of understanding, the terms used in the embodiments of this application are explained below:

[0028] (1) Digital twin is a technological concept and application paradigm that uses digital technology to construct a precise mirror image of physical entities, systems or processes in virtual space, and realizes two-way data interaction, analysis and optimization between the virtual and real worlds. Through real-time or near-real-time data connection, it enables the virtual model to dynamically reflect the state, behavior and performance of physical entities, thereby supporting a series of applications such as monitoring, simulation, prediction and optimization.

[0029] The number of any elements in the accompanying drawings is for illustrative purposes only and not as a limitation, and any naming is for distinction only and has no limiting meaning.

[0030] Against the backdrop of the rapid development of digital twin technology, depth maps, as the core data carrier for the precise mapping between the physical world and virtual space, directly impact the reliability of key scenarios such as virtual-real consistency control, dynamic simulation modeling, and real-time interaction. Depth maps quantify the spatial geometric relationships of objects in a 3D scene (such as distance, surface curvature, and topology), providing fundamental support for physics engines' rigid body motion simulation (such as collision detection), material reflection modeling (such as ray tracing), and dynamic environment perception (such as urban traffic flow simulation). They play a crucial role, especially in smart city 3D modeling and refined operation and maintenance of urban infrastructure. For example, the 3D reconstruction of large-scale urban scenes requires sub-millimeter-level precision depth information representation for building facades (such as glass curtain wall joints), underground pipe network topology (such as pipe diameter connection accuracy), and road networks (such as the geometric relationship between manhole covers and curbs). Meanwhile, urban infrastructure maintenance (such as bridge crack detection and quantification of weathering in ancient buildings) relies on high-fidelity reconstruction of microstructures using depth maps.

[0031] However, in practical applications, due to limitations in sensor hardware characteristics and environmental complexity (such as occlusion by moving objects, interference from highly reflective surfaces, and scattering noise from multiple light sources), depth maps generally suffer from three major defects: data incompleteness (missing rate >30%, forming discontinuous holes), geometric distortion (blurred object edges leading to a significant decrease in the structural similarity index, and unrecognizable sub-millimeter-level surface textures), and physical inconsistency (the repair results violate rigid body dynamics constraints, resulting in ghosting penetration phenomena). Traditional depth restoration methods (such as U-Net-based encoding and decoding architectures and GAN-driven generative adversarial networks) can improve peak signal-to-noise ratio, but they are limited by low-frequency-dominated restoration mechanisms (significant loss of edge sharpness), fragmented multimodal features (texture-geometric mismatch due to RGB and depth map semantic misalignment), and computational efficiency bottlenecks. These limitations make it difficult to meet the stringent requirements of digital twin scenarios, such as high-frequency detail preservation and physical rule compatibility (compliance with rigid body collision detection and material reflection laws).

[0032] Therefore, although related technologies have improved basic indicators through techniques such as frequency domain decomposition and multimodal fusion, insufficient semantic guidance (microstructure positioning deviation), poor compatibility of physical rules (rigid body constraint conflicts), and limited dynamic adaptability (sudden occlusion response delay) remain the core bottlenecks restricting the repair of digital twin depth maps. By enhancing the semantic relevance of high-frequency detail reconstruction, optimizing the real-time nature of physical law embedding, and improving the robustness of dynamic scenes, the geometric accuracy and system reliability of virtual-real mapping can be significantly improved, providing technical support for high-fidelity modeling and precise operation and maintenance of smart cities.

[0033] Based on the aforementioned technical deficiencies, this application needs to solve the following core technical problems:

[0034] 1) Insufficient semantic guidance: Related technologies rely on color-depth mapping assumptions or simple texture alignment, lacking high-level semantic guidance, which leads to a significant decrease in repair accuracy in complex scenes (such as reflective surfaces and vegetation occlusion).

[0035] 2) Loss of high-frequency details: Traditional interpolation and filtering methods (such as bilateral filtering and Gaussian filtering) make sub-millimeter-level structures (such as building joints and road markings) unrecognizable.

[0036] 3) Physical rule conflict: The repair results did not verify the rigid body motion law and material reflection characteristics, resulting in ghost image penetration or surface curvature distortion.

[0037] 4) Poor adaptability to dynamic scenes: Static repair strategies cannot handle transient holes (such as occlusion by moving objects), and the response latency is >100ms, which limits real-time interactive applications.

[0038] 5) Computational efficiency bottleneck: Related technologies (such as bilateral filtering) have high computational complexity, making it difficult to meet the real-time requirements (latency <50ms) at 4K resolution.

[0039] Therefore, this application provides an image reconstruction method that integrates deep semantic enhancement and frequency domain decomposition, which can obtain high-quality depth maps and provide support for digital twin scenarios.

[0040] The main protection points of this application's embodiments are as follows:

[0041] 1. Frequency Domain Decomposition and Physical Constraint Fusion Mechanism Based on Dual-Tree Complex Wavelet Transform

[0042] Specifically, during the neural network training process, the low-frequency and high-frequency components of the depth map are separated by dual-tree complex wavelet transform (DTCWT), the rigid body deformation rule library generated by finite element analysis (FEA) is transformed into a differentiable loss function, and the constraint weights are dynamically adjusted based on the semantic mask of the multi-scale feature pyramid. The joint optimization of physical rules and high-frequency details is achieved through backpropagation.

[0043] Among them are coupling optimization methods for low-frequency components and physical constraints (such as second-derivative smoothing constraints and volume conservation constraints), dynamic selection mechanisms for overcomplete dictionaries in high-frequency component sparse coding (Gabor / curved wave / DCT basis functions), and gradient backpropagation fusion methods for physical rule bases and deep learning networks.

[0044] 2. Cross-modal attention-driven dynamic semantic enhancement architecture

[0045] A cross-modal attention model is constructed with low-frequency geometric features as queries and RGB semantic features as key-value pairs. By combining the global layout information of the multi-scale feature pyramid (FP), adaptive alignment of semantic and geometric features is achieved.

[0046] This involves a lightweight multi-scale feature extraction method based on the MobileNetV3-Small backbone network, a differential processing mechanism for rigid / flexible objects (building second derivative constraint vs. vegetation curvature threshold exemption), and a joint optimization strategy for attention weights and interpretable semantic segmentation networks.

[0047] 3. Lightweight optical flow prediction and multi-scale spatiotemporal fusion system

[0048] Hardware acceleration of optical flow prediction is achieved based on a cropped version of the RAFT network. A spatiotemporal fusion strategy with dynamic weight adjustment is designed by combining motion region detection with a multi-scale feature pyramid (FP_1 / 4).

[0049] This includes methods for compressing optical flow network parameters (depth-separable convolution + INT8 quantization), a high-frequency component energy comparison and fusion mechanism based on motion masks, and methods for forward distortion and reverse consistency verification of historical frame depth maps.

[0050] 4. End-to-end acceleration framework for edge computing

[0051] By employing techniques such as memory pooling management, TensorRT layer fusion, and mixed-precision training, we achieve end-to-end real-time processing at 4K resolution in 48ms.

[0052] This includes memory reuse mechanisms for frequency domain decomposition modules and neural network computations, cross-layer computation graph optimization strategies for multi-scale feature pyramids (such as Winograd algorithm optimization), and collaborative acceleration methods for dynamic channel pruning and hardware instruction sets (such as CUDA kernel functions).

[0053] 5. Data augmentation and adversarial training mechanism based on Poisson sampling

[0054] A non-continuous hole mask is generated by Poisson disk sampling, and a dynamic adversarial training dataset is constructed by combining channel-independent multiplicative illumination perturbation.

[0055] This includes methods for coupling hole generation strategies with sensor noise models, a linkage optimization mechanism between illumination perturbation parameters and frequency domain decomposition thresholds, and a semantic prior guidance strategy based on Canny edge detection.

[0056] 6. Joint optimization of multimodal sparse representation and channel attention

[0057] The Orthogonal Matching Pursuit (OMP) algorithm is used to sparsely encode high-frequency components, and the basis function weights are dynamically adjusted by combining the channel attention mechanism.

[0058] Among them are adaptive setting methods for energy threshold filtering and residual iteration termination conditions, joint optimization algorithms for channel attention weights and sparse coefficient matrices, curve wave basis enhancement strategies for vegetation regions and DCT basis suppression mechanisms for building regions.

[0059] To further illustrate the technical solutions provided in the embodiments of this application, a detailed description is provided below in conjunction with the accompanying drawings and specific implementation methods. Although the embodiments of this application provide method operation steps as shown in the following embodiments or drawings, the method may include more or fewer operation steps based on conventional or non-inventive methods. In steps where there is no logically necessary causal relationship, the execution order of these steps is not limited to the execution order provided in the embodiments of this application.

[0060] The technical solutions provided in the embodiments of this application will be described below.

[0061] refer to Figure 1 This application provides an image reconstruction method, including the following steps:

[0062] S101: Determine the target depth map and target RGB map of the image to be processed.

[0063] Among them, the target depth map is the spatially aligned depth map, and the target RGB map is the enhanced RGB map.

[0064] S102: Process the target depth map to obtain low-frequency components, high-frequency components, edge features, and feature pyramids.

[0065] S103: Based on the target RGB image, low-frequency components, edge features, and feature pyramid, semantically enhanced deep features are obtained.

[0066] S104: Apply the target RGB image, high-frequency components, feature pyramid, and semantically enhanced deep features to generate a repair depth map that satisfies physical consistency.

[0067] S105: Apply a multi-feature pyramid to process the repair depth map and generate a target depth map that satisfies temporal consistency.

[0068] In this embodiment, firstly, a target depth map and a target RGB image of the image to be processed are determined. The target depth map is a spatially aligned depth map, and the target RGB image is an enhanced RGB image. Secondly, the target depth map is processed to obtain low-frequency components, high-frequency components, edge features, and a feature pyramid. Thirdly, based on the target RGB image, low-frequency components, edge features, and feature pyramid, semantically enhanced depth features are obtained. Then, the target RGB image, high-frequency components, feature pyramid, and semantically enhanced depth features are applied to generate a repaired depth map that satisfies physical consistency. Finally, a multi-feature pyramid is applied to process the repaired depth map to generate a target depth map that satisfies temporal consistency. The target depth map obtained by applying this image reconstruction method considers low-frequency components, high-frequency components, and edge features, and constructs a feature pyramid and considers semantically enhanced depth features, resulting in a higher quality target depth map that can better support scenarios such as digital twins.

[0069] S101 mainly involves the multimodal data alignment and preprocessing process. This process involves acquiring the image to be processed, which can be a relevant image from an urban construction scene. The target depth map and target RGB image of the image to be processed can be determined, and this process can be achieved through steps A1-A2:

[0070] A1: Obtain the original depth map and original RGB image of the image to be processed.

[0071] The process of processing the image to obtain the original depth map and the original RGB image can be found in the methods in related technologies, and will not be elaborated here.

[0072] A2: Perform spatial alignment on the original depth map to obtain a spatially aligned depth map, and enhance the original RGB image to obtain an enhanced RGB image.

[0073] Optionally, the process of obtaining a spatially aligned depth map can be achieved through steps A2-1 to A2-2:

[0074] A2-1: Construct a perspective transformation matrix based on the calibration parameters of the camera used to capture the image to be processed.

[0075] Optionally, a perspective transformation matrix can be constructed based on the camera's calibration parameters (including intrinsic parameter matrices describing the camera's internal characteristics and extrinsic parameter matrices describing the camera's position and orientation). The construction process can be found in relevant technologies and will not be elaborated here.

[0076] A2-2: Apply the perspective transformation matrix to perform a function alignment operation on the original depth map to obtain a spatially aligned depth map.

[0077] Then, the original depth map D_in is aligned using the warpPerspective function through a GPU kernel function accelerated by TensorRT. The perspective transformation matrix is ​​used as the key parameter input of the warpPerspective function, and the INTER_NEAREST interpolation method is selected to operate on the depth map D_in in the warpPerspective function to preserve the discrete characteristics of the original depth values. Finally, a depth-aligned depth map D_aligned is generated, which can also be called a spatially aligned depth map.

[0078] Optionally, the process of obtaining the enhanced RGB image can be achieved through steps A2-3 to A2-5:

[0079] A2-3: Add random noise to each color channel in the original RGB image.

[0080] Optionally, a 20% multiplicative random noise operation can be applied independently to each color channel (R, G, B) of the original RGB image to simulate lighting changes in a real scene.

[0081] A2-4: Perform histogram stretching on an RGB image with added random noise.

[0082] Specifically, histogram stretching is performed on the RGB image after noise is added to enhance image contrast.

[0083] A2-5: Perform a normalization operation on the stretched RGB image to obtain the enhanced RGB image.

[0084] For example, the stretched RGB image is normalized to adjust its pixel value range to the range of [-0.5, 0.5], and the enhanced RGB image I_enhanced is output.

[0085] It should be noted that there is no clear order between these two processes; this is just an example and does not constitute a specific limitation.

[0086] The aforementioned preprocessing process optimizes GPU memory resource operations through memory pool management technology, and simultaneously optimizes computation graph node operations using TensorRT layer fusion technology (merging Conv+BN+ReLU operations). These optimizations reduce preprocessing latency from 50ms in traditional schemes to 10ms, achieving efficient alignment and enhancement of depth RGB data at 4K resolution, providing a geometrically consistent and lighting-robust input foundation for subsequent real-time repair.

[0087] S102 mainly involves the frequency domain decomposition and multi-scale extraction process, which is primarily achieved through steps B1-B4:

[0088] B1: Perform dual-tree complex wavelet transform decomposition on the target depth map to obtain multiple sub-bands, and process the original low-frequency sub-bands in the multiple sub-bands to obtain low-frequency components.

[0089] First, a dual-tree complex wavelet transform (DTCWT) decomposition operation is performed on the spatially aligned depth map D_aligned. This operation decomposes the depth map into 7 subbands: one low-frequency approximation subband (i.e., the original low-frequency subband, which can be represented by LL) and 6 high-frequency directional subbands. The low-frequency subband LL contains the low-frequency components of the image (such as macroscopic structures like building outlines), while the 6 high-frequency directional subbands include positive and negative phase components in three directions: +15° and -15° (i.e., 165°), +45° and -45° (135°), and +75° and -75° (105°).

[0090] Secondly, the original low-frequency sub-bands in multiple sub-bands are processed to obtain low-frequency components. This process can be achieved through steps B1-1 to B1-3.

[0091] B1-1: Perform artifact suppression operation on the original low-frequency subbands in multiple subbands.

[0092] This operation utilizes the translation invariance of complex wavelets to suppress aliasing artifacts (mainly manifested as jagged edge artifacts) caused by downsampling through a phase compensation mechanism. The specific steps are as follows:

[0093] Step 1: Phase Compensation. Adjust the phase of each frequency component in the low-frequency sub-band to keep it consistent at different positions, thereby eliminating the phase discontinuity caused by downsampling.

[0094] Step 2: Smoothing. The low-frequency sub-band after phase adjustment is smoothed to further suppress aliasing artifacts.

[0095] B1-2: Perform energy threshold filtering on the original low-frequency subband after artifact suppression.

[0096] Optionally, each element is squared and summed, then sorted in descending order of energy value. Based on this, the major components with a cumulative energy percentage greater than 95% are retained, while the remaining low-energy components (less than 5%), which are typically noise, are removed.

[0097] B1-3: Perform convolution compression on the original low-frequency subband after energy filtering to obtain the low-frequency components.

[0098] Specifically, a 1x1 convolution compression operation is performed on the energy-filtered LL subband (a single-channel 2D matrix). First, the LL subband is reshaped into a 3D tensor format, then expanded to 4 channels using a zero-padding mode. Next, a 1x1 convolution kernel is used to compress the channel dimension, reducing the number of channels by a quarter while maintaining the spatial resolution. The output is a low-frequency component D_low representing the macroscopic geometric structure. This component serves as the stable basis for subsequent physical constraint optimization, preserving the large-scale structural information such as building outlines and road plans.

[0099] B2: Select the target subband that meets the conditions from multiple subbands, and process the target subband to obtain high-frequency components.

[0100] Optionally, channel pruning is performed on the six high-frequency directional subbands obtained by the dual-tree complex wavelet transform (DTCWT) decomposition (each subband corresponds to a specific direction: 15°, 45°, 75°, 105°, 135°, 165°, and each subband is an independent channel).

[0101] The target subband is defined as the subband with energy exceeding a set energy threshold among multiple subbands. The process for determining the target subband is as follows: calculate the energy distribution of each subband channel. For example, by summing the squares of all spatial location elements in the subband matrix, and then calculating the proportion of each subband channel's energy to the total high-frequency energy, remove redundant subband channels with an energy proportion of less than 5% (e.g., the 165° direction subband in the sky region usually has the lowest energy), while retaining high-energy subband channels (e.g., the 45° direction subband at the edge of buildings has the strongest response).

[0102] B2-1: Perform directional fusion operation on the target subband to obtain the fused multi-directional enhanced signal.

[0103] Perform directional fusion operation on the high-frequency subband (target subband) channels retained after pruning. For example, assign weights according to the energy ratio of each subband channel (give higher weights to high-energy subbands), and fuse multiple directional subbands into a single multidirectional enhancement signal H_fused by weighted summation.

[0104] B2-2: Perform orthogonal matching pursuit operation on the fused multi-directional enhanced signal to obtain sparse coding coefficients.

[0105] Optionally, an orthogonal matching pursuit operation is performed on the fused multi-directional enhanced signal H_fused. Specifically, the fused multi-directional enhanced signal H_fused is sparsely encoded using an overcomplete dictionary Φ (containing Gabor basis functions to capture orientation-sensitive edges, curvelet basis functions to characterize anisotropic structures, and DCT basis functions to describe global smoothing patterns) to obtain sparse encoded coefficients. This encoding process is iterated 20 times, each time selecting the basis function atom in the dictionary Φ that is most relevant to the current residual and updating the sparse coefficient set until the residual energy drops to below 10% of the initial signal energy of H_fused, effectively separating real edge components from noise components.

[0106] B2-3: Perform channel attention mechanism operation on sparse coding coefficients to obtain semantically optimized sparse coding coefficients.

[0107] The process involves performing channel attention on the sparse coding coefficients output after orthogonal matching and tracking. Specifically, a lightweight channel attention module parses the semantic information of the low-frequency component D_low (such as building outlines and road plans) to generate a spatial weight map. This map labels building areas with high weight values ​​(>0.9) and sky areas with low weight values ​​(<0.2). This weight map is then multiplied positionally by the sparse coding coefficients, enhancing coefficients at key locations such as building joints (weight >0.9) and suppressing coefficients in the sky area (weight <0.2), outputting semantically optimized sparse coding coefficients. These semantically optimized sparse coding coefficients have already been enhanced (in important regions) and suppressed (in unimportant regions) based on low-frequency semantic information.

[0108] B2-4: Reconstruct the sparse coding coefficients of semantic optimization to obtain high-frequency components.

[0109] For example, a reconstruction operation is performed on the weighted sparse coding coefficients. Specifically, the high-frequency signal is reconstructed using an overcomplete dictionary Φ to generate the final high-frequency component D_high. This component can completely preserve sub-millimeter-level details such as tile joints and leaf veins (spatial resolution up to 0.1mm), while sensor noise is efficiently suppressed (shot noise removal rate >85%), providing high-quality high-frequency input for subsequent depth map restoration.

[0110] B3: Input the low-frequency and high-frequency components into the lightweight dual-branch network to generate optimized low-frequency detail features and optimized high-frequency detail features; optimize the optimized low-frequency detail features and optimized high-frequency detail features to obtain edge features.

[0111] Optionally, the lightweight dual-branch network can be the MobileNetV3-Small backbone network, with the low-frequency component D_low and the high-frequency component D_high respectively input to the lightweight dual-branch network.

[0112] Specifically, the low-frequency component D_low (building outline / road plane) is spatially compressed by 8-stride convolution, reducing the output size to 1 / 8 of the original. Then, max pooling is used to enhance the features of the convolution result, extracting the most significant response of each 2×2 region, and finally generating a 1 / 8 scale global geometric feature F_low. This feature represents the distribution of kilometer-level building clusters and road topology at a resolution of 64×64 (when the input is 512×512).

[0113] Multi-scale detail extraction of high-frequency components D_high (tile joints / leaf veins) is performed by depthwise separable dilated convolution (dilation rate 2 / 4 / 8). The extracted high-frequency features are then subjected to sensor noise suppression processing (shot noise removal rate >85%) through nonlocal mean filtering to generate optimized high-frequency detail features F_high.

[0114] For feature optimization, firstly, semantic parsing is performed on F_low using a channel attention module (SE Block) to generate a channel weight vector W_channel_low (with the same dimension as the number of channels in F_low). Based on W_channel_low, channel weights are applied to F_low, dynamically adjusting the contribution of each channel. Then, semantic parsing is performed on F_high using the same channel attention module (SE Block) to generate a channel weight vector W_channel_high (with the same dimension as the number of channels in F_high). Based on W_channel_high, channel weights are applied to F_high, dynamically adjusting the contribution of each channel. Finally, nonlocal mean filtering is used to sharpen the edges of the weighted F_high, extracting salient edge features E_map with SSIM > 0.92.

[0115] B4: Combine edge features and optimized low-frequency detail features, and apply geometric constraints to the combination results to obtain geometrically optimized features; input the geometrically optimized features into the backbone network to construct the feature pyramid.

[0116] Optionally, E_map and F_low can be fused by feature concatenation and input into a lightweight CRF module for geometric constraint processing, outputting a geometrically optimized feature FO_geo (containing macroscopic structures of smooth wall curvature and microscopic details of edge alignment).

[0117] The geometrically optimized feature FO_geo is input into the improved MobileNetV3-Small backbone network to construct a four-scale feature pyramid:

[0118] 1 / 8 scale generation: Spatial compression is performed on FO_geo through 8-stride convolution (the convolution kernel moves 8 pixels each time), downsampling the 512×512 input to 64×64; then, feature enhancement processing is performed on the convolution result through 2×2 max pooling (extracting the maximum value of each 2×2 region), outputting FP_1 / 8 features (bottom layer features, also known as fourth layer features). This 64×64 matrix represents the distribution of kilometer-level building clusters, just like a satellite cloud image showing the urban skeleton.

[0119] 1 / 4 scale generation: The FO_geo is processed by medium-precision scanning through a depthwise separable convolution with a stride of 4 (the convolution kernel moves 4 pixels each time, and the number of channels is 128), compressing the 512×512 input to 128×128, and outputting FP_1 / 4 features (also known as third-layer features). This 128×128 matrix locates medium-sized objects (such as vehicles) with a boundary error of <2 pixels.

[0120] 1 / 2 scale generation: FO_geo is processed by component-level parsing through stride 2 dilated convolution (the convolution kernel moves 2 pixels each time, with a dilation rate of 2). The receptive field is expanded by interval sampling, reducing the dimensionality of the 512×512 input to 256×256, and outputting FP_1 / 2 features (also known as second-layer features). This 256×256 matrix captures architectural defects (such as cracks in car windows) with a precision of 5mm.

[0121] Original image scale generation: FO_geo is processed to preserve details by 8-dilation convolution (convolution kernel elements are sampled at 8-pixel intervals) to maintain a 512×512 resolution; then, the high-frequency component D_high is processed to enhance the edges by pixel-level multiplication, and FP_full features (also known as the first layer features) are output. This 512×512 matrix achieves sub-millimeter detail reconstruction with an accuracy of 0.1mm (such as tile joints).

[0122] This application addresses the problem that traditional methods rely on simple color-depth mapping and are prone to semantic misalignment and structural distortion in complex scenes. It constructs a cross-modal attention model, deeply integrates RGB semantics and depth geometric features, and uses a semantic segmentation network to achieve differentiated processing of scene categories. This significantly improves the repair accuracy of complex scenes (such as reflective surfaces and vegetation occlusion), greatly enhances cross-modal feature alignment capability, and significantly optimizes geometric-texture consistency.

[0123] Furthermore, addressing the issues of local structural distortion caused by single-scale features and high computational resource consumption of traditional methods, a multi-scale feature pyramid combined with dilated convolution covers everything from global layout to local details; lightweight network architecture and hardware acceleration technology are deeply integrated. This significantly enhances multi-granularity scene reconstruction capabilities and greatly improves detail integrity; computational resource consumption is significantly reduced, and high-resolution processing efficiency is optimized through breakthroughs.

[0124] Regarding S103, the main focus is on the process of deep semantic enhancement and cross-modal fusion. Optionally, the feature pyramid is a multi-scale feature pyramid, which can be implemented as follows:

[0125] C1: Based on low-frequency components and bottom-level features in the feature pyramid, determine the weighted query vector; perform multi-level feature extraction on the target RGB image based on the set network to obtain semantic features at each layer; process at least one layer of features in the feature pyramid through feature concatenation and convolution to obtain the medium-sized object location key and detail enhancement value; perform noise suppression processing on each layer of features in the feature pyramid to obtain the initial deep features for semantic enhancement.

[0126] Specifically, the low-frequency component D_low is compressed using 1x1 convolution to generate the initial query vector Q; Q and FP_1 / 8 are fused by channel-wise multiplication to enhance the building area perception capability based on the global layout information of FP_1 / 8, and the weighted Q_weighted is output to guide the attention mechanism to focus on key areas.

[0127] Multi-level feature extraction processing of the RGB image I_enhanced is performed using the ResNet-18 network to obtain semantic features from layer 1 to layer 4. Noise suppression processing is applied to the features of the four layers FP_full, FP_1 / 2, FP_1 / 4, and FP_1 / 8 using learnable channel attention, and a unified output FP_fused is output for subsequent attention calculation.

[0128] Optionally, the process of determining the positioning keys and detail enhancement values ​​for medium-sized objects can be implemented through steps C1-1 to C1-2:

[0129] C1-1: By combining features and performing 3*3 convolution, key vector generation is performed on the third layer features in the feature pyramid and the third layer features of the set network to obtain the location key of the medium-sized object.

[0130] C1-2: By combining features and performing 1*1 convolution, value vector generation is performed on the first layer features in the feature pyramid and the first layer features of the network to obtain detail enhancement values.

[0131] Specifically, key vector generation is performed on the FP_1 / 4 and ResNet layer3 features through feature concatenation and 3x3 convolution to output the medium-sized object localization key K; value vector generation is performed on the FP_full and ResNet layer1 features through feature concatenation and 1x1 convolution to output the detail enhancement value V.

[0132] C2: Perform cross-modal attention calculation on the weighted query vector, medium-sized object location key, and detail enhancement value to obtain the first fusion feature. Then, perform geometric fidelity processing on the first fusion feature and low-frequency components to obtain the intermediate semantic enhancement feature.

[0133] Optionally, cross-modal attention calculation is performed on Q_weighted, K, and V through Scaled Dot-Product Attention to generate fused feature A; geometric fidelity processing is performed on A and the original depth feature D_low by adding the residuals to preserve the geometric structure of the original depth map, while introducing RGB semantic information to output intermediate semantic enhancement feature F_attention.

[0134] C3: Perform edge enhancement processing on the target layer features and edge images of the feature pyramid to obtain the second fusion feature; calculate the wall integrity mask through building category damage, and calculate the leaf contour mask through vegetation edge loss;

[0135] Optionally, edge enhancement processing is performed on FP_1 / 4 and E_map through feature concatenation, and edge information extracted from RGB images is injected into deep features to generate a fused feature F_edge_enhanced, which is used to construct the input of the semantic segmentation model architecture DeepLabv3+. Global supervision processing is performed on the deep output of DeepLabv3+ through building category loss calculation to optimize the building structure in the high-level semantic space and generate a wall integrity mask M_wall. Local supervision processing is performed on the shallow output of DeepLabv3+ through vegetation edge loss calculation to optimize the vegetation edge details in the low-level feature space and generate a leaf outline mask M_leaf.

[0136] C4: Perform high-dimensional fusion processing on the initial semantically enhanced deep features, intermediate semantically enhanced features, wall integrity mask, leaf outline mask, and initial semantically enhanced deep features to obtain semantically enhanced deep features.

[0137] Optionally, the attention fusion feature F_attention, semantic mask M_wall / M_leaf, and multi-scale FP feature FP_fused are concatenated along the channel dimension to perform high-dimensional fusion processing, generating a unified representation F_high_dim; F_high_dim is then compressed through 1x1 convolution to remove redundant information, outputting F_reduced; F_reduced is then enhanced with key information through global semantic weighting, adjusting the contribution of each channel based on the semantic weights provided by FP_1 / 8, outputting the final semantically enhanced feature F_fused, which serves as the core input for subsequent physical constraint optimization and high-frequency reconstruction.

[0138] S104 mainly involves the process of physical constraint optimization and high-frequency reconstruction, which can optionally be implemented through the following steps:

[0139] D1: The semantically enhanced deep features and features from each layer of the feature pyramid are concatenated along the channel dimension using a lightweight decoder, and the concatenated features are gradually upsampled to the original image resolution through multiple deconvolution layers to obtain a preliminary repaired depth map.

[0140] Specifically, a lightweight decoder concatenates the features FP_1 / 8, FP_1 / 4, and FP_full of F_fused and FP at all levels along the channel dimension to achieve cross-level information fusion, which is used for subsequent upsampling to generate high-resolution output. Multiple deconvolutional layers are used to progressively upsample the concatenated features to the original image resolution (512×512) to restore the spatial dimension and map the low-resolution features back to the original image size. The initial repaired depth map D_init is output for subsequent physical constraint and high-frequency detail enhancement processing.

[0141] D2: Perform rigid region mask extraction on the preliminary repair depth map to obtain the intermediate repair depth map.

[0142] D2-1: Perform rigid region mask extraction on the preliminary repair depth map to determine the target region that needs to be subject to rigid constraints, constrain the parameters of the target region, and perform differential modeling on the target region to obtain rigid body dynamic constraints.

[0143] Specifically, the motion rule base of typical rigid bodies (such as cubes and cylinders) is modeled using finite element analysis (FEA) to pre-generate their deformation patterns under collision and gravity, and physical constraints such as curvature continuity and volume conservation are defined to guide the model in learning a depth distribution that conforms to real-world laws. Based on this pre-generated rule base, the building-type region mask M_wall output from the semantically guided repair stage and the global features FP_1 / 8 of the multi-scale feature pyramid are jointly analyzed to extract object regions with rigid behavior (such as walls and roofs), providing structural priors for subsequent geometric consistency modeling. Then, the extracted rigid regions are subjected to geometric continuity enforcement through second-derivative smoothing constraints to suppress wall... The model identifies phenomena such as body fracture and uneven roof surfaces, outputting the rigid structural feature F_rigid as the core input for physical consistency modeling. Combining the structural priors provided by F_rigid, the model extracts rigid region masks from the preliminary repair depth map D_init, identifying regions R requiring rigid constraints to prevent incorrect modeling of flexible regions such as vegetation. Furthermore, the model constrains the spatial second derivative of region R using an L1 norm loss term, enhancing the geometric continuity of building-type regions. Finally, integrating the aforementioned pre-generated finite element rule library, the model differentiates between different rigid regions through a dynamic weight adjustment mechanism: a planar deformation mode is used to enhance surface smoothness in building regions, while a shell deformation mode is applied to vehicle components to enhance local curvature continuity.

[0144] D2-2: Perform material property analysis on the target RGB image and construct a reference depth that conforms to the laws of optical physics based on the analyzed material properties; segment the high reflectivity region through the third layer of features in the feature pyramid to generate a mask for the high reflectivity surface; for the region marked by the mask, minimize the error between the preliminary repair depth map and the theoretical reflection depth to obtain the optical path tracing constraint.

[0145] Specifically, the enhanced RGB image I_enhanced output from the preprocessing stage is processed by material property analysis (this image comes from the multimodal data alignment and preprocessing output of the first stage) to obtain the specular reflection coefficient k_s and roughness σ. Based on the analyzed material properties, the theoretical reflection depth D_theory under the virtual light source is generated by the ray casting algorithm to construct a reference depth that conforms to the laws of optical physics. High reflectivity regions are segmented by the mesoscale feature FP_1 / 4 in the feature pyramid to generate a mask M that identifies highly reflectivity surfaces such as glass and metal. For the regions identified by the mask M, the error between the preliminary repair depth D_init output from the feature decoding stage and the theoretical reflection depth D_theory generated by the ray casting algorithm is minimized by the L2 norm loss, forcing the reflection angle error to be controlled within <0.5°.

[0146] D2-3: The network parameters of the integrated physical loss function of rigid body dynamics constraints and optical path tracing constraints are optimized by directional propagation, and the intermediate repair depth map is determined based on the processed parameters.

[0147] Optionally, the network parameters can be optimized by backpropagation to improve the combined physical loss of rigid body dynamics constraints and optical path tracing constraints, so that the model output converges in a physically reasonable direction; the intermediate optimization result D_phys is output as the basic input for the high-frequency detail enhancement stage.

[0148] D3: Sparse coding is performed on the high-frequency components to obtain the denoised high-frequency detail features.

[0149] Optionally, sparse coding is performed on the high-frequency components to obtain a combination of preliminarily denoised high-frequency signals and semantically optimized basis functions. Based on the combination of basis functions and the selected sparse coefficients, signal reconstruction is performed to obtain denoised high-frequency detail features.

[0150] Specifically, noise suppression and edge preservation are achieved through nonlocal sparse representation and adaptive basis function optimization: First, a complete dictionary Φ is constructed using multi-basis function fusion technology, including Gabor (direction-sensitive), curvelet (anisotropic), and DCT (global smoothing) basis functions, covering the requirements for multi-directional high-frequency feature representation. Then, the high-frequency component D_high is sparsely encoded using the Orthogonal Matching Pursuit (OMP) algorithm, iteratively solving the optimization objective of "minimizing the error between the high-frequency component and dictionary reconstruction under sparsity constraints." The most relevant atoms to the residual are selected and the sparse coefficients are updated sequentially until the residual energy drops to 10% of the initial signal energy or the maximum number of iterations (20) is reached. A significant sparse coefficient preservation strategy is used to filter the sparse coefficients output by OMP, retaining significant coefficients with an energy percentage >5% and generating a preliminary denoised high-frequency signal. Simultaneously, a channel attention mechanism is used to dynamically weight the contributions of the basis functions in the dictionary Φ and output a semantically optimized combination of basis functions.

[0151] Signal reconstruction is performed based on the optimized basis function combination and the selected sparse coefficients to generate the final denoised high-frequency detail D_high', which significantly improves the signal-to-noise ratio and edge sharpness of the sub-millimeter structure, eliminates the blurring and distortion problem in traditional methods, and greatly reduces the computational complexity.

[0152] D4: By jointly optimizing the intermediate restoration depth map and the high-frequency detail features of the denoised data through pixel-level weighted fusion, a restoration depth map that meets physical consistency is obtained.

[0153] Specifically, the physical constraint optimization result D_phys and the optimized high-frequency details D_high' are jointly optimized by pixel-level weighted fusion, simultaneously constraining physical laws (rigid body dynamics + optical path tracing loss) and high-frequency reconstruction errors, and finally outputting the depth map D_repaired.

[0154] The physically consistent repair depth map D_repaired is represented as follows:

[0155] D repaired (x,y)=f θ (F fused D high ,FP)f θ In a neural network model, D_repaired represents the model's output during training, while the loss function (including physical constraints) is dynamically calculated during training. That is, after each forward propagation generates D_repaired, its derivative is immediately calculated, and the loss is backpropagated to adjust the model parameters. Therefore, D_repaired is progressively optimized during training. In a depth map, x and y typically represent the horizontal and vertical coordinate axes of the image, i.e., the pixel positions. The second derivative corresponds to the curvature change of the depth value in space, used to measure the smoothness of the surface.

[0156] This application addresses the problem that traditional methods suffer from sub-millimeter level detail blurring due to low-frequency dominance in restoration, and the restoration results violate physical laws (such as ghosting and reflection distortion). It separates high-frequency details from low-frequency structures using frequency domain decomposition technology, specifically optimizing edge sharpness and noise suppression. Furthermore, it introduces physical constraints (rigid body dynamics and optical path tracing) to force the restoration results to conform to real-world laws. This significantly enhances the ability to preserve high-frequency details, greatly improves edge sharpness, and significantly improves physical consistency, eliminating geometric distortion and optical artifacts.

[0157] S105 mainly involves the process of dynamic scene compensation and real-time repair, which can be achieved through the following steps:

[0158] E1: Taking the repaired depth map and the target depth map of the previous frame as input, the optical flow field storing pixel displacement vectors is obtained by processing the original RAFT network.

[0159] Optionally, the repaired depth map and the target depth map of the previous frame to be processed are used as dual-frame inputs, and motion vector extraction processing is performed on the dual-frame inputs to obtain the optical flow field that stores the pixel displacement vectors.

[0160] Specifically, lightweight optical flow prediction achieves efficient motion estimation for dynamic scenes through a cropped RAFT network architecture optimization and hardware acceleration. It uses the current frame repair result D_repaired (providing complete repaired depth information) and the historical frame depth map D_prev (data from the previous frame's output D_Out cached by the system; D_repaired is directly reused during the first frame processing) as dual-frame inputs. Temporal motion correlations are extracted by compressing the parameters of the original RAFT network (removing 6 redundant convolutional layers, reducing the number of channels from 256 to 128, and reducing the number of parameters by 80%). The computational cost is reduced to 1 / 4 by replacing the standard convolutions with depth-separable convolutions in the compressed network. TensorRT is applied to the network weights and activation values. INT8 quantization (with key channels retaining FP16 and combined with dynamic range calibration) compresses the optical flow prediction latency from 15ms to 5ms; at the same time, optical flow field initialization is performed on the global layout information of the multi-scale feature pyramid FP_1 / 8 to enhance the large-scale motion capture capability; finally, the optimized network architecture is used to extract motion vectors from the dual-frame input and output the optical flow field F_flow that stores pixel displacement vectors, providing a high-precision motion compensation basis for spatiotemporal fusion.

[0161] In one scenario, when processing the first frame of a video sequence (without historical data), the repair result D_repaired of the current frame is directly used as the initial value D_prev. Subsequently, the temporally consistent depth map D_out output from the previous frame is cached as D_prev for use in the next frame.

[0162] E2: Based on the optical flow field, the invisible areas of the repair depth map will be determined and filled based on the repair depth map.

[0163] Specifically, based on the optical flow field E_flow, the historical frame depth map D_prev (the output D_out of the previous frame from the system cache) is first geometrically aligned by a forward warp operation, and then mapped to the current frame coordinate system (using bilinear interpolation to preserve geometric continuity). The pixel-level displacement relationship is then marked with occlusion regions by optical flow reverse consistency verification to identify regions that are invisible due to object movement.

[0164] For occluded areas, the current frame's repair result D_repaired is used to fill the area directly, thus eliminating motion artifacts caused by misaligned historical information.

[0165] E3: Based on the mesoscale features of the feature pyramid and the hole mask, the visible area of ​​the repair depth map is repaired to obtain the fused weight map.

[0166] Specifically, the current frame repair result D_repaired and the historical frame warping result D_prev_warped in the non-occluded area are adaptively fused using the mesoscale feature FP_1 / 4 (used to locate dynamic objects such as vehicles and vegetation) and the spatial mask M_mask (to identify the area to be repaired). This generates a pixel-level fusion weight map Weight_map, which is used to control the contribution ratio of the current frame and the historical frame in the non-occluded area.

[0167] Optionally, the hole M_mask can be obtained by performing a Poisson disk sampling operation on the spatially aligned depth map D_aligned to simulate a hole region of 30%-50%, and creating a hole mask M_mask based on the sampling results.

[0168] E4: By fusing the weighted map, the invisible and visible regions after processing are merged to obtain a temporally consistent target depth map.

[0169] Specifically, the non-occluded regions are weighted and fused by the weight map Weight_map, and combined with the independent filling results of the occluded regions, a complete temporally consistent depth map D_out is synthesized and output. This achieves efficient real-time processing at 4K resolution, significantly improving the dynamic hole repair effect, effectively suppressing inter-frame jitter, and preserving high-frequency details.

[0170] This application addresses the issues of high response latency and inability to handle transient occlusion and motion blur in traditional static restoration strategies. It designs a lightweight optical flow prediction module combined with a multi-scale spatiotemporal fusion strategy and a dynamic weight adjustment mechanism to balance information from historical and current frames. This significantly improves the efficiency of dynamic scene processing, meeting the requirements for high-resolution real-time interaction; the transient hole restoration effect is significantly optimized, and motion artifacts and inter-frame jitter are effectively suppressed.

[0171] like Figure 2 As shown, based on the same inventive concept as the above image reconstruction method, this application embodiment also provides an image reconstruction apparatus, including a determination unit 21 and a data processing unit 22.

[0172] The determining unit 21 is used to: determine the target depth map and the target RGB map of the image to be processed; wherein, the target depth map is a spatially aligned depth map, and the target RGB map is an enhanced RGB map;

[0173] Data processing unit 22 is used to: process the target depth map to obtain low-frequency components, high-frequency components, edge features, and feature pyramids;

[0174] The data processing unit 22 is also used to: obtain semantically enhanced deep features based on the target RGB image, low-frequency components, edge features, and feature pyramid;

[0175] The data processing unit 22 is also used to: apply the target RGB image, high-frequency components, feature pyramid and semantically enhanced deep features to generate a repair depth map that satisfies physical consistency;

[0176] The data processing unit 22 is also used to: apply a multi-feature pyramid to process the repair depth map and generate a target depth map that satisfies temporal consistency.

[0177] In one alternative implementation, the data processing unit 22 is specifically used for:

[0178] Perform dual-tree complex wavelet transform decomposition on the target depth map to obtain multiple sub-bands, and process the original low-frequency sub-bands in the multiple sub-bands to obtain low-frequency components;

[0179] Target subbands that meet the conditions are selected from multiple subbands, and the target subbands are processed to obtain high-frequency components; wherein, the target subband is the subband with energy greater than a set energy threshold among multiple subbands;

[0180] The low-frequency and high-frequency components are input into a lightweight dual-branch network to generate optimized low-frequency detail features and optimized high-frequency detail features. The optimized low-frequency detail features and optimized high-frequency detail features are then further optimized to obtain edge features.

[0181] Edge features and optimized low-frequency detail features are spliced ​​together, and geometric constraints are applied to the splicing result to obtain geometrically optimized features. The geometrically optimized features are then input into a backbone network to construct a feature pyramid.

[0182] In one alternative implementation, the data processing unit 22 is specifically used for:

[0183] Perform artifact suppression on the original low-frequency subbands in multiple subbands;

[0184] Perform energy threshold filtering on the original low-frequency subband after artifact suppression;

[0185] Perform convolution compression on the original low-frequency subband after energy filtering to obtain the low-frequency components.

[0186] In one alternative implementation, the data processing unit 22 is specifically used for:

[0187] Perform directional fusion operation on the target sub-band to obtain the fused multi-directional enhanced signal;

[0188] Perform orthogonal matching pursuit operation on the fused multi-directional enhanced signal to obtain sparse coding coefficients;

[0189] By performing channel attention on the sparse coding coefficients, semantically optimized sparse coding coefficients are obtained.

[0190] The high-frequency components are obtained by reconstructing the sparse coding coefficients of semantic optimization.

[0191] In one alternative implementation, the data processing unit 22 is specifically used for:

[0192] Based on low-frequency components and bottom-level features in the feature pyramid, a weighted query vector is determined; multi-level feature extraction is performed on the target RGB image based on the defined network to obtain semantic features at each layer; at least one layer of features in the feature pyramid is processed through feature concatenation and convolution to obtain the medium-sized object location key and detail enhancement value; noise suppression is performed on each layer of features in the feature pyramid to obtain the initial deep features for semantic enhancement.

[0193] Cross-modal attention computation is performed on the weighted query vector, medium-sized object location key, and detail enhancement value to obtain the first fusion feature. Then, geometric fidelity processing is performed on the first fusion feature and the low-frequency component to obtain the intermediate semantic enhancement feature.

[0194] Edge enhancement processing is performed on the target layer features and edge images of the feature pyramid to obtain the second fused features; the wall integrity mask is calculated by building category damage, and the leaf contour mask is calculated by vegetation edge loss;

[0195] The initial semantic enhancement deep features, intermediate semantic enhancement features, wall integrity mask, leaf outline mask, and initial semantic enhancement deep features are fused in a high dimension to obtain the semantic enhancement deep features.

[0196] In one alternative implementation, the data processing unit 22 is specifically used for:

[0197] By combining features and performing 3*3 convolution, key vector generation is performed on the third layer features in the feature pyramid and the third layer features of the set network to obtain the location key for medium-sized objects.

[0198] By combining features and performing 1*1 convolution, value vector generation is performed on the first layer features in the feature pyramid and the first layer features of the defined network to obtain detail enhancement values.

[0199] In one alternative implementation, the data processing unit 22 is specifically used for:

[0200] A lightweight decoder is used to concatenate the semantically enhanced deep features and the features of each layer of the feature pyramid along the channel dimension. Then, multiple deconvolution layers are used to gradually upsample the concatenated features to the original image resolution to obtain a preliminary repaired depth map.

[0201] Rigid region masking is performed on the preliminary repair depth map to obtain the intermediate repair depth map.

[0202] Sparse coding of high-frequency components yields denoised high-frequency detail features;

[0203] By jointly optimizing the intermediate restoration depth map and the high-frequency detail features of the denoised data through pixel-level weighted fusion, a restoration depth map that satisfies physical consistency is obtained.

[0204] In one alternative implementation, the data processing unit 22 is specifically used for:

[0205] Rigid region masking is performed on the preliminary repair depth map to determine the target region that needs to be subject to rigid constraints. The parameters of the target region are constrained, and the target region is differentiated and modeled to obtain rigid body dynamic constraints.

[0206] The target RGB image is processed by material property analysis, and a reference depth that conforms to the laws of optical physics is constructed based on the analyzed material properties. High reflectivity regions are segmented by the third layer of features in the feature pyramid to generate a mask for the high reflectivity surface. For the regions marked by the mask, the error between the preliminary repair depth map and the theoretical reflection depth is minimized to obtain the optical path tracing constraint.

[0207] The network parameters of the integrated physical loss function of rigid body dynamics constraints and optical path tracing constraints are optimized by directional propagation, and the intermediate repair depth map is determined based on the processed parameters.

[0208] In one alternative implementation, the data processing unit 22 is specifically used for:

[0209] Sparse coding is performed on the high-frequency components to obtain a preliminary denoised high-frequency signal and a semantically optimized basis function combination. Based on the basis function combination and the selected sparse coefficients, signal reconstruction is performed to obtain the denoised high-frequency detail features.

[0210] In one alternative implementation, the data processing unit 22 is specifically used for:

[0211] The repaired depth map and the target depth map of the previous frame to be processed are used as inputs. The original RAFT network is processed to obtain the optical flow field that stores the pixel displacement vector.

[0212] Based on the optical flow field, the invisible areas of the repair depth map will be identified and filled based on the repair depth map;

[0213] The visible areas of the repair depth map are repaired based on the mesoscale features of the feature pyramid and the hole mask to obtain the fused weight map.

[0214] By fusing the weighted map, the invisible and visible regions after processing are merged to obtain a temporally consistent target depth map.

[0215] In one alternative implementation, the data processing unit 22 is specifically used for:

[0216] The repaired depth map and the target depth map of the previous frame to be processed are used as dual-frame inputs, and motion vector extraction processing is performed on the dual-frame inputs to obtain the optical flow field that stores the pixel displacement vectors.

[0217] In one alternative implementation, the determining unit 21 is specifically used for:

[0218] Obtain the original depth map and original RGB image of the image to be processed;

[0219] Perform spatial alignment on the original depth map to obtain a spatially aligned depth map, and then enhance the original RGB image to obtain an enhanced RGB image.

[0220] In one alternative implementation, the determining unit 21 is specifically used for:

[0221] Construct a perspective transformation matrix based on the calibration parameters of the camera used to capture the image to be processed;

[0222] By applying the perspective transformation matrix, a function alignment operation is performed on the original depth map to obtain a spatially aligned depth map;

[0223] The original RGB image is enhanced to obtain the enhanced RGB image, including:

[0224] Add random noise to each color channel of the original RGB image;

[0225] Perform histogram stretching on the RGB image after adding random noise;

[0226] Normalize the stretched RGB image to obtain the enhanced RGB image.

[0227] The image reconstruction apparatus proposed in this application adopts the same inventive concept as the image reconstruction method described above and can achieve the same beneficial effects, so it will not be described again here.

[0228] Based on the same inventive concept as the image reconstruction method described above, this application also provides an electronic device, which may specifically be a desktop computer, portable computer, smartphone, tablet computer, personal digital assistant (PDA), server, etc. Figure 3 As shown, the electronic device may include a processor 301 and a memory 302.

[0229] Processor 301 can be a general-purpose processor, such as a central processing unit (CPU), digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0230] Memory 302, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memory, magnetic disk, optical disk, etc. Memory is any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. Memory 302 in this embodiment may also be a circuit or any other device capable of implementing storage functions for storing program instructions and / or data.

[0231] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned computer storage medium can be any available medium or data storage device that a computer can access, including but not limited to: mobile storage devices, random access memory (RAM), magnetic storage (e.g., floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical storage (e.g., CDs, DVDs, BDs, HVDs, etc.), and semiconductor storage (e.g., ROMs, EPROMs, EEPROMs, non-volatile memory (NAND flash), solid-state drives (SSDs)) and other media capable of storing program code.

[0232] Alternatively, if the integrated units described above in this application are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods of the various embodiments of this application. The aforementioned storage medium includes: mobile storage devices, random access memory (RAM), magnetic memory (e.g., floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical memory (e.g., CDs, DVDs, BDs, HVDs, etc.), and semiconductor memory (e.g., ROMs, EPROMs, EEPROMs, non-volatile memory (NAND flash), solid-state drives (SSDs), etc.) and other media capable of storing program code.

[0233] Based on the same inventive concept, this application also provides a computer program product, which includes computer program code that, when run on a computer, causes the computer to execute any of the image reconstruction methods discussed above. Since the principle by which the above computer program product solves the problem is similar to that of image reconstruction, the implementation of the above computer program product can be referred to the implementation of the method, and repeated details will not be elaborated further.

[0234] The above embodiments are only used to provide a detailed description of the technical solutions of this application. However, the description of the above embodiments is only for the purpose of helping to understand the methods of the embodiments of this application and should not be construed as a limitation on the embodiments of this application. Any changes or substitutions that can be easily conceived by those skilled in the art should be covered within the protection scope of the embodiments of this application.

Claims

1. A method of image reconstruction, characterized by, The method comprises the following steps: determine the target depth map and the target RGB image of the image to be processed; wherein the target depth map is a spatially aligned depth map, and the target RGB image is an enhanced RGB image; processing the target depth map to obtain low-frequency components, high-frequency components, edge features and feature pyramids; based on the target RGB image, the low-frequency components, the edge features and the feature pyramids, a semantically enhanced depth feature is obtained; applying the target RGB image, the high-frequency components, the feature pyramids and the semantically enhanced depth feature to generate a repaired depth map that satisfies physical consistency; applying the multi-feature pyramid to the repaired depth map to generate a target depth map that satisfies temporal consistency.

2. The method of claim 1, wherein, The processing of the target depth map to obtain low-frequency components, high-frequency components, edge features and feature pyramids comprises: performing a dual-tree complex wavelet transform decomposition operation on the target depth map to obtain a plurality of subbands, and processing the original low-frequency subband in the plurality of subbands to obtain the low-frequency components; selecting target subbands that meet the conditions from the plurality of subbands, and processing the target subbands to obtain the high-frequency components; wherein the target subbands are subbands in the plurality of subbands whose energy is greater than a set energy threshold; input the low-frequency components and the high-frequency components into a lightweight dual-branch network respectively to generate optimized low-frequency detail features and optimized high-frequency detail features; optimize the optimized low-frequency detail features and the optimized high-frequency detail features to obtain the edge features; splicing the edge features and the optimized low-frequency detail features, and performing geometric constraint processing on the splicing result to obtain geometric optimization features; input the geometric optimization features into a set backbone network to construct a feature pyramid.

3. The method of claim 2, wherein, The processing of the original low-frequency subband in the plurality of subbands to obtain the low-frequency components comprises: performing an artifact suppression operation on the original low-frequency subband in the plurality of subbands; performing an energy threshold filtering operation on the original low-frequency subband after the artifact suppression operation; performing a convolution compression operation on the original low-frequency subband after the energy filtering to obtain the low-frequency components.

4. The method of claim 2, wherein, The processing of the target subband to obtain the high-frequency components comprises: performing a direction fusion operation on the target subband to obtain a fused multi-direction enhancement signal; performing an orthogonal matching pursuit operation on the fused multi-direction enhancement signal to obtain sparse coding coefficients; performing a channel attention mechanism operation on the sparse coding coefficients to obtain semantically optimized sparse coding coefficients; performing a reconstruction operation on the semantically optimized sparse coding coefficients to obtain the high-frequency components.

5. The method of claim 1, wherein, The feature pyramid is a multi-scale feature pyramid; the obtaining of the semantically enhanced depth feature based on the target RGB image, the low-frequency components, the edge features and the feature pyramids comprises: determine a weighted query vector based on the low-frequency component and the bottom-layer features in the feature pyramid; perform multi-level feature extraction on the target RGB image based on a preset network to obtain semantic features of each layer; perform processing on at least one layer of features in the feature pyramid through feature splicing and convolution to obtain a medium object positioning key and a detail enhancement value; and perform noise suppression processing on the features of each layer of the feature pyramid to obtain an initial depth feature with enhanced semantics. perform cross-modal attention calculation processing on the weighted query vector, the medium object positioning key and the detail enhancement value to obtain a first fusion feature, and perform geometric fidelity processing on the first fusion feature and the low-frequency component to obtain an intermediate semantic enhancement feature; perform edge enhancement processing on the target layer features of the feature pyramid and the edge image to obtain a second fusion feature; calculate a wall integrity mask through building category damage, and calculate a leaf contour mask through vegetation edge loss; perform high-dimensional fusion processing on the initial depth feature with enhanced semantics, the intermediate semantic enhancement feature, the wall integrity mask, the leaf contour mask and the initial depth feature with enhanced semantics to obtain a depth feature with enhanced semantics.

6. The method of claim 5, wherein, The feature pyramid is a four-scale feature pyramid, and the processing of the features of at least one layer in the feature pyramid through feature splicing and convolution to obtain a medium physical positioning key and a detail enhancement value includes: perform key vector generation processing on the third layer features in the feature pyramid and the third layer features of the preset network through feature splicing and 3*3 convolution to obtain a medium object positioning key; perform value vector generation processing on the first layer features in the feature pyramid and the first layer features of the preset network through feature splicing and 1*1 convolution to obtain a detail enhancement value.

7. The method of claim 1, wherein, The application of the target RGB image, the high-frequency component, the feature pyramid and the depth feature with enhanced semantics to generate a repaired depth map satisfying physical consistency includes: perform channel dimension splicing on the depth feature with enhanced semantics and the features of each layer of the feature pyramid through a lightweight decoder, and gradually up-sample the spliced features to the original image resolution through multiple deconvolution layers to obtain a preliminary repaired depth map; perform rigid region mask extraction processing on the preliminary repaired depth map to obtain an intermediate repaired depth map; perform sparse coding processing on the high-frequency component to obtain a denoised high-frequency detail feature; perform joint optimization on the intermediate repaired depth map and the denoised high-frequency detail feature through pixel-level weighted fusion to obtain a repaired depth map satisfying physical consistency.

8. The method of claim 7, wherein, The rigid region mask extraction processing on the preliminary repaired depth map to obtain an intermediate repaired depth map includes: perform rigid region mask extraction processing on the preliminary repaired depth map to determine a target region to which a rigid constraint needs to be applied, perform constraint processing on the parameters of the target region, and perform differential modeling processing on the target region to obtain a rigid body dynamics constraint; The target RGB image is subjected to material attribute analysis processing, and a reference depth conforming to optical physical laws is constructed based on the analyzed material attributes; a high-reflective region is segmented by a third layer of features in the feature pyramid, and a mask of a high-reflective surface is generated; an error minimization process is performed on the preliminary repaired depth map and the theoretical reflection depth for the region identified by the mask, to obtain a light path tracking constraint; The network parameter optimization process is performed on the comprehensive physical loss function of the rigid body dynamics constraint and the light path tracking constraint by directional propagation, and an intermediate repaired depth map is determined based on the processed parameters.

9. The method of claim 7, wherein, The sparse coding process is performed on the high-frequency component to obtain a denoised high-frequency detail feature, including: The sparse coding process is performed on the high-frequency component to obtain a preliminary denoised high-frequency signal and a semantic optimized basis function combination, and a signal reconstruction process is performed based on the basis function combination and the screened sparse coefficients to obtain a denoised high-frequency detail feature.

10. The method of claim 1, wherein, The feature pyramid is applied to process the repaired depth map to generate a target depth map that meets the timing consistency, including: The repaired depth map and the target depth map of the previous frame of image to be processed are taken as inputs, and a flow field storing pixel displacement vectors is obtained by processing the original RAFT network; Based on the flow field, an invisible region of the repaired depth map is determined, and the repaired depth map is filled based on the repaired depth map; The visible region of the repaired depth map is repaired according to the mesoscale feature of the feature pyramid and the hole mask to obtain a fusion weight map; The invisible region and the visible region after processing are fused by the fusion weight map to obtain a target depth map that meets the timing consistency.

11. The method of claim 10, wherein, The repaired depth map and the target depth map of the previous frame of image to be processed are taken as inputs, and a flow field storing pixel displacement vectors is obtained by processing the original RAFT network, including: The repaired depth map and the target depth map of the previous frame of image to be processed are taken as inputs, and a flow field storing pixel displacement vectors is obtained by processing the original RAFT network, including:

12. The method of claim 1, wherein, The target depth map and the target RGB image of the image to be processed are determined, including: An original depth map and an original RGB image of the image to be processed are obtained; A spatial alignment operation is performed on the original depth map to obtain a spatially aligned depth map, and an enhanced RGB image is obtained by enhancing the original RGB image.

13. The method of claim 12, wherein, A spatial alignment operation is performed on the original depth map to obtain a spatially aligned depth map, including: A perspective transformation matrix is constructed based on the calibration parameters of the camera used to capture the image to be processed; The perspective transformation matrix is applied to perform a function alignment operation on the original depth map to obtain a spatially aligned depth map; The original RGB image is enhanced to obtain an enhanced RGB image, including: Random noise is added to each color channel of the original RGB image; A histogram stretching operation is performed on the RGB image with random noise added; A normalization operation is performed on the stretched RGB image to obtain an enhanced RGB image.

14. An image reconstruction apparatus, characterized by comprising: including: A determining unit is configured to determine a target depth map and a target RGB map of an image to be processed, wherein the target depth map is a spatially aligned depth map, and the target RGB map is an enhanced RGB map; A data processing unit is configured to process the target depth map to obtain a low-frequency component, a high-frequency component, an edge feature, and a feature pyramid; The data processing unit is further configured to obtain a semantically enhanced depth feature based on the target RGB map, the low-frequency component, the edge feature, and the feature pyramid; The data processing unit is further configured to apply the target RGB map, the high-frequency component, the feature pyramid, and the semantically enhanced depth feature to generate a repaired depth map satisfying physical consistency; The data processing unit is further configured to apply the multi-feature pyramid to process the repaired depth map to generate a target depth map satisfying temporal consistency.

15. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the steps of the method in any one of claims 1 to 13.

16. A computer-readable storage medium having stored thereon computer program instructions, wherein, The computer program instructions are executed by the processor to implement the steps of the method in any one of claims 1 to 13.