A visual backbone network based on frequency domain decoupling and microstructure compensation and an image processing method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-08
- Publication Date
- 2026-08-11
AI Technical Summary
[0005]因此,本发明提供了一种基于频域解耦与微结构补偿的视觉骨干网络及图像处理方法,解决现有方法在空间域内进行特征提取与融合,不能有效区分图像中由低频结构信息与高频纹理细节所主导的不同成分,导致在重建或压缩任务中出现结构模糊或细节丢失的问题,现有方法采用统一的上采样策略,忽略不同频率成分在空间连续性与细节恢复方面的差异,影响重建图像的清晰度与真实感问题
Smart Images

Figure CN122550982A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and computer vision technology, and in particular to a visual backbone network and image processing method based on frequency domain decoupling and microstructure compensation. Background Technology
[0002] Visual backbone networks and their image processing methods are important foundational technologies in the fields of computer vision, image understanding, and intelligent perception. Through multi-layer convolution, downsampling, and feature aggregation, hierarchical representation of the spatial structure of images is achieved. By introducing multi-scale feature fusion, attention mechanisms, and encoder-decoder structures, and by introducing normalization, rearrangement, or channel alignment strategies in the data preprocessing stage, the stability and efficiency of network training and inference can be improved.
[0003] However, existing methods still have shortcomings. Existing methods perform feature extraction and fusion in the spatial domain, which cannot effectively distinguish between different components in an image dominated by low-frequency structural information and high-frequency texture details. This leads to problems such as structural blurring or loss of details in reconstruction or compression tasks. Existing methods adopt a uniform upsampling strategy, ignoring the differences in spatial continuity and detail recovery of different frequency components, which affects the clarity and realism of the reconstructed image. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a visual backbone network and image processing method based on frequency domain decoupling and microstructure compensation, which solves the problem that existing methods cannot effectively distinguish between different components in an image dominated by low-frequency structural information and high-frequency texture details when performing feature extraction and fusion in the spatial domain. This leads to structural blurring or loss of details in reconstruction or compression tasks. Existing methods use a uniform upsampling strategy, ignoring the differences in spatial continuity and detail recovery of different frequency components, which affects the clarity and realism of the reconstructed image.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides a visual backbone network and image processing method based on frequency domain decoupling and microstructure compensation, comprising, The input image is collected and organized into a three-dimensional tensor. A pixel coordinate system is established, and the dynamic range is set based on the bit depth of the input image. The three-dimensional tensor is normalized to obtain a normalized pixel tensor. The periodic rearrangement ratio is set based on empirical rules, and the periodic rearrangement operation is performed in combination with the normalized pixel tensor to obtain the output tensor. A three-dimensional source index is established for the output tensor, and a lossless injection tensor is constructed. The lossless injection tensor is used as the input of the encoder layer 0. The structural path features and high-frequency residual tensors are calculated and fused to obtain multi-scale features. The fused features at each scale are spatially flattened to construct a token matrix. The query matrix and key matrix are obtained through linear projection and rotational position encoding. The gating weights are calculated based on the multi-scale features. The geometric consistency strength is calculated based on the query matrix and key matrix. The geometric consistency strength is fused with the gating weights to obtain stable output features. The highest resolution decoding feature map is constructed from the stable output features. The texture compensation vector is extracted, and the texture compensation vector is added to the highest resolution decoding feature map to obtain the compensation features. The reconstructed image is obtained by combining the dynamic range.
[0007] As a preferred embodiment of the visual backbone network and image processing method based on frequency domain decoupling and microstructure compensation described in this invention, wherein: the normalization processing of the three-dimensional tensor to obtain a normalized pixel tensor includes: Obtain the target output type of the task through the API interface, record the data source type of this task, and collect the input image from the corresponding data channel; The input image is composed into a three-dimensional tensor. The dynamic range is calculated based on the bit depth of the input image. The three-dimensional tensor is divided by the dynamic range to obtain a normalized pixel tensor.
[0008] As a preferred embodiment of the visual backbone network and image processing method based on frequency domain decoupling and microstructure compensation described in this invention, wherein: the step of establishing a three-dimensional source index for the output tensor and constructing a lossless injection tensor includes: The periodic rearrangement ratio is set based on empirical rules. The square of the periodic rearrangement ratio is calculated, and the square is multiplied by the number of channels of the normalized pixel tensor to obtain the number of rearrangement channels. The output tensor is obtained by combining the rearrangement size. Based on the channel index expansion relationship defined in the input image, a three-dimensional source index is obtained; A lossless injection tensor is constructed based on the 3D source index and the output tensor.
[0009] As a preferred embodiment of the visual backbone network and image processing method based on frequency domain decoupling and microstructure compensation described in this invention, wherein: the calculation and fusion of structural path features and high-frequency residual tensors to obtain multi-scale features includes: Define the lossless injection tensor as the input of the 0th layer of the encoder, and set the number of encoder layers to 1. ,right The input of the layer is physically folded to obtain Layer folding tensor; The folded tensor is layer normalized to obtain a normalized tensor. Based on the normalized tensor, the structural path features are obtained through linear projection. High-frequency residual tensors are extracted from folded tensors using local convolution residuals; The high-frequency residual tensor and structural path features are superimposed element by element to obtain... Layer fusion characteristics; Will The layer fusion feature is set as The input of the layer is used to repeat the above operation, traversing... Each layer, obtained The fused features of each layer are stacked horizontally to generate multi-scale features.
[0010] As a preferred embodiment of the visual backbone network and image processing method based on frequency domain decoupling and microstructure compensation described in this invention, wherein: the fusion of geometric consistency strength and gating weights to obtain stable output features includes: Spatially flatten the fused features at each scale in the multi-scale features to obtain a token matrix, construct a two-dimensional position index of the token matrix, define the angular frequency, construct a rotation angle for the two-dimensional position index of the token matrix, and obtain the query matrix and key matrix after injecting rotation position encoding; Spatial mean is calculated for the fused features at each scale in the multi-scale features to obtain the channel summary vector. The channel autocorrelation matrix is calculated using a normalized cosine similarity matrix on the channel summary vector. The elements in the autocorrelation matrix are summed and averaged. The mean is then processed by the Sigmoid function to generate gate weights. The spatial geometric correlation matrix is calculated based on the query matrix and key matrix after the injection rotation position encoding. The average value of the spatial geometric correlation matrix is obtained by calculating the geometric consistency strength by taking the row direction. The spatial gating weight is obtained by multiplying the geometric consistency strength and the gating weight. The product of the spatial gating weights and the fused features is calculated to obtain the enhanced features. The sum of the enhanced features and the fused features is then calculated to obtain the stable output features.
[0011] As a preferred embodiment of the visual backbone network and image processing method based on frequency domain decoupling and microstructure compensation described in this invention, wherein: the construction of the highest resolution decoding feature map for stable output features includes: Extract encoder The stable output characteristics of the layer are set as the initial value for decoding. Linear projection moments are applied to the initial value for decoding to generate the channel alignment matrix. Set a folding factor, apply the inverse folding operator to the channel alignment matrix to obtain upsampled features, and then map them using the standard inverse PixelShuffle formula to obtain the mapped values; Extract encoder layer The stable output features of the layer are set as jump connection features. Linear projection is used on the jump connection features to generate a jump connection feature tensor. The jump-connected feature tensor and the upsampled feature are added together to obtain intermediate features, which are then normalized to obtain normalized intermediate features. Perform a 3×3 two-dimensional convolution on the normalized intermediate features to obtain a refined residual. Then, perform an addition operation on the refined residual and the normalized intermediate features to obtain the highest resolution decoded feature map.
[0012] As a preferred embodiment of the visual backbone network and image processing method based on frequency domain decoupling and microstructure compensation described in this invention, wherein: the extraction of texture compensation vector includes: Local feature blocks are constructed using local receptive field modeling, and spatial average pooling is performed on the local feature blocks to obtain local microstructure fingerprint vectors. The gradient magnitude is calculated for the highest resolution decoded feature map, the positions with gradient magnitude greater than the gradient magnitude threshold are selected, the local feature blocks corresponding to the positions are extracted, and the channel direction mean is statistically analyzed for the local feature blocks to obtain the texture prior vector. The vectors are arranged horizontally to obtain the generative texture prior knowledge base. The similarity between the local microstructure fingerprint vector and the texture prior vector is calculated using the cosine similarity formula. The similarity is then normalized to obtain a normalized weight. This weight is then multiplied by the local microstructure fingerprint vector and summed to obtain the texture compensation vector.
[0013] As a preferred embodiment of the visual backbone network and image processing method based on frequency domain decoupling and microstructure compensation described in this invention, wherein obtaining the compensation features and combining them with the dynamic range to obtain the reconstructed image includes: The result of multiplying the texture compensation vector by the scalar weight is added to the highest resolution decoded feature map to obtain the compensation feature; Calculate the mean of the compensation feature, and subtract the mean from the compensation feature to obtain the stable compensation feature; The stable compensation features are divided into a first feature subset and a second feature subset along the channel dimension; The first feature subset is entered into the smoothing baseline path, and local feature preprocessing and bilinear interpolation upsampling are performed sequentially to obtain the constrained smoothing component. The second feature subset enters the sharp detail path, and spatial resolution is improved by channel expansion mapping combined with sub-pixel rearrangement (PixelShuffle) technology to obtain rearranged sharp components; the high-frequency detail convolutional layer that generates rearranged sharp components is processed by progressive zero-initialization. The rearranged sharp components and the constrained smooth components are weighted and summed to obtain the fused feature map; The product of the fused feature map and the dynamic range is calculated to obtain the reconstructed image.
[0014] The beneficial effects of this invention are as follows: This invention constructs the highest resolution decoding feature map from stable output features, extracts texture compensation vectors, performs addition operations between the texture compensation vectors and the highest resolution decoding feature map to obtain compensation features, and combines the dynamic range to obtain the reconstructed image; it achieves coordinated enhancement of structure and detail, improves the stability of multi-scale features, and enhances the clarity of the reconstructed image. Attached Figure Description
[0015] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a flowchart of the visual backbone network and image processing method based on frequency domain decoupling and microstructure compensation in Example 1.
[0017] Figure 2 This is a schematic diagram of the spatial gating attention mechanism of the visual backbone network and image processing method based on frequency domain decoupling and microstructure compensation in Example 1. Detailed Implementation
[0018] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0019] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0020] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0021] Example 1, referring to Figure 1 and Figure 2 This is the first embodiment of the present invention, which provides a visual backbone network and image processing method based on frequency domain decoupling and microstructure compensation, including the following steps: S1. Collect the input image and organize it into a three-dimensional tensor, establish a pixel coordinate system, and set the dynamic range based on the bit depth of the input image. Normalize the three-dimensional tensor to obtain a normalized pixel tensor. Set the periodic rearrangement ratio based on empirical rules, and perform periodic rearrangement operation in combination with the normalized pixel tensor to obtain the output tensor. Establish a three-dimensional source index for the output tensor and construct a lossless injection tensor. Specifically, the 3D tensor is normalized to obtain a normalized pixel tensor, including: The pixels of the input image are arranged into a three-dimensional tensor in the order of height-width-channel, and a pixel coordinate system is defined, including defining the top left corner as the origin, increasing horizontally to the right, and decreasing vertically downwards. The dynamic range is calculated based on bit depth, using the following formula: , in, For dynamic range, For depth; Dividing the 3D tensor by the dynamic range yields the normalized pixel tensor.
[0022] By performing unified 3D tensor modeling and normalization on the input images, and under the premise that the task target and data source type are explicitly distinguished, numerical consistency across imaging modalities is achieved. By completing dynamic range alignment before the image enters subsequent processing, images from different depths and different sensors are comparable in the same numerical space, fundamentally avoiding feature bias problems caused by differences in imaging conditions. By clarifying the construction method of the pixel coordinate system, the semantics of spatial location remain stable and continuous in the subsequent encoding and decoding process, which is conducive to the long-term preservation of structural information. This lays a reliable numerical and spatial foundation for subsequent multi-scale modeling and fine reconstruction.
[0023] Furthermore, a three-dimensional source index is established for the output tensor to construct a lossless injection tensor, including: The periodic rearrangement factor is set based on empirical rules. The spatial size (width and height) of the normalized pixel tensor is divided by the periodic rearrangement factor. If the result is positive, the result is set as the rearrangement size and proceeded to the next operation; otherwise, no operation is performed. Calculate the square of the periodic rearrangement ratio, multiply the square by the number of channels of the normalized pixel tensor to obtain the number of rearrangement channels, and combine it with the rearrangement size to obtain the output tensor; Based on the channel index expansion relationship defined in the input image, a three-dimensional source index is obtained. The formula is: , in, To output the tensor channel index, For the original channel index, For line offset within the space block, For column offset within the space block, The periodic rearrangement rate; The lossless injection tensor is constructed based on the 3D source index and output tensor, using the following formula: , in, To inject tensors without loss, To output a tensor, The pixel position represents the row index, column index, and channel index of the output tensor.
[0024] By constructing a 3D source index and introducing a lossless injection mechanism, this step achieves structural rearrangement between the spatial and channel dimensions without introducing information compression or discarding. This approach breaks through the limitation of traditional rearrangement, which only focuses on changes in data form and ignores the traceability of the source. It enables features at any position in the output tensor to be back-mapped to the specific spatial block and channel source of the original image, thereby significantly enhancing the interpretability and stability of the system. This lossless injection method provides higher-dimensional structural freedom for the subsequent encoding stage, allowing spatial neighborhood relationships to be implicitly integrated into the channel expression, effectively improving the carrying capacity of complex textures and detailed structures.
[0025] S2. The lossless injection tensor is used as the input of the encoder layer 0. The structural path features and high-frequency residual tensor are calculated and fused to obtain multi-scale features. The fused features at each scale are spatially flattened to construct a token matrix. The query matrix and key matrix are obtained through linear projection and rotation position encoding. The gating weights are calculated based on the multi-scale features. The geometric consistency strength is calculated based on the query matrix and key matrix. The geometric consistency strength is fused with the gating weights to obtain stable output features. Specifically, structural path features and high-frequency residual tensors are calculated and fused to obtain multi-scale features, including: Define the lossless injection tensor as the input of the 0th layer of the encoder, and set the number of encoder layers to 1. The current layer number is The 0th layer is layer; right The input of the layer is physically folded with a folding factor of 2 to obtain... The folding tensor of the layer is used to obtain the shape relationship, and the formula is: , , , , in, For layer The folding tensor, For vertical indexing, For horizontal indexing, and The vertical and horizontal offsets within the folded block are determined by the folding factor. For layer Input, For layer height, For layer width, For layer The number of channels, , as well as For layer Height, width, and number of channels; Layer normalization is performed on the folded tensor to obtain a normalized tensor. Based on the normalized tensor, structural path features (low frequency) are obtained through linear projection. High-frequency residual tensors are extracted from folded tensors using local convolution residuals (two-dimensional convolution with a kernel size of 3×3); The high-frequency residual tensor and structural path features are superimposed element by element to obtain... Layer fusion characteristics; Will The layer fusion feature is set as The input of the layer is used to repeat the above operation, traversing... Each layer, obtained The fused features of each layer are stacked horizontally to generate multi-scale features.
[0026] By explicitly distinguishing between structural path features and high-frequency residual features during the encoding stage and constructing multi-scale features through layer-by-layer fusion, this step achieves collaborative modeling of low-frequency structural information and high-frequency detail information. Structural path features focus on the overall contour and geometric continuity, while high-frequency residual features accurately depict local changes and texture abrupt changes. The two form a complementary relationship during the layer-by-layer fusion process, avoiding the problem of detail loss or noise amplification that is easily caused by single-path modeling. Through the horizontal stacking of multiple scales, the system can simultaneously retain global consistency and local sensitivity under different receptive fields, so that the subsequent attention and decoding processes have a richer and more stable feature foundation.
[0027] Furthermore, the geometric consistency strength is fused with the gating weights to obtain stable output features, including: Spatially flatten the fused features at each scale in the multi-scale feature set to obtain a token matrix, and construct a two-dimensional position index for the token matrix, using the following formula: , , in, and The token in the token matrix The horizontal and vertical coordinates of each token The modulo operation represents taking the remainder after division; Define the angular frequency, construct the rotation angle for the two-dimensional position index of the token matrix, and combine the query matrix and key matrix to obtain the query matrix and key matrix after injecting the rotation position encoding. The formula is as follows: , , , , , in, For dimension angular frequency, For a single-head dimension, it represents the ratio of the number of channels to the number of attention heads in the shape relationship. and The rotation angle for the horizontal and vertical coordinates. For querying the matrix, The key matrix, For attention head index, and The initial query matrix and key matrix are obtained by linear projection from the token matrix; Spatial mean is taken for the fused features at each scale in the multi-scale features to obtain the channel summary vector. After "summation and averaging", the missing steps are supplemented: "...summation and averaging of the elements in the autocorrelation matrix, dividing the features into two parts along the channel dimension, using the parameterless gating mechanism (SimpleGate) of channel feature multiplication to activate, and combining the mean with the Sigmoid function to generate gating weights." The spatial geometric relevance matrix is calculated based on the query matrix and key matrix after injection rotation position encoding, using the following formula: , in, This is the spatial geometric correlation matrix. For transpose, This is the scaling factor; The geometric consistency strength is obtained by averaging the spatial geometric correlation matrix along the row direction, as shown in the formula: , in, For geometric consistency strength, For layer The query location index in For layer The number of tokens in the matrix For layer The key position index in; The spatial gating weight is obtained by multiplying the geometric consistency strength and the gating weight. The product of the spatial gating weights and the fused features is calculated to obtain the enhanced features. The sum of the enhanced features and the fused features is then calculated to obtain the stable output features.
[0028] By jointly modulating geometric consistency strength and channel gating weights, this step constructs a stable feature enhancement mechanism that is simultaneously constrained by spatial structure and channel statistical characteristics. Geometric consistency strength is used to characterize the intrinsic correlation stability between spatial locations, while gating weights reflect the importance of different channels in the overall representation. The fusion of the two effectively suppresses anomalous responses that are spatially inconsistent but numerically prominent. This mechanism makes feature enhancement no longer dependent on a single attention result, but rather carried out under the dual constraints of geometric rationality and semantic importance, thereby significantly improving the robustness and consistency of output features in complex scenarios.
[0029] S3. Construct the highest resolution decoding feature map for the stable output features, extract the texture compensation vector, perform an addition operation between the texture compensation vector and the highest resolution decoding feature map to obtain the compensation features, and combine the dynamic range to obtain the reconstructed image. Specifically, constructing the highest resolution decoded feature map from stable output features includes: Extract encoder The stable output characteristics of the layer are set as the initial decoding values. Linear projection moments are applied to the initial decoding values to generate a channel alignment matrix. The number of channels in the channel alignment matrix is limited by the following formula: , in, For the number of channels, For layer Decoding features in the channel alignment matrix, For layer Initial values for decoding; Set a folding factor, apply the inverse folding operator to the channel alignment matrix to obtain upsampled features, and then map them using the standard inverse PixelShuffle formula to obtain the mapped values; Extract encoder layer The stable output features of the layer are set as jump connection features. Linear projection is used on the jump connection features to generate a jump connection feature tensor. The jump-connected feature tensor and the upsampled feature are added together to obtain intermediate features, which are then normalized to obtain normalized intermediate features. Perform a 3×3 two-dimensional convolution on the normalized intermediate features to obtain a refined residual. Then, perform an addition operation on the refined residual and the normalized intermediate features to obtain the highest resolution decoded feature map.
[0030] By introducing channel alignment and inverse folding mechanisms in the decoding stage, and combining them with skip-connected features for step-by-step fusion, this step achieves stable reconstruction of high-resolution feature maps. Channel alignment ensures that features of different scales have consistent semantic dimensions before fusion, while inverse folding and upsampling effectively restore the spatial detail distribution. The introduction of skip-connected features enables direct interaction between high-level semantic information and low-level spatial details, avoiding the blurring problem commonly found in deep decoding. Residual refinement further compensates for local errors, enabling the final decoded feature map to maintain the correct overall structure while possessing higher detail restoration capabilities.
[0031] Furthermore, the texture compensation vector is extracted, including: Local feature blocks are constructed using local receptive field modeling, and spatial average pooling is performed on the local feature blocks to obtain local microstructure fingerprint vectors. The gradient magnitude is calculated for the highest resolution decoded feature map. Locations with gradient magnitude greater than the gradient magnitude threshold (set based on statistical analysis) are selected, and the local feature blocks corresponding to the locations are extracted. The mean values of the channel directions of the local feature blocks are statistically analyzed to obtain the texture prior vectors. The vectors are arranged horizontally to obtain the generative texture prior knowledge base. The similarity between the local microstructure fingerprint vector and the texture prior vector is calculated using the cosine similarity formula. The similarity is then normalized to obtain a normalized weight. This weight is then multiplied by the local microstructure fingerprint vector and summed to obtain the texture compensation vector.
[0032] By constructing local microstructure fingerprints and introducing generative texture priors, this step provides an explicit texture compensation mechanism. Local microstructure fingerprints can stably describe the statistical features within a region, while texture priors are derived from the real structural distribution of high-gradient regions. The similarity measure between the two provides a clear basis for the compensation process rather than random enhancement. This approach effectively avoids the problem of artifacts easily introduced by traditional texture enhancement methods, ensuring that texture compensation only applies to structurally reliable regions, thereby improving clarity while maintaining the consistency of the overall texture style.
[0033] Furthermore, the compensation features are obtained, and combined with the dynamic range to obtain the reconstructed image, including: The result of multiplying the texture compensation vector by a scalar weight (set based on historical experimental experience) is added to the highest resolution decoded feature map (the feature vector corresponding to the pixel position) to obtain the compensation feature. Calculate the mean of the compensation feature, and subtract the mean from the compensation feature to obtain the stable compensation feature; The stable compensation features are divided into a first feature subset and a second feature subset along the channel dimension; The first feature subset is entered into the smooth baseline path, and local feature preprocessing and bilinear interpolation upsampling are performed sequentially to obtain smooth structural components. For the second feature subset, a sharp detail path is entered. Spatial resolution is enhanced through channel augmentation mapping combined with pixel shaving. In implementation, progressive zero-initialization is used for the high-frequency detail convolutional layers that generate sharp components to ensure training stability. The second feature subset undergoes spatial resolution enhancement through channel augmentation mapping combined with pixel shaving to obtain rearranged sharp components; the first feature subset is upsampled using bilinear interpolation to obtain constrained smooth components. The rearranged sharp components and the constrained smooth components are weighted and summed to obtain the fused feature map; The product of the fused feature map and the dynamic range is calculated to obtain the reconstructed image.
[0034] By stabilizing the compensation features and combining frequency decomposition and differential upsampling strategies, this step achieves an adaptive balance between sharpness and smoothness in the reconstructed image. The sharp and smooth components are upsampled using different methods, which allows edge details to be fully magnified while the background area remains continuous and natural. Finally, through dynamic range recovery, the reconstruction results in the feature space are accurately mapped back to the image space, thereby obtaining a reconstructed image that is significantly better than existing methods in terms of visual quality and numerical consistency. This demonstrates the comprehensive technical advantages of this invention in complex image reconstruction tasks.
[0035] In summary, this invention constructs a highest-resolution decoding feature map from stable output features, extracts texture compensation vectors, and performs an addition operation between the texture compensation vectors and the highest-resolution decoding feature map to obtain compensation features. Combined with dynamic range, the reconstructed image is obtained; thus, it achieves synergistic enhancement of structure and detail, improves the stability of multi-scale features, and enhances the clarity of the reconstructed image.
[0036] Experimental results show that the method of this invention achieves a PSNR of 35.85 dB (corresponding to MSE = 2.6 × 10⁻⁶) on the large-scale refined image restoration benchmark dataset LSDIR. -4On the CLEVR dataset, the PSNR reached 39.21 dB (corresponding to MSE=1.2×10). -4 The reconstruction accuracy is excellent.
[0037] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A visual backbone network and image processing method based on frequency domain decoupling and microstructure compensation, characterized in that: include, The input image is collected and organized into a three-dimensional tensor. A pixel coordinate system is established, and the dynamic range is set based on the bit depth of the input image. The three-dimensional tensor is normalized to obtain a normalized pixel tensor. The periodic rearrangement ratio is set based on empirical rules, and the periodic rearrangement operation is performed in combination with the normalized pixel tensor to obtain the output tensor. A three-dimensional source index is established for the output tensor, and a lossless injection tensor is constructed. The lossless injection tensor is used as the input of the encoder layer 0. The structural path features and high-frequency residual tensors are calculated and fused to obtain multi-scale features. The fused features at each scale are spatially flattened to construct a token matrix. The query matrix and key matrix are obtained through linear projection and rotational position encoding. The gating weights are calculated based on the multi-scale features. The geometric consistency strength is calculated based on the query matrix and key matrix. The geometric consistency strength is fused with the gating weights to obtain stable output features. The highest resolution decoding feature map is constructed from the stable output features. The texture compensation vector is extracted, and the texture compensation vector is added to the highest resolution decoding feature map to obtain the compensation features. The reconstructed image is obtained by combining the dynamic range.
2. The visual backbone network and image processing method based on frequency domain decoupling and microstructure compensation as described in claim 1, characterized in that: The normalization process of the three-dimensional tensor to obtain a normalized pixel tensor includes: The input image is composed into a three-dimensional tensor. The dynamic range is calculated based on the bit depth of the input image. The three-dimensional tensor is divided by the dynamic range to obtain a normalized pixel tensor.
3. The visual backbone network and image processing method based on frequency domain decoupling and microstructure compensation as described in claim 2, characterized in that: The process of establishing a three-dimensional source index for the output tensor and constructing a lossless injection tensor includes: The periodic rearrangement ratio is set based on empirical rules. The square of the periodic rearrangement ratio is calculated, and the square is multiplied by the number of channels of the normalized pixel tensor to obtain the number of rearrangement channels. The output tensor is obtained by combining the rearrangement size. Based on the channel index expansion relationship defined in the input image, a three-dimensional source index is obtained; A lossless injection tensor is constructed based on the 3D source index and the output tensor.
4. The visual backbone network and image processing method based on frequency domain decoupling and microstructure compensation as described in claim 3, characterized in that: The computational structure path features and high-frequency residual tensors are fused to obtain multi-scale features, including: Define the lossless injection tensor as the input of the 0th layer of the encoder, and set the number of encoder layers to 1. ,right The input of the layer is physically folded to obtain Layer folding tensor; The folded tensor is layer normalized to obtain a normalized tensor. Based on the normalized tensor, the structural path features are obtained through linear projection. High-frequency residual tensors are extracted from folded tensors using local convolution residuals; The high-frequency residual tensor and structural path features are superimposed element by element to obtain... Layer fusion characteristics; Will The layer fusion feature is set as The input of the layer is used to repeat the above operation, traversing... Each layer, obtained The fused features of each layer are stacked horizontally to generate multi-scale features.
5. The visual backbone network and image processing method based on frequency domain decoupling and microstructure compensation as described in claim 4, characterized in that: The process of fusing geometric consistency strength with gating weights to obtain stable output features includes: Spatially flatten the fused features at each scale in the multi-scale features to obtain a token matrix, construct a two-dimensional position index of the token matrix, define the angular frequency, construct a rotation angle for the two-dimensional position index of the token matrix, and obtain the query matrix and key matrix after injecting rotation position encoding; Spatial mean is calculated for the fused features at each scale in the multi-scale feature set to obtain the channel summary vector. The normalized cosine similarity matrix is calculated for the channel summary vector to obtain the channel autocorrelation matrix. The elements in the channel autocorrelation matrix are summed and averaged, and the resulting mean is input into the Sigmoid function to generate the gating weights. The spatial geometric correlation matrix is calculated based on the query matrix and key matrix after the injection rotation position encoding. The average value of the spatial geometric correlation matrix is obtained by calculating the geometric consistency strength by taking the row direction. The spatial gating weight is obtained by multiplying the geometric consistency strength and the gating weight. The product of the spatial gating weights and the fused features is calculated to obtain the enhanced features. The sum of the enhanced features and the fused features is then calculated to obtain the stable output features.
6. The visual backbone network and image processing method based on frequency domain decoupling and microstructure compensation as described in claim 5, characterized in that: The construction of the highest resolution decoding feature map from stable output features includes: Extract encoder The stable output characteristics of the layer are set as the initial value for decoding. Linear projection moments are applied to the initial value for decoding to generate the channel alignment matrix. Set a folding factor, apply the inverse folding operator to the channel alignment matrix to obtain upsampled features, and map them using the standard inverse PixelShuffle formula to obtain the mapped values; Extract encoder layer The stable output features of the layer are set as jump connection features. Linear projection is used on the jump connection features to generate a jump connection feature tensor. The jump-connected feature tensor and the upsampled feature are added together to obtain intermediate features, which are then normalized to obtain normalized intermediate features. Perform a 3×3 two-dimensional convolution on the normalized intermediate features to obtain a refined residual. Then, perform an addition operation on the refined residual and the normalized intermediate features to obtain the highest resolution decoded feature map.
7. The visual backbone network and image processing method based on frequency domain decoupling and microstructure compensation as described in claim 6, characterized in that: The extraction of the texture compensation vector includes: Local feature blocks are constructed using local receptive field modeling, and spatial average pooling is performed on the local feature blocks to obtain local microstructure fingerprint vectors. The gradient magnitude is calculated for the highest resolution decoded feature map, the positions with gradient magnitude greater than the gradient magnitude threshold are selected, the local feature blocks corresponding to the positions are extracted, and the channel direction mean is statistically analyzed for the local feature blocks to obtain the texture prior vector. The vectors are arranged horizontally to obtain the generative texture prior knowledge base. The similarity between the local microstructure fingerprint vector and the texture prior vector is calculated using the cosine similarity formula. The similarity is then normalized to obtain a normalized weight. This weight is then multiplied by the local microstructure fingerprint vector and summed to obtain the texture compensation vector.
8. The visual backbone network and image processing method based on frequency domain decoupling and microstructure compensation as described in claim 7, characterized in that: The obtained compensation features, combined with the dynamic range, yield a reconstructed image, including: The result of multiplying the texture compensation vector by the scalar weight is added to the highest resolution decoded feature map to obtain the compensation feature; Calculate the mean of the compensation feature, and subtract the mean from the compensation feature to obtain the stable compensation feature; The stable compensation features are decomposed using a frequency decomposition method based on channel gradient energy statistics to obtain sharp component channels and smooth component channels. Subpixel rearrangement upsampling is applied to the sharp component channel to obtain the rearranged sharp component, and continuous constraint interpolation upsampling is applied to the smooth component channel to obtain the constrained smooth component; the rearranged sharp component and the constrained smooth component are weighted and summed to obtain the fused feature map. The product of the fused feature map and the dynamic range is calculated to obtain the reconstructed image.