An image segmentation interpretation method and system based on hierarchical high-frequency padding

CN122821132APending Publication Date: 2026-09-25BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611028108.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-10
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

本发明针对现有视觉Transformer分割方法由单尺度特征构建的特征金字塔因简单重采样造成的细粒度层级高频缺失、方向边界响应失真以及纹理伪影和语义无关高频干扰难以区分的问题,通过层级高频构造、方向动态滤波和下一粗层级语义调制的协同处理,实现对边界、细小目标和局部结构相关高频信息的有效补充与选择性抑制,从而提升图像分割解译结果的边界清晰度、区域一致性和整体鲁棒性

Benefits of technology

[0078]本发明,对比现有技术,具有以下优点:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821132A_ABST
    Figure CN122821132A_ABST
Patent Text Reader

Abstract

The application discloses a hierarchical high-frequency filling-based image segmentation and interpretation method and system, which comprises the following steps: carrying out block embedding on an input image and extracting multi-level ViT deep features; carrying out Haar wavelet decomposition on the features of each level, combining image anchoring direction filtering and next-level semantic modulation to implement hierarchical high-frequency filling, and constructing a four-level alignment feature pyramid; carrying out multi-scale fusion, feature integration and pixel-level classification on the filled features to obtain a segmentation and interpretation result. The system comprises an image block embedding subsystem, a multi-level ViT feature extraction subsystem, a hierarchical high-frequency filling expansion subsystem, a multi-scale pyramid fusion subsystem and a classification output subsystem. The application can inhibit texture artifacts and semantic irrelevant high frequencies, relieve high-frequency loss and frequency role distortion in a feature pyramid, and improve segmentation boundary definition, regional consistency and overall accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an image segmentation and interpretation method and system, specifically to an image segmentation and interpretation method and system based on hierarchical high-frequency filling, belonging to the field of computer image processing technology. Background Technology

[0002] Image segmentation is a core task in computer vision, aiming to assign each pixel in an image to a specific semantic category. It has wide applications in fields such as autonomous driving, medical imaging, remote sensing interpretation, industrial quality inspection, and scene understanding. This task not only requires models to identify the category of objects in an image, but also to recover object boundaries, fine structures, and internal consistency of regions at the pixel level. Therefore, it places high demands on the ability to represent multi-scale features and preserve boundary details.

[0003] Visual Transformers, capable of modeling long-range dependencies through self-attention mechanisms and obtaining strong global contextual representations, have become crucial backbone networks for dense prediction tasks such as image segmentation. While standard visual Transformer models are structurally simple and possess strong global modeling capabilities, the original network typically outputs only single-scale block-level features, lacking the multi-level spatial hierarchy inherent in convolutional neural networks. To adapt to segmentation decoders, existing methods often construct a feature pyramid after the backbone network, transforming Transformer features into multi-level resolution features.

[0004] Existing pyramid construction methods mostly rely on resampling operations such as transposed convolution, interpolated upsampling, or pooling downsampling. While these operations can change the feature space size, from a frequency domain perspective, their main function is to copy, rearrange, or filter existing spectra, making it difficult to independently generate high-frequency evidence consistent with target boundaries, small target structures, and semantic regions. Although the resulting fine-grained layers offer higher resolution, they may lack detailed responses aligned with semantic boundaries, and adjacent layers can easily exhibit mutually scaled spectral variations, making it difficult to establish a clear frequency division from coarse to fine.

[0005] Existing frequency-aware segmentation methods and anti-oversmoothing methods recognize the importance of high-frequency responses for boundary preservation and attempt to enhance detail representation through frequency enhancement, local filtering, or feature reconstruction. However, these methods often process high-frequency information based on local texture or response intensity, making it difficult to determine which directional high frequencies should appear at the current pyramid level and which textures or resampling artifacts should be suppressed. Therefore, how to construct structurally consistent directional high frequencies after size expansion and utilize a coarser semantic context for selection remains an unresolved problem in conventional visual Transformer segmentation frameworks.

[0006] In summary, developing an image segmentation and interpretation method capable of performing hierarchical high-frequency filling, constructing boundary-related directional details, and suppressing texture artifacts within a standard visual Transformer image segmentation framework is of great value for improving the multi-scale expressive power and segmentation accuracy of standard visual Transformers in dense prediction tasks. It is also a key issue that urgently needs to be addressed in the current development of image segmentation technology. Summary of the Invention

[0007] The purpose of this invention is to address the shortcomings and deficiencies of existing technologies by creatively proposing an image segmentation and interpretation method based on hierarchical high-frequency filling. This invention addresses the problems of fine-grained hierarchical high-frequency loss, directional boundary response distortion, and difficulty in distinguishing texture artifacts and semantically irrelevant high-frequency interference caused by simple resampling in existing visual Transformer segmentation methods, which construct feature pyramids from single-scale features. Through the collaborative processing of hierarchical high-frequency construction, directional dynamic filtering, and the next coarse-level semantic modulation, this invention effectively supplements and selectively suppresses high-frequency information related to boundaries, small targets, and local structures, thereby improving the boundary clarity, regional consistency, and overall robustness of the image segmentation and interpretation results.

[0008] The overall process of this invention is as follows: Figure 1 As shown, the implementation uses the standard visual Transformer as the global feature extraction backbone. First, the input image is divided into image blocks and encoded as token features, and multi-level deep features are extracted from multiple encoder layers. Then, the multi-level deep features are scaled to form a four-level aligned feature pyramid. Haar wavelet decomposition is performed on each level of aligned features to obtain low-frequency components and multiple directional high-frequency components. The directional filtering module is used to predict the spatial dynamic convolution kernel based on the current directional high frequency, directional prior, the next coarse-level semantic conditions, and optional image-side evidence to construct candidate directional high frequencies. Then, the next-level modulation module generates directional spatial gating to perform semantic filtering and redundancy suppression on the candidate high frequencies. Finally, the final segmentation and interpretation results are output through inverse wavelet reconstruction, residual fusion, multi-scale pyramid fusion, and pixel-level classification.

[0009] To achieve the above objectives, the present invention employs the following technical solutions.

[0010] A hierarchical high-frequency filling-based image segmentation and interpretation method includes the following steps:

[0011] Step 1: Perform block embedding on the input image.

[0012] Specifically, the input image is divided into non-overlapping image blocks. A convolutional layer with a kernel size of 16×16 and a stride of 16 is used to linearly project these blocks, resulting in a block embedding feature map. Positional encoding is then added to preserve the spatial location information of the image blocks. The final result is a block embedding feature map with dimensions one-sixteenth the size of the original image.

[0013] Step 2: Extract multi-level deep features from the block embedding feature map.

[0014] Specifically, a standard visual Transformer encoder is used to process the block embedding feature maps. Each encoding layer includes layer normalization, a multi-head self-attention module, a multilayer perceptron module, and residual connections. The multi-head self-attention module calculates global relevance through query, key, and value features, while the multilayer perceptron module enhances feature representation through channel expansion, non-linear activation, and channel recovery. The outputs of the encoders at layers 3, 6, 9, and 12 are taken, and the token sequence is restored to a spatial feature map, which serves as the fourth-level deep feature map for subsequent processing.

[0015] Furthermore, the encoder comprises 12 Transformer coding layers, each with an embedding dimension of 768, 12 multi-head self-attention heads, each with a dimension of 64, and a multilayer perceptron hidden layer with a dimension of 3072. Each coding layer employs a pre-normalized structure, and its attention and feedforward computations can be expressed as follows:

[0016]

[0017]

[0018]

[0019]

[0020]

[0021]

[0022] in, This refers to the coding layer number; , For the first Layer input and output tokens; , , and These are layer normalization, activation, normalized weights, and multi-head splicing, respectively. , and and , , The whole and the first The query, key, and value of each size. The first serial number; , , , , and It is a linear mapping matrix; This is the scaling factor; , These are the intermediate features of multi-head self-attention output and attention residual, respectively.

[0023] Furthermore, the outputs of layers 3, 6, 9, and 12 are taken and rearranged into spatial feature maps, resulting in... , , , With a 512×512 input, the spatial size of the four-level deep features is 32×32, and the number of channels is 768.

[0024] Step 3: Expand the size of multi-level deep features and perform high-frequency filling of the layers.

[0025] Specifically, the four levels of deep features are scaled to form a shared pyramid: the output of the third layer is upsampled four times, the output of the sixth layer is upsampled twice, the output of the ninth layer maintains an identity mapping, and the output of the twelfth layer is downsampled twice, resulting in four levels of aligned features from fine to coarse.

[0026] Specifically, Haar wavelet decomposition is performed on each level of aligned features, decomposing the features into a low-frequency component and three high-frequency components in the LH, HL, and HH directions. For each high-frequency sub-band in a direction, a kernel conditional feature is constructed, consisting of the current high-frequency component, the next coarse-level semantic condition, and image-side evidence. Then, the directional filtering module predicts the spatial dynamic convolution kernel, and performs position-by-position local recombination on the high-frequency sub-band in the direction to form candidate high-frequency responses in the direction.

[0027] Specifically, the next-layer modulation module receives the refined features from the next coarse-level layer, and generates spatial gating in three directions (LH, HL, and HH) through convolution and Tanh activation. It retains semantically supported high-frequency components while attenuating texture-driven, artifact-driven, or semantically irrelevant redundant high-frequency components. Then, it performs inverse Haar wavelet reconstruction by combining the modulated high-frequency components in the three directions with the projected low-frequency components, and adds this to the original aligned features through learnable residual scaling to obtain the refined features for the current level.

[0028] Specifically, the high-frequency filling of the hierarchy is performed in order from coarse to fine. The coarsest level does not have a semantic condition for the next coarse level, so the modulation gate of its next level takes the identity gate; the remaining levels all use the refined features of the next coarse level as semantic conditions.

[0029] Furthermore, the fourth-level deep features are transformed into fourth-level aligned features from fine to coarse. ,in These correspond to 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the input image size, respectively.

[0030]

[0031]

[0032]

[0033]

[0034] Where Deconv2 represents A transposed convolution with a stride of 2, represented by MaxPool2. Max pooling, express Convolution channel projection, with a uniform channel number of 256.

[0035] like Figure 3 As shown in (a), Haar wavelet decomposition is performed on each level of alignment feature to obtain one low-frequency component and three high-frequency components in three directions:

[0036]

[0037] in, The hierarchical index representing the alignment feature, with values ​​of 1, 2, 3, and 4; Indicates the first Level alignment features, Represents the Haar wavelet decomposition operator; Indicates the first Low-frequency components are used to preserve key semantic and contour information; , and They represent the first The high-frequency components in the LH, HL, and HH directions are used to characterize edge, texture, and detail responses in different directions.

[0038] like Figure 3 As shown in (b), the hierarchical high-frequency filling is in accordance with to Proceed from coarse to fine. If Refined features from the next coarse level Generate semantic conditions ;like ,but For each directional sub-band Constructing nuclear conditional features ,in Optional image-side evidence.

[0039] like Figure 3 As shown in (c), the directional filtering module is based on the kernel condition characteristics. Predicting spatial dynamic convolution kernels:

[0040]

[0041]

[0042]

[0043] in, Represents a low-pass dynamic convolution kernel. Represents Qualcomm dynamic convolution kernels. Represents a hybrid convolution kernel; As a direction prior, Norm0 denotes zero-mean normalization. and There are two learnable convolutional layers. To balance the weights, specifically, in this embodiment, the spatial dynamic convolution kernel size is 3×3, and the orientation prior includes horizontal, vertical, and diagonal templates.

[0044] Furthermore, spatial dynamic convolution uses a corresponding 3×3 kernel to reconstruct high-frequency directions at each spatial location, and obtains candidate high-frequency directions using residuals:

[0045]

[0046] like Figure 3 As shown in (d), the next-level modulation module uses the refined features of the next coarse-level layer to generate directional spatial gating and obtain the modulated directional high frequency. :

[0047]

[0048]

[0049] Among them, when hour, An all-1 identity gating method is used; when the gating sizes are inconsistent, alignment is achieved through bilinear interpolation. Specifically, the low-frequency components and the modulated directional high-frequency components are reconstructed using inverse Haar wavelets, and the refined features of the current level are obtained through residual connections.

[0050]

[0051] in, Indicates inverse Haar wavelet reconstruction, The learnable residual scaling coefficients are used. Through this residual structure, the module preserves the original aligned feature semantics while supplementing only the high-frequency details after directional filtering and semantic gating. When a region lacks coarse-level semantic support, the corresponding high-frequency frequencies are attenuated through gating, thereby reducing the interference of texture noise and interpolation artifacts on the segmentation boundaries. This leads to... , , and Fourth-level high-frequency filling feature.

[0052] Step 4: Perform multi-scale pyramid fusion on the features after high-frequency filling at the hierarchical level.

[0053] Specifically, multi-scale pyramid fusion includes pyramid pooling, lateral connections, top-down fusion, and feature refinement. Pyramid pooling takes the coarsest-level features as input and obtains global and local context through multiple average pooling branches at different scales; lateral connections use 1×1 convolutions to unify the number of channels; top-down fusion starts from the coarsest level, upsamples the current level features to the size of adjacent finer levels, and adds them to its lateral features; finally, 3×3 convolutions are used to refine the fused features of each level.

[0054] Furthermore, the fourth-level features after high-frequency filling are analyzed. , , , Multi-scale feature pyramid fusion is performed. This step includes three parts: pyramid pooling, top-down fusion, and feature refinement. The deepest feature, to It represents a higher resolution hierarchical feature.

[0055] Furthermore, pyramid pooling... For input, use , , and Four average pooling branches are used to extract multi-scale context. Each branch is processed... After convolution and bilinear upsampling, compared with the original splicing, and through Convolution yields the deepest fused feature G_4:

[0056]

[0057]

[0058] in, Indicates the pooling scale as Branch output, This represents the deepest fusion feature after pyramid pooling.

[0059] Furthermore, , , respectively Lateral convolution yields , , and with As the starting point for top-down fusion, it is upsampled step by step and added to the corresponding lateral features, and then... Convolutional refinement yields multi-level pyramid features. , , , :

[0060]

[0061]

[0062] in, Indicates the first Lateral features of the layer This indicates the fusion and refinement of the first... Layered pyramid features are used for subsequent classification output.

[0063] Step 5: Integrate and classify the multi-level pyramid feature maps and output the results.

[0064] Specifically, all pyramid feature maps are upsampled to the highest resolution level through bilinear interpolation and concatenated along the channel dimension. They are then integrated into a unified segmentation feature through a 3×3 convolutional bottleneck layer. Subsequently, a 1×1 convolution is used to map the feature maps to the channel space corresponding to the number of categories, resulting in a segmentation prediction heatmap. Finally, the prediction heatmap is upsampled to the input image size and the probability of each pixel category is obtained through Softmax. The resulting image segmentation interpretation is then output.

[0065] Furthermore, , and All samples were upsampled to The dimensions are concatenated along the channel dimension and fused using a 3×3 convolutional bottleneck layer. Then, a 1×1 convolution is used to obtain a category prediction heatmap, which is upsampled to the input image size.

[0066]

[0067]

[0068]

[0069]

[0070] in, The number of channels equals the number of semantic categories. Represents pixels Belongs to the The probability of a class This is the final segmentation and interpretation diagram.

[0071] Furthermore, to achieve the objectives of this invention, based on the above method, this invention also proposes an image segmentation and interpretation system based on hierarchical high-frequency filling, including an image block embedding module, a multi-level feature extraction module, a hierarchical high-frequency filling module, a multi-scale pyramid fusion module, and a classification output module.

[0072] The image block embedding module is used to divide the input image into image blocks and convert them into block embedding features containing position codes.

[0073] The multi-level feature extraction module is used to obtain multi-level deep features with global context through a visual Transformer encoder;

[0074] The hierarchical high-frequency filling module is used to expand multi-level deep features into four-level pyramid features, and to form hierarchical refined features through wavelet domain directional high-frequency construction and next-level semantic modulation.

[0075] The multi-scale pyramid fusion module is used to perform pyramid pooling, top-down fusion, and local refinement on hierarchical refined features to generate multi-scale fused features.

[0076] The classification output module is used to integrate multi-scale fusion features and perform pixel-level classification, outputting the final image segmentation and interpretation results.

[0077] Beneficial effects

[0078] Compared with the prior art, the present invention has the following advantages:

[0079] 1. This invention compensates for the lack of frequency structure in single-scale Transformer features by using a hierarchical frequency filling paradigm, thereby obtaining a multi-level expression suitable for dense prediction.

[0080] 2. This invention introduces Haar wavelet domain hierarchical high-frequency filling and uses directional dynamic kernels to construct candidate high-frequency components, supplementing frequency evidence related to boundaries and fine structures that cannot be repaired by simple resampling.

[0081] 3. This invention uses the next layer of semantics for spatial selection, which concentrates high-frequency responses on semantic boundaries and effective structural regions, while suppressing texture artifacts and irrelevant edges, thus helping to improve the consistency of segmented regions. Attached Figure Description

[0082] Figure 1 This is a general overview of the process of this invention.

[0083] Figure 2 This is a schematic flowchart of the method of the present invention;

[0084] Figure 3 This is a schematic diagram of the neural network structure used in the method of the present invention;

[0085] Figure 4 This is a schematic diagram of the system composition of the present invention. Detailed Implementation

[0086] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.

[0087] like Figure 2 As shown, an image segmentation and interpretation method based on hierarchical high-frequency filling includes the following steps:

[0088] Step 1: Perform block embedding on the input image.

[0089] Furthermore, the input image is processed using a convolutional kernel with a size of 16×16 and a stride of 16, resulting in a block embedding feature map with dimensions one-sixteenth that of the original image. The input image has 3 channels, and the output block embedding feature map has 768 channels.

[0090] Step 2: Extract multi-level deep features from the block embedding feature map.

[0091] Furthermore, the encoder comprises 12 Transformer coding layers, each with an embedding dimension of 768, 12 multi-head self-attention heads, each attention head with a dimension of 64, and a multilayer perceptron hidden layer with a dimension of 3072. The l-th coding layer adopts a pre-normalized structure, and its attention and feedforward computation relationship is described in step 2 above.

[0092] in, This refers to the coding layer number; , For the first Layer input and output tokens; , , and These are layer normalization, activation, normalized weights, and multi-head splicing, respectively. , and and , , The whole and the first The query, key, and value of each size. The first serial number; , , , , and It is a linear mapping matrix; This is the scaling factor; , These are the intermediate features of multi-head self-attention output and attention residual, respectively.

[0093] Furthermore, the outputs of layers 3, 6, 9, and 12 are taken and rearranged into spatial feature maps, resulting in... , , , With a 512×512 input, the spatial size of the four-level deep features is 32×32, and the number of channels is 768.

[0094] Step 3: Expand the size of multi-level deep features and perform high-frequency filling of the layers.

[0095] Furthermore, the four-level deep features are transformed into four-level aligned features Pi0 from fine to coarse, where i=1,2,3,4 correspond to 1 / 4, 1 / 8, 1 / 16 and 1 / 32 of the input image size, respectively. The specific scale transformation relationship is described in step 3 above.

[0096] Where Deconv2 represents A transposed convolution with a stride of 2, represented by MaxPool2. Max pooling, express Convolution channel projection, with a uniform channel number of 256.

[0097] like Figure 3 As shown in (a), Haar wavelet decomposition is performed on each level of alignment feature to obtain a low-frequency component and three high-frequency components in three directions. The wavelet decomposition relationship is shown in step 3 above.

[0098] in, The hierarchical index representing the alignment feature, with values ​​of 1, 2, 3, and 4; Indicates the first Level alignment features, Represents the Haar wavelet decomposition operator; Indicates the first Low-frequency components are used to preserve key semantic and contour information; , and They represent the first The high-frequency components in the LH, HL, and HH directions are used to characterize edge, texture, and detail responses in different directions.

[0099] like Figure 3 As shown in (b), the high-frequency filling of the hierarchy is carried out in order from coarse to fine. The semantic conditions of the non-coarsest level are generated by refining the features of the next coarsest level, and the semantic conditions of the coarsest level are set to zero. For each directional sub-band, a kernel condition feature is constructed, which consists of the high frequency of the current direction, the semantic conditions, and the optional image side evidence. The construction relationship is shown in step 3 above.

[0100] like Figure 3 As shown in (c), the directional filtering module predicts the spatial dynamic convolution kernel based on the kernel condition feature Uib, and the calculation relationship is shown in step 3 above.

[0101] in, Represents a low-pass dynamic convolution kernel. Represents Qualcomm dynamic convolution kernels. Represents a hybrid convolution kernel; As a direction prior, Norm0 denotes zero-mean normalization. and There are two learnable convolutional layers. To balance the weights, specifically, in this embodiment, the spatial dynamic convolution kernel size is 3×3, and the orientation prior includes horizontal, vertical, and diagonal templates.

[0102] Furthermore, spatial dynamic convolution uses the corresponding 3×3 kernel to reconstruct the high frequency of the direction at each spatial location, and uses the residual form to obtain the candidate high frequency of the direction. The specific calculation relationship is described in step 3 above.

[0103] like Figure 3 As shown in (d), the next-level modulation module uses the refined features of the next coarse level to generate directional spatial gating and obtain the modulated directional high frequency.

[0104] Specifically, when i=4, g4b is gated with all 1s; when the gating sizes are inconsistent, they are aligned by bilinear interpolation. In particular, the low-frequency components and the modulated directional high frequencies are reconstructed using inverse Haar wavelets, and the refined features of the current level are obtained through residual connections. The inverse wavelet reconstruction relationship is described in step 3 above.

[0105] in, Indicates inverse Haar wavelet reconstruction, The learnable residual scaling coefficients are used. Through this residual structure, the module preserves the original aligned feature semantics while supplementing only the high-frequency details after directional filtering and semantic gating. When a region lacks coarse-level semantic support, the corresponding high-frequency frequencies are attenuated through gating, thereby reducing the interference of texture noise and interpolation artifacts on the segmentation boundaries. This leads to... , , and Fourth-level high-frequency filling feature.

[0106] Step 4: Perform multi-scale pyramid fusion on the features after high-frequency filling at the hierarchical level.

[0107] Furthermore, the fourth-level features after high-frequency filling are analyzed. , , , Multi-scale feature pyramid fusion is performed. This step includes three parts: pyramid pooling, top-down fusion, and feature refinement. The deepest feature, to It represents a higher resolution hierarchical feature.

[0108] Furthermore, pyramid pooling... For input, use , , and Four average pooling branches are used to extract multi-scale context. Each branch is processed... After convolution and bilinear upsampling, it is concatenated with the original Y4, and then... Convolution yields the deepest fused features The calculation relationship is shown in step 4 above.

[0109] in, Indicates the pooling scale as Branch output, This represents the deepest fusion feature after pyramid pooling.

[0110] Furthermore, , , respectively Convolution yields , , Using G4 as the starting point for top-down fusion, the features are upsampled at each level and added to the corresponding lateral features, then refined through 3×3 convolution to obtain multi-level pyramid features. , , , The fusion relationship is described in step 4 above.

[0111] in, Indicates the first Lateral features of the layer This indicates the fusion and refinement of the first... Layered pyramid features are used for subsequent classification output.

[0112] Step 5: Integrate and classify the multi-level pyramid feature maps and output the results.

[0113] Furthermore, , and All samples were upsampled to Dimensions, spliced ​​together in the channel dimension and passed through The convolutional bottleneck layers are fused and then utilized. Convolution yields a category prediction heatmap, which is then upsampled to the input image size. The classification output relationship is described in step 5 above.

[0114] in, The number of channels equals the number of semantic categories. Represents pixels Belongs to the The probability of a class This is the final segmentation and interpretation diagram.

[0115] like Figure 4 As shown, an image segmentation and interpretation system based on hierarchical high-frequency filling includes a multi-level feature extraction subsystem M1, a hierarchical high-frequency filling subsystem M2, a multi-scale pyramid fusion subsystem M3, and a classification output subsystem M4.

[0116] The multi-level feature extraction subsystem M1 is used to extract deep features;

[0117] The hierarchical high-frequency fill extension subsystem M2 is used for hierarchical frequency fill reconstruction;

[0118] The multi-scale pyramid fusion subsystem M3 is used for pyramid pooling and top-down fusion.

[0119] The classification output subsystem M4 is used for feature integration and pixel-level classification.

[0120] The connections between the above subsystems are as follows:

[0121] The input is connected to the input terminal of the multi-level ViT feature extraction subsystem M1;

[0122] The output of the multi-level feature extraction subsystem M1 is connected to the input of the hierarchical high-frequency filling subsystem M2;

[0123] The output of the hierarchical high-frequency filling subsystem M2 is connected to the input of the multi-scale pyramid fusion subsystem M3;

[0124] The output of the multi-scale pyramid fusion subsystem M3 is connected to the input of the classification output subsystem M4;

[0125] The output of the classification output subsystem M4 is connected to the output segmentation and interpretation results.

[0126] This invention discloses an image segmentation and interpretation method and system based on hierarchical high-frequency filling. By using hierarchical high-frequency filling, the high-frequency information lost in the ViT network is restored. A directional filtering module is used to fill the high frequencies, and a modulation module at the next level is used to optimize the high-frequency distribution. While enhancing the high-frequency details at the boundaries, the method effectively suppresses noise inside the object, thereby improving the intra-class consistency and inter-class separability of the segmentation results. This achieves high-precision and robust image segmentation and interpretation based on the ViT model.

Claims

1. An image segmentation and interpretation method based on hierarchical high-frequency filling, characterized in that, Includes the following steps: Step 1: Perform block embedding processing on the input image to obtain block embedding features containing spatial location information; Step 2: Input the block embedding features into the visual Transformer encoder to extract multi-level deep features from multiple encoder layers; Step 3: Expand the size of the multi-level deep features to form four-level aligned features, and perform high-frequency filling on each level of aligned features to obtain four-level refined features. Step 4: Perform multi-scale pyramid fusion on the four-level refined features to generate a multi-level pyramid feature map; Step 5: Integrate and classify the multi-level pyramid feature maps at the pixel level, and output the final image segmentation and interpretation results.

2. The method as described in claim 1, characterized in that, The block embedding process in step 1 includes: using a convolutional layer with a kernel size of 16×16 and a stride of 16 to project non-overlapping image blocks onto the input image, and adding positional encoding to obtain block embedding features whose length and width are both one-sixteenth of the input image.

3. The method as described in claim 1, characterized in that, The visual Transformer encoder in step 2 includes multiple Transformer coding layers. Each coding layer includes layer normalization, a multi-head self-attention module, a multilayer perceptron module, and residual connections. Features are extracted from the outputs of layers 3, 6, 9, and 12, and the token sequence is restored to a spatial feature map to obtain four levels of deep features.

4. The method as described in claim 1, characterized in that, The size expansion in step 3 includes: quadrupling upsampling of the output of layer 3, upsampling of the output of layer 6 by two times, maintaining the identity mapping of the output of layer 9, downsampling of the output of layer 12, and forming a four-level alignment feature from fine to coarse through channel projection.

5. The method as described in claim 1, characterized in that, The hierarchical high-frequency filling in step 3 includes: performing Haar wavelet decomposition on the i-th level alignment feature to obtain low-frequency components and high-frequency components in three directions: in, , Indicates the first Level alignment features, This represents the Haar wavelet decomposition operator. Indicates the first Low-frequency components, , and These represent the high-frequency components in the LH, HL, and HH directions, respectively.

6. The method as described in claim 5, characterized in that, The hierarchical high-frequency filling is performed in order from coarse to fine. For each directional high-frequency component, a kernel condition feature is constructed that includes the current directional high frequency, the next coarse-level semantic condition, and optional image-side evidence. The spatial dynamic convolution kernel is predicted by the directional filtering module to reassemble the directional high frequency position by position and generate candidate directional high-frequency responses.

7. The method as described in claim 6, characterized in that, The hierarchical high-frequency filling also includes a next-level modulation process; the next-level modulation process uses the refined next-level coarse-level features to generate spatial gating in three directions: LH, HL, and HH, performs semantic filtering on the candidate direction high frequencies, and reconstructs the filtered direction high frequencies and low-frequency components using inverse Haar wavelets, obtaining the current level refined features through residual connections: in, Indicates the first Level-by-level refining features , and This represents the directional high-frequency component after semantically gated modulation. This represents the inverse Haar wavelet reconstruction operator. This represents the learnable residual scaling factor.

8. The method as described in claim 1, characterized in that, Step 4, multi-scale pyramid fusion, includes pyramid pooling, top-down fusion, and feature refinement. Pyramid pooling takes the deepest level of refined features as input and then... , , and Average pooling branch extracts multi-scale context; top-down fusion upsamples coarse-level features and adds them to corresponding fine-level lateral features; feature refinement is performed through... Convolution generates multi-level pyramid feature maps.

9. The method as described in claim 1, characterized in that, The integration and classification in step 5 includes: upsampling the multi-level pyramid feature map to a uniform spatial size and stitching it together in the channel dimension, integrating it into segmentation features through a convolutional layer, and then mapping it to the category channel through a 1×1 convolution to obtain a segmentation prediction heatmap. The segmentation prediction heatmap is then upsampled to the input image size and the pixel-level segmentation interpretation result is output.

10. An image segmentation and interpretation system based on hierarchical high-frequency filling, characterized in that, The method includes an image patch embedding module, a multi-level feature extraction module, a hierarchical high-frequency filling module, a multi-scale pyramid fusion module, and a classification output module; the image patch embedding module is used to perform step 1 of the method according to any one of claims 1 to 9; the multi-level feature extraction module is used to perform step 2 of the method; the hierarchical high-frequency filling module is used to perform step 3 of the method; the multi-scale pyramid fusion module is used to perform step 4 of the method; and the classification output module is used to perform step 5 of the method.