An image segmentation method and system

By constructing a spatial-frequency domain collaborative boundary distribution network and utilizing a multi-scale boundary distribution generation module and a state-space semantic enhancement module, the problem of boundary segmentation instability under complex imaging conditions is solved, achieving stable, coherent, and accurate target boundary segmentation, thus improving the accuracy and consistency of image segmentation.

CN122090066BActive Publication Date: 2026-06-26CHONGQING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHONGQING UNIV
Filing Date
2026-04-23
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing image segmentation methods struggle to achieve stable, coherent, and accurate target boundary segmentation under complex imaging conditions, especially in scenarios with low contrast, weak boundaries, diverse shapes, and strong interference. Boundary breaks and incomplete contours result in insufficient utilization of frequency domain information, and limited receptive field and global semantic consistency.

Method used

A spatial-frequency domain collaborative boundary distribution network is constructed. The boundary distribution map is predicted by a multi-scale boundary distribution generation module and joint supervision in the spatial and frequency domains is applied. Combined with a state-space semantic enhancement module and a boundary distribution guidance decoding module, the boundary distribution map is approximated in terms of spatial location and spectral structure, local high-frequency interference is suppressed, global semantic consistency is enhanced, and the boundary-guided cross-scale decoder strengthens the boundary-related response.

Benefits of technology

Under complex imaging conditions such as low contrast, diverse shapes, and interference from reflections and folds, accurate and complete polyp segmentation results were achieved while maintaining reasonable computational overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122090066B_ABST
    Figure CN122090066B_ABST
Patent Text Reader

Abstract

The application relates to the field of computer vision, and particularly discloses an image segmentation method and system, wherein a boundary distribution map is predicted by a multi-scale boundary distribution generation module as prior knowledge, a space-frequency domain joint supervision strategy is adopted, the boundary distribution map simultaneously approximates a real boundary in a spatial position and a spectral structure, false edge responses caused by local high-frequency interference are effectively inhibited, and the continuity and stability of the boundary are improved; an effective receptive field is expanded by a state space semantic enhancement module, and global semantic consistency is enhanced; a boundary-guided cross-scale decoder is combined with a boundary feature enhancement module, boundary-related responses are strengthened in multi-resolution fusion, and collaborative optimization of region positioning and boundary refinement is realized, so that accurate and complete polyp segmentation results are obtained under complex imaging conditions such as low contrast, various morphologies, and interference such as reflection and wrinkles, and reasonable calculation overhead is maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision, and more particularly to an image segmentation method and system. Background Technology

[0002] Image segmentation is one of the core tasks in computer vision, aiming to divide an image into multiple regions with specific semantic meanings. Its accuracy directly affects subsequent advanced vision tasks, such as object recognition and scene understanding. In many applications such as medical image analysis, autonomous driving, and remote sensing image interpretation, accurate target localization and boundary delineation are crucial.

[0003] Taking colonoscopy image analysis in medical imaging as an example, polyp segmentation is a crucial step in the early screening and diagnosis of colorectal cancer. However, polyps in real colonoscopy images often exhibit low contrast, diverse shapes, and blurred boundaries, accompanied by strong interfering factors such as reflections, folds, mucus, and bubbles, making polyp segmentation a significant challenge. This scenario vividly illustrates the typical difficulties of image segmentation under complex imaging conditions.

[0004] In existing technologies, deep learning-based image segmentation methods mainly employ encoder-decoder structures. For example, U-Net and its variants (such as UNet++ and ResUNet) improve segmentation performance to some extent by fusing high-level semantics with low-level details through skip connections. To address the issues of target scale variations and background similarity, subsequent methods have introduced multi-scale context modules, attention mechanisms, and Transformer structures. For instance, some methods utilize parallel partial decoders and inverse attention to progressively correct segmented regions; others enhance cross-domain robustness through pyramidal Transformers; and still others combine hybrid attention with Transformer structures to focus on both local texture and global morphology. Furthermore, state-space models (SSMs), due to their potential in long-range dependency modeling, have also been explored for segmentation tasks, such as VM-UNet.

[0005] However, the aforementioned existing technologies still generally suffer from the following shortcomings when dealing with scenarios involving low contrast, weak boundaries, diverse shapes, and strong interference (such as noise, bright reflections, and complex textures). These problems are particularly pronounced when using colonoscopy for polyp segmentation:

[0006] 1. Boundary breaks and incomplete contours: Existing methods typically treat boundaries as binary contours or auxiliary branches, applying supervision only in the spatial domain and lacking modeling of boundary signal stability. Under local high-frequency interference such as reflections and wrinkled textures, skip connections in the decoding process are prone to injecting noise, leading to spurious responses near the boundaries, resulting in contour breaks, outward expansion, or false detections, making it difficult to output coherent and accurate target contours;

[0007] 2. Insufficient utilization of frequency domain information: Most existing methods only perform boundary supervision in the spatial domain, without considering the stability and consistency of the boundary structure in the frequency domain. The true polyp boundary corresponds to the stable high-frequency structure of the image, but existing techniques fail to effectively utilize frequency domain constraints to suppress abnormal high-frequency peaks caused by local noise, resulting in inaccurate boundary localization under weak contrast or strong interference conditions, and the segmentation results lack structural integrity;

[0008] 3. Limited receptive field and global semantic consistency: Traditional CNNs have a slow effective receptive field growth, making it difficult to explicitly capture large-scale structural relationships and long-range dependencies. They fall short in handling large targets or scenes requiring high contextual information. Although Transformers can improve global modeling, their computational complexity is high, and they remain susceptible to interference from local pseudo-edge responses. In colonoscopy images, this manifests as the network's difficulty in distinguishing polyps from similar background textures, or the appearance of inconsistent or discontinuous predictions near boundaries. Summary of the Invention

[0009] This invention provides an image segmentation method and system, which solves the technical problem of how to obtain stable, coherent and accurate target boundary segmentation under complex imaging conditions, while maintaining reasonable computational overhead.

[0010] To address the above technical problems, this invention provides an image segmentation method, comprising: constructing a spatial-frequency domain cooperative boundary distribution network, wherein the spatial-frequency domain cooperative boundary distribution network includes an encoding module, a multi-scale boundary distribution generation module, and a boundary distribution guidance decoding module.

[0011] The encoding module uses an encoder to extract multi-scale features from the input image;

[0012] The multi-scale boundary distribution generation module predicts a continuous boundary distribution map from the partial-scale features output by the encoding module, and applies joint spatial and frequency domain supervision to the boundary distribution map;

[0013] The boundary distribution guidance decoding module is used to enhance the multi-scale features through the state space semantic enhancement module to obtain jump connection features, and to perform multi-level decoding with the boundary distribution map as a priori to output a polyp segmentation mask.

[0014] Furthermore, the joint supervision in the spatial and frequency domains is as follows: a combination of spatial regression loss and frequency domain consistency loss is applied to the boundary distribution map; the spatial regression loss is the average value of the pixel-by-pixel mean square error between the predicted boundary distribution map and the ideal boundary distribution map; the frequency domain consistency loss is the difference in amplitude spectrum between the predicted boundary distribution map and the ideal boundary distribution map after two-dimensional fast Fourier transform.

[0015] Furthermore, the multi-scale boundary distribution generation module includes:

[0016] A multi-branch, multi-scale receptive block is used to capture the boundary context of different receptive fields for the partial scale features output by the encoding module, thereby obtaining multi-scale receptive features.

[0017] The boundary aggregation module is used to perform multi-round interactive fusion of the multi-scale receptive features to obtain a boundary feature map.

[0018] The boundary quality regression module is used to perform boundary confidence regression and geometric recalibration on the boundary feature map and the shallowest features in the multi-scale receptive features to obtain the boundary quality regression feature map.

[0019] The upsampling module is used to upsample the boundary quality regression feature map to the original image size to obtain the boundary distribution map.

[0020] Furthermore, the boundary aggregation module upsamples the deep features in the multi-scale receptive features and multiplies them element-wise with the shallow features to construct boundary feature maps of different information granularities at each level. The resulting boundary feature maps are then concatenated, compressed by convolution, and processed by the convolutional block attention module to output the boundary feature map.

[0021] The boundary quality regression module obtains a boundary confidence map by passing the boundary feature map through a Sigmoid function. This boundary confidence map is then used to generate multiple feature maps with different focuses through multiple branches. Simultaneously, the shallowest feature in the multi-scale receptive features is concatenated with the compressed features output by the boundary aggregation module and then convolved to obtain a concatenated feature map. This concatenated feature map is multiplied by multiple feature maps with different focuses, and the multiplication results are concatenated and then convolved. Finally, it is added to the boundary feature map to obtain the boundary quality regression feature map.

[0022] Furthermore, the boundary distribution guidance decoding module includes a state space semantic enhancement module, a multi-level decoding part, and a segmentation part;

[0023] The state space semantic enhancement module is used to perform state space modeling on the multi-scale features output by the encoding module along the row and column directions to obtain enhanced skip connection features;

[0024] The multi-level decoding part is used to recover the resolution step by step by using the boundary distribution map as a priori and combining the enhanced skip connection features to obtain multi-level decoding features;

[0025] The segmentation part is used to obtain a polyp segmentation mask by interacting with the low-resolution semantic branch and the high-resolution boundary detail branch and residual fusion of the multi-level decoded features.

[0026] Furthermore, the processing flow of the state space semantic enhancement module is mathematically quantified as follows:

[0027] ,

[0028] in, , These represent the input and output characteristics of the module, respectively. express Convolution operation, The operation representing the coordinate attention branch of the state-space model. This represents the operation of concatenating multiple feature tensors along the channel dimension;

[0029] The operation flow of the coordinate attention branch in the state-space model is as follows:

[0030] The input features are sequentially scanned along the row and column directions using a state-space model to obtain one-dimensional row description vectors and one-dimensional column description vectors.

[0031] The one-dimensional row description vector and the one-dimensional column description vector are concatenated in the spatial dimension, and then the intermediate representation is obtained by convolution dimensionality reduction.

[0032] The intermediate representation is split into two components along the row and column directions. Each component is then subjected to a 1×1 convolution and a sigmoid activation to obtain the corresponding row-direction attention map and column-direction attention map.

[0033] The input features are recalibrated point by point using the row direction attention map and the column direction attention map.

[0034] Furthermore, the multi-level decoding section employs a multi-level boundary-guided cross-scale decoder for step-by-step decoding, and the decoding process is represented as follows:

[0035] ,

[0036] ,

[0037] ,

[0038] in, This is a low-resolution output characteristic. Features for high-resolution output The current layer decodes the output features. These are the output features of the previous layer's decoding. For the feature enhancement operation of the boundary feature enhancement module, For the enhanced skip connection features, For the boundary distribution map, For average pooling operation, This indicates element-wise multiplication. Indicates an upsampling operation. This indicates a convolution operation.

[0039] Furthermore, the operation of the boundary feature enhancement module is as follows:

[0040] The features obtained by convolving the enhanced skip connection features are multiplied with the features obtained by upsampling the boundary distribution map. The result of the multiplication is then used to extract attention through the channel attention module and the spatial attention module. The extracted attention is then residually connected with the features obtained by convolving the enhanced skip connection features to obtain the output features.

[0041] Furthermore, the overall loss function of the space-frequency domain cooperative boundary distribution network includes segmentation loss and boundary distribution loss; the boundary distribution loss includes spatial domain loss and frequency domain loss.

[0042] The spatial domain loss calculates the mean square error of the predicted boundary distribution map and the ideal boundary distribution map pixel by pixel, and calculates the average error only for pixels whose error is greater than a set threshold.

[0043] The frequency domain loss is used to perform two-dimensional fast Fourier transforms on the predicted boundary distribution map and the ideal boundary distribution map respectively, and the difference in their amplitude spectra is compared.

[0044] The segmentation loss is the sum of weighted binary cross-entropy loss and weighted cross-union ratio loss.

[0045] The present invention also provides an image segmentation system, the key feature of which is: it includes a network construction unit for constructing a spatial-frequency domain cooperative boundary distribution network in the image segmentation method.

[0046] This invention provides an image segmentation method and system that uses a multi-scale boundary distribution generation module to predict a boundary distribution map as a priori, and employs a spatial-frequency domain joint supervision strategy to make the boundary distribution map approximate the real boundary in both spatial location and spectral structure. This effectively suppresses pseudo-edge responses caused by local high-frequency interference and improves the continuity and stability of the boundary. The state-space semantic enhancement module expands the effective receptive field and enhances global semantic consistency. By combining a boundary-guided cross-scale decoder with a boundary feature enhancement module, the boundary-related response is strengthened in multi-resolution fusion, achieving synergistic optimization of region localization and boundary refinement. Thus, under complex imaging conditions such as low contrast, diverse morphology, and interference from reflections and wrinkles, accurate and complete polyp segmentation results are obtained while maintaining reasonable computational overhead. Attached Figure Description

[0047] Figure 1 This is a schematic diagram of an image segmentation method provided in an embodiment of the present invention;

[0048] Figure 2 This is a structural diagram of the Space-Frequency Domain Cooperative Boundary Distribution Network (SFBD-Net) provided in an embodiment of the present invention;

[0049] Figure 3 This is a structural diagram of the Boundary Aggregation Module (BAM) provided in an embodiment of the present invention;

[0050] Figure 4 This is a structural diagram of the boundary quality regression (BQR) module provided in an embodiment of the present invention;

[0051] Figure 5 This is a structural diagram of the State Space Semantic Enhancement Module (S2FEM) provided in an embodiment of the present invention;

[0052] Figure 6 This is a structural diagram of the X-direction attention branch pool (X SSM Pool) and the Y-direction attention branch pool (Y SSM Pool) provided in an embodiment of the present invention. Figure 6 (a) shows the structure of the X-direction attention branch pool (X SSM Pool). Figure 6 (b) shows the structure of the Y-direction attention branch pool (Y SSM Pool);

[0053] Figure 7 This is a structural diagram of the Boundary Guided Cross-Scale Decoder (BGCSD) provided in an embodiment of the present invention;

[0054] Figure 8 This is a structural diagram of the boundary feature enhancement module (BFE) provided in an embodiment of the present invention;

[0055] Figure 9 This is a visualization result of the segmentation of static image data and dynamic frame-by-frame data obtained from experiments provided in this embodiment of the invention;

[0056] Figure 10 This is a comparison of the segmentation results of each model after ablation, provided in an embodiment of the present invention.

[0057] Figure 11 This is a comparison of the effective receptive field and attention heatmap results of each model after ablation, provided in the embodiments of the present invention.

[0058] Figure 12 This is a sensitivity analysis of the frequency domain loss weight λ provided in the embodiments of the present invention. Figure 12 (a) is the mDice coefficient analysis diagram. Figure 12 (b) is a graph showing the mIoU coefficient analysis.

[0059] Figure 13 The different embodiments of the present invention provide Segmentation comparison and boundary comparison under frequency domain loss Figure 13 In the middle (a), the comparison is segmented. Figure 13 (b) shows the comparison of BDMs from the same input group. Detailed Implementation

[0060] The embodiments of the present invention are described in detail below with reference to the accompanying drawings. The embodiments are given for illustrative purposes only and should not be construed as limiting the present invention. The accompanying drawings are for reference and illustration only and do not constitute a limitation on the scope of patent protection of the present invention, because many changes can be made to the present invention without departing from the spirit and scope of the present invention.

[0061] This invention first provides an image segmentation method, taking the segmentation of colonoscopy images in medical imaging as an example. The segmentation principle of this method is as follows: Figure 1 As shown, this method segments colonoscopy images by constructing and training a State-space Frequency-guided Boundary Distribution Network (SFBD-Net) for colonoscopy polyp segmentation. SFBD-Net is based on an encoder-decoder structure, using a boundary distribution map (BDM) as a prior, spatial and frequency domain supervision for refinement, and finally using the BDM as a prior-guided decoder to form a boundary distribution map closed-loop collaboration (BDM-CLC). The boundary distribution map closed-loop collaboration consists of three steps:

[0062] First, a boundary distribution map (BDM) is generated on the encoded features as a boundary prior;

[0063] The boundary distribution map (BDM) was then stabilized to ensure its consistency and reliability under conditions of interference and low contrast.

[0064] Finally, in the decoding stage, the boundary distribution map (BDM) is used to guide cross-scale fusion, suppressing false edges and enhancing true boundaries, thereby outputting more accurate segmentation results.

[0065] Figure 2 This is a detailed structural diagram of the space-frequency domain cooperative boundary distribution network. (See diagram below.) Figure 2 As shown, the network includes an encoding module, a multi-scale boundary distribution generation module (MFBGM), and a boundary distribution guidance decoding module (BGCSD).

[0066] The encoding module uses an encoder to extract multi-scale features, with the feature resolution halved step by step from shallow to deep.

[0067] The Multi-Scale Boundary Distribution Generation Module (MFBGM) includes a multi-branch multi-scale receptive block (including multiple multi-scale receptive blocks RFB), a boundary aggregation module (BAM), a boundary quality regression module (BQR), and an upsampling module. The multi-branch multi-scale receptive block (RFB) extracts the boundary context of different receptive fields of some scale features in the multi-scale features. Then, the boundary aggregation module (BAM) performs multiple rounds of interactive fusion between high- and low-resolution features. Finally, the boundary quality regression module (BQR) outputs a continuous boundary distribution map (BDM) (upsampled to the original image size). This BDM is simultaneously subject to joint spatial-frequency domain supervision: a combination of spatial regression loss and frequency domain consistency loss is applied to the boundary distribution map (BDM).

[0068] The Boundary Distribution Guided Decoding Module (BGCSD) includes a State Space Semantic Enhancement Module (S2FEM), a multi-level decoding part, and a segmentation part.

[0069] The State-Space Semantic Enhancement Module (S2FEM) models the state space along the x and y directions, outputting enhanced skip (jump connection) features with a larger effective receptive field and global semantic consistency. The multi-level decoding part uses the boundary distribution map (BDM) output by the Multi-Scale Boundary Distribution Generation Module (MFBGM) as a priori, combining it with the skip features enhanced by the State-Space Semantic Enhancement Module (S2FEM) to progressively restore resolution. At each decoding level, the Boundary Feature Enhancement Module (BFE) uses the boundary distribution map (BDM) to enhance the skip features with channel and spatial attention, highlighting boundary-related responses. The segmentation part interacts with the output features of the decoding part through a dual-branch process (low-resolution semantic branch and high-resolution boundary detail branch) and residual fusion, ultimately outputting a polyp segmentation mask of the same size as the input.

[0070] During the encoding stage, SFBD-Net uses the PVTv2 backbone network to extract multi-scale features from the input image (raw colonoscopy image). This embodiment takes four scales—1 / 4, 1 / 8, 1 / 16, and 1 / 32 resolution—as examples to obtain the corresponding four-scale features. , , and .

[0071] To explicitly characterize polyp boundaries at multiple semantic levels, SFBD-Net designed a Boundary Distribution Map (BDM)-CLC. First, a continuous boundary distribution map (BDM) is predicted using a multi-scale boundary distribution generation module (MFBGM). Then, during the decoding phase, the boundary-guided cross-scale decoder (BGCSD) uses this boundary distribution map (BDM) to explicitly guide the cross-scale fusion process. The MFBGM receives three scale features from mid-to-high levels (for example, 1 / 8, 1 / 16, and 1 / 32, corresponding to the features of the last three layers). , , First, multi-branch RFB is used to extract contextual information from different receptive fields. Then, the BAM module performs multi-round interactive fusion between high- and low-resolution features. Finally, the BQR module outputs the predicted boundary distribution map (BDM). Furthermore, to ensure that the predicted boundary distribution map (BDM) simultaneously approximates the ideal boundary distribution in both the spatial and frequency domains, this embodiment proposes a joint spatial-frequency domain supervision strategy. This strategy applies a combination constraint (weighted summation) of spatial regression loss and frequency domain consistency loss to the boundary distribution map (BDM), resulting in a more clearly defined, structurally coherent, and stable predicted boundary. The spatial regression loss is the average pixel-wise mean square error between the predicted boundary distribution map and the ideal boundary distribution map; the frequency domain consistency loss is the difference in amplitude spectra between the predicted boundary distribution map and the ideal boundary distribution map after two-dimensional fast Fourier transform.

[0072] Considering the importance of semantics at each layer of the encoder for polyp and background discrimination, SFBD-Net embeds a state-space semantic enhancement module (S2FEM) in the Skip connection. Through convolutional branches and state-space coordinate attention (SSM-CA), it collaboratively models local texture and long-range dependencies, providing a semantic representation with a larger effective receptive field that combines coarse and fine granularity for subsequent decoding.

[0073] In the decoding phase, SFBD-Net proposes a Boundary Distribution Guided Decoding Module (BGCSD), which embeds the generated Boundary Distribution Map (BDM) as an explicit prior into the multi-scale decoding process. This allows for the continuous utilization of boundary information guided by the BDM at different resolutions, achieving collaborative recovery of semantic localization and boundary details. To further enhance boundary sensitivity, a Boundary Feature Enhancement Module (BFE) is introduced into the BGCSD module to strengthen the response to realistic contours.

[0074] (1) Multi-scale boundary distribution generation module (MFBGM)

[0075] In colonoscopy polyp segmentation, interference is rampant, and significant differences in polyp scale can easily lead to boundary response breaks or offsets. To provide stable and interpretable boundary priors, this invention designs a multi-scale boundary distribution generation module (MFBGM) to predict continuous boundary distribution maps (BDM), and uses this as a prior input for the decoding part to enhance boundary localization and contour coherence.

[0076] MFBGM uses the features from the last three layers of the encoder as input. : 1 / 8, : 1 / 16, (1 / 32), such as Figure 2 As shown, the data sequentially passes through a multi-branch, multi-scale receptive block (RFB), a boundary aggregation module (BAM), and a boundary quality regression module (BQR), ultimately upsampling to the original image scale to output a boundary distribution map (BDM). The multi-scale receptive block (RFB) is responsible for capturing the boundary context under different receptive fields. The boundary aggregation module (BAM) is responsible for filtering out inconsistent noise across multiple scales and strengthening stable boundary patterns. The boundary quality regression module (BQR) performs confidence regression and geometric recalibration on the boundary response, ensuring that the output boundary distribution map (BDM) remains continuous even under strong disturbances.

[0077] Three multiscale receptive blocks (RFBs) for input , , After processing, multi-scale perceptual features are obtained:

[0078] ,

[0079] in, This represents the operation of multi-scale receptive blocks.

[0080] Figure 3 The diagram shows the structure of the Boundary Aggregation Module (BAM). Figure 3 As shown, the Boundary Aggregation Module (BAM) supports... , , The processing flow is as follows:

[0081] First, deep features Upsampling to the shallowest layer features Using the same scale, boundary feature maps containing different information granularities are constructed step by step by multiplying adjacent features:

[0082] ,

[0083] in, This indicates element-wise multiplication. Indicates an upsampling operation. This indicates a convolution operation.

[0084] Obtain boundary feature maps with different information granularities , and Then, the three feature maps are concatenated by channels and compressed by convolutional blocks to obtain compressed features:

[0085] ,

[0086] in, This indicates a channel splicing operation.

[0087] Finally, the compressed features are processed using attention blocks (CBAM) to focus more on important features and suppress unnecessary features, thus focusing more on boundary information, resulting in a boundary feature map:

[0088] ,

[0089] in, This represents the attention block operation, which enhances key features and suppresses irrelevant features by sequentially combining channel attention and spatial attention: First, spatial information is aggregated through global average pooling and global max pooling, and channel attention weights are generated through a shared multilayer perceptron. These weights are then multiplied with the original features channel by channel to achieve channel recalibration. Subsequently, global average pooling and global max pooling are applied again in the channel dimension. The resulting two two-dimensional feature maps are concatenated and then processed through a convolutional layer and activated by a sigmoid function to generate spatial attention weights. These weights are then multiplied with the weighted feature map position by position, and the final output is a feature representation that simultaneously focuses on discriminative channels and key spatial regions.

[0090] Figure 4 The diagram shows the structure of the Boundary Quality Regression (BQR) module. Figure 4 As shown, the Boundary Quality Regression (BQR) module applies the shallowest layer features. Compression features and boundary feature map The processing flow is as follows:

[0091] First, the boundary feature map is processed by Sigmoid. Transform into a boundary confidence graph:

[0092] ,

[0093] in, This represents the Sigmoid operation (normalization operation).

[0094] Subsequently, the boundary confidence graph is divided into four branches. Converted into four feature maps with different focus areas:

[0095] ,

[0096] in, This represents the Sobel boundary detection operator. Strengthen the high-confidence boundary response to help ensure boundary continuity. Suppress non-boundary areas, such as wrinkles and foam. This refers to a region with uncertain boundaries. Provide boundary geometry cues to make the BDM more closely match the real geometric contours.

[0097] At the same time, the shallowest features Compression features After convolution, the concatenated features are obtained:

[0098] .

[0099] Then, the feature maps will be stitched together. Multiplying each feature map by one of the four feature maps with different focuses yields four intermediate features:

[0100] .

[0101] Finally, the four intermediate features are concatenated and then convolved. The resulting convolution is then compared with the boundary feature map. Adding them together, we obtain the boundary quality regression feature map:

[0102] .

[0103] In the spatial-frequency joint supervision strategy of the multi-scale boundary distribution generation module (MFBGM), the construction process of the ideal BDM is as follows: first, extract the target boundary set S from the ground truth image (GT), and calculate the boundary values ​​for each pixel. The Euclidean distance to the boundary is then calculated, and this distance is mapped to the boundary distribution map using a Gaussian function.

[0104] (2) Boundary Distribution Guide Decoding Module (BGCSD)

[0105] To enhance the encoder's ability to model fine-grained structures and long-range dependencies, this invention proposes a State-Space Semantic Enhancement Module (S2FEM). This module integrates multi-scale local perception and orientation-aware global context modeling: on the one hand, it preserves local texture and boundary details through convolutional branches; on the other hand, it captures long-range dependencies between rows and columns using coordinate attention branches based on the State-Space Model (SSM), thereby expanding the effective receptive field, improving the discriminativeness and semantic integrity of features, and providing more comprehensive global information support for subsequent decoding. The overall structure of this module is as follows: Figure 5 As shown, its processing flow can be mathematically represented as follows:

[0106] ,

[0107] in, , These represent the input and output characteristics of the module, respectively. express Convolution operation, This represents the operation of the coordinate attention branch (SSM CA) in a state-space model (SSM). This represents the operation of concatenating multiple feature tensors along the channel dimension.

[0108] like Figure 5 As shown, the processing flow of the coordinate attention branch (SSM CA) in the state-space model (SSM) is as follows:

[0109] First, the input features respectively adopt as Figure 6 The X-direction attention branch pool (X SSM Pool) shown Figure 6 (a) and the Y-direction attention branch pool (Y SSM Pool) Figure 6 Attention is extracted from (b) to obtain the corresponding one-dimensional description. , The X SSM Pool and Y SSM Pool enhance the receptive field. The SSM Pool leverages the strength of SSM in modeling long-range dependencies, capturing long-range dependencies along the X and Y spatial directions to obtain a one-dimensional description. , ,in (i) represents the overall importance of the i-th row after scanning. (j) represents the overall importance of column j after scanning. When using SSM for calculation, rows or columns in the feature map are treated as sequences and calculated using the SSM formula, as shown below:

[0110] ,

[0111] in, Indicates the current time step Input sequence elements (input features) (a vector of pixel values ​​in a specific row or column of a dataset). The hidden state (latent vector) at the current time step is used to compress and remember historical input information. It is the hidden state of the previous time step; This represents the output of the current time step (the result after processing by SSM). This represents the state transition matrix, controlling the hidden state from... arrive The way things evolve determines the retention and attenuation of historical information; This represents the input mapping matrix, which maps the current input... The amount of update mapped to the hidden state; The output mapping matrix represents the hidden states. Mapped to output ; To directly pass the matrix, the input... It is directly superimposed on the output, and can usually be set to zero or equal to... Add them together.

[0112] In the SSM Pool of this invention, Initialized to 0.5. Initialize to 0 (i.e. By scanning the row or column sequence using this recursive formula, a one-dimensional vector describing the importance of the entire row / column is finally obtained.

[0113] Subsequently, by utilizing the relationship between directional information and channels, and The representation is concatenated in the spatial dimension and then reduced in dimensionality using a shared 3×3 convolution to obtain an intermediate representation:

[0114] ,

[0115] in, It is a convolutional block consisting of 3×3 convolution, BatchNormal (batch normalization), and ReLU.

[0116] Then Split along the x-axis and y-axis dimensions into and Where r is the reduction ratio of the control block size, and H, W, and C are the input features, respectively. The height, width, and number of channels are then used to generate row-direction attention maps and column-direction attention maps by performing two 1×1 convolutions and passing them through a sigmoid function.

[0117] ,

[0118] in This represents the Sigmoid activation function (normalization operation).

[0119] Finally, using and For input features Perform point-by-point recalibration and output features. Each position Calculated by the following formula:

[0120] ,

[0121] in, For channel indexing, These represent the positions in the height and width directions, respectively. Acting on the All columns of the row, Acting on the All rows of the column are multiplied together via a broadcast mechanism to achieve fine-grained weighting of the input features. This operation does not require... and Instead of performing channel stitching, this method directly utilizes the product of two directional attention maps to highlight important spatial locations while preserving precise location information. Final output While preserving the original local texture, long-range dependencies along rows and columns are incorporated, thereby achieving a larger effective receptive field and stronger global semantic consistency.

[0122] For the skip path, the State Space Semantic Enhancement Module (S2FEM) can significantly reduce the dominance of local texture noise, making subsequent cross-scale fusion more dependent on stable structures rather than accidental textures, which explains the reduction of boundary fractures and the improvement of generalization from a mechanistic perspective.

[0123] To fully utilize the boundary priors of the Boundary Distribution Map (BDM) and enhance the cross-scale aggregation capability during the decoding stage, this invention designs a Boundary Guided Cross-Scale Decoder (BGCSD), the overall structure of which is as follows: Figure 7 As shown, BGCSD employs a dual-path interaction of a decoder branch and a skip branch: the decoder branch receives the decoded features from the previous layer and generates the semantic representation at the current scale; the skip branch takes the encoded features enhanced by S2FEM as input and obtains boundary-enhanced features under the action of the boundary feature enhancement module (BFE). The two paths, guided by BDM, perform gating and cross-scale fusion, enabling the model to focus more on high-response boundary regions while performing region localization.

[0124] The entire process can be expressed by the following formula:

[0125] ,

[0126] in, The current layer decodes the output features. These are the output features of the previous layer's decoding. This is a low-resolution output characteristic. For high-resolution output features, we have:

[0127] ,

[0128] ,

[0129] in, For the feature enhancement operations of the Boundary Feature Enhancement Module (BFE), For the enhanced skip feature, For the boundary distribution map, This is for average pooling operations.

[0130] The structure of the Boundary Feature Enhancement Module (BFE) is as follows: Figure 8 As shown, its input consists of the skip features enhanced by the State Space Semantic Enhancement Module (S2FEM) and the corresponding Boundary Distribution Map (BDM). Specifically, the features obtained by convolving the enhanced skip features are multiplied by the features obtained by upsampling the Boundary Distribution Map (BDM). The multiplication result is then processed by channel attention and spatial attention modules for attention extraction. The extracted attention is then residually concatenated with the features obtained by convolving the enhanced skip features to obtain the output features.

[0131] Considering the significant selectivity of boundary information in both the channel and spatial dimensions, this invention cascades channel attention and spatial attention modules within the boundary feature enhancement module (BFE) to form a lightweight boundary enhancement unit: channel attention is used to select discriminative channels relevant to the boundary, and spatial attention is used to further focus on high-response regions of the boundary distribution map (BDM). A residual connection is used between the BFE output and the original skip features to avoid over-suppressing semantic information and improve training stability.

[0132] Boundary Guided Cross-Scale Decoder (BGCSD) achieves complementarity of features at different resolutions through bi-branch interaction and cross-scale fusion. At the same time, it uses the Boundary Feature Enhancement (BFE) module to explicitly enhance boundary-related responses during the decoding process, providing richer boundary information and lower noise feature representations for subsequent cross-scale aggregation.

[0133] (3) Loss function

[0134] To ensure that the boundary distribution map (BDM) is accurate in pixel location and stable in spectral structure, this invention designs a hybrid loss to jointly optimize the matching between the spatial and frequency domains. The total loss function of this invention consists of two parts: a segmentation loss used to supervise the final segmentation mask. and the boundary distribution loss used to supervise the boundary distribution map (BDM). The total loss is defined as follows:

[0135] .

[0136] To improve the accuracy of boundary distribution map (BDM) predictions and their structural consistency with the ideal BDM, this invention designs a hybrid loss function to jointly optimize the matching between the spatial and frequency domains. Boundary distribution loss. Defined as:

[0137] ,

[0138] in, For spatial domain loss, For frequency domain loss, To balance the weights.

[0139] For spatial domain loss First, given the predicted boundary distribution map (BDM) and the ideal boundary distribution map (BDM), calculate the mean square error (MSE) pixel by pixel:

[0140] ,

[0141] in, and These represent the pixel values ​​of the i-th pixel in the given predicted boundary distribution map (BDM) and ideal boundary distribution map (BDM), respectively.

[0142] Since the ideal boundary map (BDM) approaches zero in most non-boundary regions, directly applying the full map's MSE will result in a large number of easy samples dominating, diluting the effective gradient near the boundary. Therefore, a threshold is set. Backpropagation is only performed on pixels with larger errors. Finally, spatial domain loss... The definition is as follows:

[0143] ,

[0144] in, Indicates that the error is greater than A set of pixel indices.

[0145] Frequency domain loss aims to constrain the structural spectrum of the boundary distribution map (BDM). A two-dimensional FFT (Fast Fourier Transform) is performed on the predicted boundary distribution map (BDM) and the ideal boundary distribution map (BDM), and their spectral differences are compared. Frequency domain loss The definition is as follows:

[0146] ,

[0147] ,

[0148] in, These are the Predicted Boundary Distribution Map (BDM) and the Ideal Boundary Distribution Map (BDM), respectively. The results of performing two-dimensional FFT (Fast Fourier Transform) on the predicted boundary distribution map (BDM) and the ideal boundary distribution map (BDM) are shown respectively. Indicates taking the complex range, This represents the L1 norm. This loss term forces the predicted boundary map (BDM) to match the high-frequency energy distribution of the ideal boundary map (BDM) in the frequency domain, thereby suppressing abnormal high-frequency peaks caused by reflections, wrinkles, etc., and reducing false boundary responses.

[0149] Segmentation loss The final polyp segmentation mask, used for supervision, employs a combination of weighted BCE (Binary Cross-Entropy) and weighted IoU (Intersection over Union) to simultaneously ensure pixel-level classification stability and region overlap quality. Segmentation Loss Designed as follows:

[0150] ,

[0151] in, For weighted binary cross-entropy loss, The weighted average loss is calculated by combining the results.

[0152] Weighted binary cross-entropy loss Weighted average loss Defined as:

[0153] ,

[0154] ,

[0155] in, This represents the probability value predicted by the network. For true labels (0 or 1). Pixel-level weights, subscripts Indicates the first in the image line, number A pixel in the column.

[0156] Under the closed-loop cooperation mechanism of the Boundary Distribution Map (BDM), the gradient generated by the segmentation loss is backpropagated to the decoder and the prior branch of the Boundary Distribution Map (BDM), prompting the network to learn the solution that "the region prediction should be consistent with the stable boundary prior", thereby further improving the overall segmentation effect.

[0157] Corresponding to the above method embodiments, this invention also provides an image segmentation system, including:

[0158] A network construction unit is used to construct the space-frequency domain cooperative boundary distribution network.

[0159] The embodiments described in this invention can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with the implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0160] Computer programs for implementing the methods and systems of the present invention may be written in any combination of one or more programming languages ​​and stored in a computer-readable storage medium. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0161] Computer-readable storage media can be tangible media that may contain or store computer programs for use by or in conjunction with an instruction execution system, apparatus, or device. Computer-readable storage media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, optical disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0162] In summary, polyps in real colonoscopy images often exhibit low contrast, diverse morphologies, and blurred boundaries, accompanied by interference from reflections, folds, and mucus, making existing methods prone to contour breaks and false detections. This invention proposes a frequency-domain guided boundary distribution network (SFBD-Net) for colonoscopy polyp segmentation. This network predicts boundary distribution maps (BDMs) on multi-scale features using a multi-scale boundary distribution generation module (MFBGM) and uses these as boundary priors. A joint spatial-frequency domain supervision strategy is proposed to simultaneously constrain the learning of the boundary distribution map (BDM) from both boundary location and spectral structure levels, making it more stable and coherently approximate the real boundary distribution. Skip connections introduce a state-space semantic enhancement module (S2FEM) to strengthen skip features, expand the effective receptive field, and improve global semantic consistency. At the decoding end, a boundary-guided cross-scale decoder (BGCSD) is designed, combined with a boundary feature enhancement module (BFE), to repeatedly utilize the boundary distribution map (BDM) prior to strengthen boundary-related responses during multi-resolution fusion, thereby achieving coordinated optimization of region localization and boundary refinement.

[0163] (4) Experiment and Results

[0164] To comprehensively evaluate the segmentation model SFBD-Net designed in this invention, the following six metrics are used for evaluation: average Dice coefficient (mDICE), average intersection-over-union ratio (mIOU), weighted F-measure (the harmonic mean of weighted precision and recall), and so on. S-measure (predicted structural similarity to the true value) ), Maximum E-measure (enhanced alignment at the optimal threshold), The first five metrics are the mean absolute error (MAE) and the mean absolute error (MAE). For the first five metrics, a higher value indicates better performance; while for the last metric (MAE), a lower value is better.

[0165] The performance of a model is evaluated by considering both recall and precision, and is defined as follows:

[0166] ,

[0167] in, and Recall and precision, respectively, and the adjustment coefficient. Set it to 1.

[0168] The formula used to measure the structural similarity between the predicted results and the true annotations is as follows:

[0169] ,

[0170] in, These represent the similarity measures at the object-level and region-level, respectively, with parameter α being a balance coefficient set to 0.5.

[0171] Simultaneous calculations are performed at both the pixel and image levels, defined as follows:

[0172]

[0173] in, This represents the enhancement alignment matrix between local pixel values ​​and the image-level mean; H and W are the height and width of the input, respectively. This represents the maximum value across all images.

[0174] MAE is used to calculate the absolute error between the prediction result and the ground truth (GT), with a value range of [0–1]. The calculation formula is as follows:

[0175]

[0176] in For GT pixels, For the predicted pixels, m is the total number of pixels.

[0177] This experiment evaluated the model on five publicly available datasets: Kvasir-SEG, CVC-ClinicDB, CVC-ColonDB, CVC-300, and ETIS. Kvasir-SEG and CVC-ClinicDB were used for training and in-domain testing, while CVC-ColonDB, CVC-300, and ETIS were used only for cross-dataset testing to examine generalization ability. To further verify the model's robustness in more clinically relevant video scenarios, additional evaluation was performed on the PolypGen Video Sequence dataset. This dataset consists of 23 colonoscopy video sequences, containing 2225 pixel-level annotated images, exhibiting more significant motion blur, glare interference, and interference from medical instruments and other objects.

[0178] All experiments were implemented in the PyTorch framework and trained on two NVIDIA RTX3090 GPUs. Input images were uniformly scaled to 352×352. The Adam optimizer was used with an initial learning rate of 1e-4, which was gradually decayed to 1e-7 over 100 epochs using cosine annealing. The batch size was set to 36. Data augmentation techniques such as random flipping and rotation were employed during training. All comparison methods were reproduced under the same data partitioning, input resolution, and evaluation metric settings to ensure fair comparison.

[0179] The experiments compare SFBD-Net with various representative polyp segmentation methods, including classic CNN architectures, Transformer / hybrid architectures, and recent high-performance networks. Specific comparison models include U-Net, UNet++, SSFormer, CFANet, VM-Unet, NPD-Net, PraNet-V2, and DALA. Static data performance results are shown in Tables 1, 2, and 3.

[0180] Table 1 Performance of the Kvasir-SEG dataset

[0181]

[0182] Table 2 Performance of the CVC-ClinicDB dataset

[0183]

[0184] Table 3. Performance Comparison of Cross-Dataset Generalization Experiments

[0185]

[0186] As shown in Tables 1 and 2, the SFBD-Net designed in this embodiment achieves an mDice of 0.930 and an mIoU of 0.883 on the Kvasir-SEG dataset, and an mDice of 0.925 and an mIoU of 0.877 on the CVC-ClinicDB dataset, ranking best or tied for best in multiple metrics. Table 3 shows that in cross-dataset tests, it achieves mDice / mIoU of 0.807 / 0.726, 0.903 / 0.837, and 0.805 / 0.727 on the CVC-ColonDB, CVC-300, and ETIS datasets, respectively, outperforming most of the comparison methods and demonstrating strong generalization ability.

[0187] The segmentation visualization results of the static image data and dynamic frame-by-frame data obtained from the experiment are as follows: Figure 9 As shown. From Figure 9 It can be seen that in scenes with reflections, folds, or blurred boundaries, most methods are prone to producing phenomena such as broken contours, outward expansion of boundaries, or missegmentation of folds as polyps. In contrast, SFBD-Net can more accurately fit the true boundaries of polyps and maintain more complete regional coherence in scenes with small polyps or low contrast. Figure 9 The lower half shows that SFBD-Net can better distinguish between polyp and non-polyp regions in each frame of video data compared to other models, and can also achieve good segmentation results in the presence of highlights, blur, and other interference.

[0188] The test results for the PolypGen Video Sequence dataset are shown in Table 4. As can be seen from Table 4, SFBD-Net achieves the highest mDice and mIoU, and the lowest value on MAE.

[0189] Table 4 Performance comparison on the PolypGen Video Sequence dataset

[0190]

[0191] To verify the effectiveness of each key component, this embodiment conducted module ablation experiments on the Kvasir-SEG and CVC-ClinicDB datasets. The ablation results are shown in Table 5.

[0192] The experiment defined the baseline model as one that retained only the backbone and basic decoding structure, excluding the key components BDM-CLC, S2FEM, and BFE. Apart from the model structure, all other training and inference settings (data partitioning, input resolution, loss function, optimizer, number of training epochs, etc.) remained consistent, serving as an ablation control. Subsequently, each module was gradually added to examine its contribution to segmentation accuracy and boundary quality.

[0193] Table 5 SFBD-Net Ablation Experiments

[0194]

[0195] The baseline achieved mDice / mIoU of 0.896 / 0.844 and 0.881 / 0.835 on Kvasir-SEG and CVC-ClinicDB, respectively. Adding BDM-CLC (with BDSG) further improved mDice to 0.917 and 0.910, and mIoU to 0.864 and 0.859, demonstrating that explicit BDM modeling and its supervision significantly enhance boundary localization capabilities. Further introducing S2FEM (without BFE) improved mDice to 0.923 and 0.917, indicating that the larger effective receptive field brought by global state-space modeling helps suppress background interference such as wrinkles / reflections and improve target consistency. Finally, adding BFE (SFBD-Net) at the decoding end resulted in optimal model performance, validating the crucial role of boundary enhancement mechanisms in improving boundary fit and detail recovery during cross-scale fusion.

[0196] The segmentation comparison results of each model after ablation are as follows: Figure 10 As shown, white represents true positives, red represents false positives, and blue represents false negatives. Figure 10As shown, the introduction of BDM-CLC enhances the contour continuity at weak boundaries; the addition of S2FEM gradually expands the effective receptive field of the model from local texture to cover the entire polyp region.

[0197] The results of the comparison of the effective receptive field and attention heatmap of each model after ablation are as follows: Figure 11 As shown, the selected structure is the last decoding module of the decoder. Figure 11 As shown, with the introduction of BDM-CLC, S2FEM, and BFE, attention gradually shifts from the scattered background response to the polyp, resulting in a more complete focus on the polyp as a whole.

[0198] Ablation experiments show that introducing boundary distribution modeling and adding frequency domain consistency constraints improves the model segmentation effect, indicating that frequency domain constraints can suppress pseudo high-frequency edges caused by reflection and texture, and improve contour breakage and undersegmentation problems.

[0199] Figure 12 For the sensitivity analysis of the frequency domain loss weight λ, Figure 12 (a) shows the mDice coefficient analysis. Figure 12 (b) shows the mIoU coefficient analysis. Figure 12 As shown, when λ=0.03, SFBD-Net achieves optimal or near-optimal overall performance on the Kvasir-SEG and CVC-ClinicDB datasets.

[0200] Figure 13 For different Segmentation comparison and boundary comparison under frequency domain loss Figure 13 In the middle (a), the comparison is shown, where white represents true positives, red represents false positives, and blue represents false negatives; Figure 13 (b) shows the BDM comparison within the same input group. Figure 13 It can be seen that not adding frequency domain constraints will lead to breakage in boundary segmentation, while an excessively large λ may overemphasize spectral consistency and thus affect region segmentation. Therefore, subsequent experiments will use λ=0.03 by default.

[0201] Furthermore, under the same input resolution (3×352×352), the theoretical computational cost (FLOPs) and parameter size (Params) of each method were statistically analyzed, as shown in Table 6.

[0202] Table 6. Statistical Comparison of Parameters with Strong Baseline Model

[0203]

[0204] As shown in Table 6, SFBD-Net has 15.73G FLOPs and 29.99M parameters, which is acceptable among most strong baseline methods. Combined with its performance improvement on the Kvasir-SEG and CVC-ClinicDB datasets, it can be seen that the boundary prior closed-loop design proposed in this invention can achieve higher boundary quality and overall segmentation performance without significantly increasing computational overhead.

[0205] In summary, this invention provides a comprehensive evaluation of SFBD-Net on six publicly available datasets. Experimental results show that:

[0206] 1. The spatial-frequency domain collaborative constraint strategy proposed in this invention can effectively suppress local high-frequency interference such as reflection and wrinkles, so that the predicted boundary distribution map (BDM) is close to the real boundary in terms of both location and spectral structure, thereby reducing contour breaks and false edges.

[0207] 2. The State Space Semantic Enhancement Module (S2FEM) expands the effective receptive field by modeling the state space along the x / y direction, enhancing global semantic consistency, and significantly improving region integrity, especially in low-contrast and small polyp scenes.

[0208] 3. During the decoding process, BGCSD repeatedly uses BDM priors for gating and cross-scale fusion, and combines BFE to explicitly enhance the boundary-related response, so that the model can still output a coherent and fitting polyp contour under complex dynamic conditions such as motion blur and mechanical interference.

[0209] In summary, this invention effectively addresses the problems of boundary breakage, false detection, and insufficient generalization ability in existing technologies under complex imaging conditions by constructing a BDM closed-loop collaboration (BDM-CLC) of "boundary prior generation—stability supervision—prior-guided decoding". Experimental data on six public datasets fully verify the significant advantages of the proposed spatial-frequency domain collaborative constraint method in improving polyp segmentation accuracy, boundary quality, and clinical robustness.

[0210] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.

Claims

1. An image segmentation system, characterized in that, This includes a space-frequency domain cooperative boundary distribution network, which comprises an encoding module, a multi-scale boundary distribution generation module, and a boundary distribution guidance decoding module. The encoding module uses an encoder to extract multi-scale features from the input image; The multi-scale boundary distribution generation module predicts a continuous boundary distribution map from the partial-scale features output by the encoding module, and applies joint spatial and frequency domain supervision to the boundary distribution map. The joint spatial and frequency domain supervision involves applying a combination of spatial regression loss and frequency domain consistency loss to the boundary distribution map. The spatial regression loss is the average pixel-wise mean square error between the predicted boundary distribution map and the ideal boundary distribution map. The frequency domain consistency loss is the difference in amplitude spectra between the predicted boundary distribution map and the ideal boundary distribution map after undergoing two-dimensional fast Fourier transform. The boundary distribution guidance decoding module is used to enhance the multi-scale features through the state space semantic enhancement module to obtain jump connection features, and to perform multi-level decoding with the boundary distribution map as a priori to output a polyp segmentation mask. The boundary distribution guidance decoding module includes a state space semantic enhancement module, a multi-level decoding part, and a segmentation part; The state space semantic enhancement module is used to perform state space modeling on the multi-scale features output by the encoding module along the row and column directions to obtain enhanced skip connection features; The multi-level decoding part is used to recover the resolution step by step by using the boundary distribution map as a priori and combining the enhanced skip connection features to obtain multi-level decoding features; The segmentation part is used to obtain a polyp segmentation mask by interacting with the low-resolution semantic branch and the high-resolution boundary detail branch and by residual fusion of the multi-level decoded features. The processing flow of the state space semantic enhancement module is mathematically represented as follows: , in, , These represent the input and output characteristics of the module, respectively. express Convolution operation, The operation representing the coordinate attention branch of the state-space model. This represents the operation of concatenating multiple feature tensors along the channel dimension; The operation flow of the coordinate attention branch in the state-space model is as follows: The input features are sequentially scanned along the row and column directions using a state-space model to obtain one-dimensional row description vectors and one-dimensional column description vectors. The one-dimensional row description vector and the one-dimensional column description vector are concatenated in the spatial dimension, and then the intermediate representation is obtained by convolution dimensionality reduction. The intermediate representation is split into two components along the row and column directions. Each component is then subjected to a 1×1 convolution and a sigmoid activation to obtain the corresponding row-direction attention map and column-direction attention map. The input features are recalibrated point by point using the row direction attention map and the column direction attention map.

2. The image segmentation system according to claim 1, characterized in that, The multi-scale boundary distribution generation module includes: A multi-branch, multi-scale receptive block is used to capture the boundary context of different receptive fields for the partial scale features output by the encoding module, thereby obtaining multi-scale receptive features. The boundary aggregation module is used to perform multi-round interactive fusion of the multi-scale receptive features to obtain a boundary feature map. The boundary quality regression module is used to perform boundary confidence regression and geometric recalibration on the boundary feature map and the shallowest features in the multi-scale receptive features to obtain the boundary quality regression feature map. The upsampling module is used to upsample the boundary quality regression feature map to the original image size to obtain the boundary distribution map.

3. The image segmentation system according to claim 2, characterized in that: The boundary aggregation module upsamples the deep features in the multi-scale receptive features and multiplies them element-wise with the shallow features to construct boundary feature maps of different information granularities. The resulting boundary feature maps are then concatenated, compressed by convolution, and processed by the convolutional block attention module to output the boundary feature maps. The boundary quality regression module obtains a boundary confidence map by passing the boundary feature map through a Sigmoid function. This boundary confidence map is then used to generate multiple feature maps with different focuses through multiple branches. Simultaneously, the shallowest feature in the multi-scale receptive features is concatenated with the compressed features output by the boundary aggregation module and then convolved to obtain a concatenated feature map. This concatenated feature map is multiplied by multiple feature maps with different focuses, and the multiplication results are concatenated and then convolved. Finally, it is added to the boundary feature map to obtain the boundary quality regression feature map.

4. The image segmentation system according to claim 1, characterized in that, The multi-level decoding section employs a multi-level boundary-guided cross-scale decoder for step-by-step decoding. The decoding process is represented as follows: , , , in, This is a low-resolution output characteristic. Features for high-resolution output The current layer decodes the output features. These are the output features of the previous layer's decoding. For the feature enhancement operation of the boundary feature enhancement module, For the enhanced skip connection features, For the boundary distribution map, For average pooling operation, This indicates element-wise multiplication. Indicates an upsampling operation. This indicates a convolution operation.

5. The image segmentation system according to claim 4, characterized in that, The boundary feature enhancement module operates as follows: The features obtained by convolving the enhanced skip connection features are multiplied with the features obtained by upsampling the boundary distribution map. The result of the multiplication is then used to extract attention through the channel attention module and the spatial attention module. The extracted attention is then residually connected with the features obtained by convolving the enhanced skip connection features to obtain the output features.

6. The image segmentation system according to claim 1, characterized in that: The overall loss function of the space-frequency domain cooperative boundary distribution network includes segmentation loss and boundary distribution loss; the boundary distribution loss includes spatial domain loss and frequency domain loss. The spatial domain loss calculates the mean square error of the predicted boundary distribution map and the ideal boundary distribution map pixel by pixel, and calculates the average error only for pixels whose error is greater than a set threshold. The frequency domain loss is used to perform two-dimensional fast Fourier transforms on the predicted boundary distribution map and the ideal boundary distribution map respectively, and the difference in their amplitude spectra is compared. The segmentation loss is the sum of weighted binary cross-entropy loss and weighted cross-union ratio loss.

7. An image segmentation method, characterized in that: The image to be segmented is input into an image segmentation system according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Tampered region positioning method based on boundary attention mechanism

    CN118470112A

  • Through-the-wall radar human body behavior recognition method and recognition system based on spatial-temporal characteristics

    CN121348263A