A Gastrointestinal Polyp Segmentation Method Based on Spatial-Frequency Feature Interaction and Boundary Enhancement Network
By introducing space-frequency feature interaction and boundary enhancement networks into the polyp segmentation method, the processing challenges of irregular shapes and blurred boundaries are solved, and the accuracy and efficiency of polyp segmentation are significantly improved.
Patent Information
- Application Number
- CN202411107329.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-13
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-08-13
AI Technical Summary
Existing polyp segmentation methods are challenging in dealing with irregular shapes and blurred boundaries, and missed detection and misdiagnosis are common, affecting the early screening and diagnosis of colorectal cancer.
A digestive tract polyp segmentation method based on space-frequency feature interaction and boundary enhancement network is adopted, and the prediction results are gradually refined through the spatial and frequency feature interaction module (SFFI) and boundary enhancement module (BE), combined with deep supervised learning.
Effectively extract spatial and frequency characteristics, enhance the accuracy of boundary detection, overcome the challenges of traditional methods in dealing with irregular shapes and blurred boundaries, and significantly improve the accuracy and efficiency of polyp segmentation.
Smart Images

Figure CN119006818B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer-aided medical diagnosis, and particularly relates to a method for segmenting digestive tract polyps based on spatial-frequency feature interaction and boundary enhancement network. Background Art
[0002] Colorectal cancer (CRC) has an extremely high incidence rate globally and is one of the main causes of cancer deaths. Research shows that the incidence and mortality rates of CRC are showing a sharp upward trend. Therefore, it is particularly important to explore effective methods for early screening and diagnosis. Endoscopy is the standard tool for screening polyps and can provide information on the appearance and location of polyps so that doctors can remove them before they deteriorate. Research shows that early endoscopy can reduce the incidence of colorectal cancer by 30%. Therefore, early screening and removal of polyps are crucial for the prevention and treatment of CRC.
[0003] However, in clinical diagnosis, due to the heavy burden on the medical system, insufficient resources, the cumbersome and somewhat subjective polyp segmentation work, the phenomena of missed detection and misdiagnosis are relatively common. Research shows that the proportion of missed detection of colonic polyps during endoscopy is as high as 27%. To improve the efficiency and accuracy of endoscopy, researchers have begun to use computer-aided diagnosis systems to automatically detect polyps through medical image analysis techniques. Among them, polyp segmentation technology can help doctors more accurately identify and extract polyp regions, thereby improving the efficiency of diagnosis and treatment.
[0004] In recent years, deep learning techniques have made remarkable progress in the field of medical image analysis, providing new solutions for polyp segmentation. Methods based on the "encoder-decoder" structure have shown excellent performance in feature extraction and generating segmentation results. However, these methods mainly focus on spatial information while ignoring the equally important frequency information. Frequency information can complement spatial information and provide more comprehensive lesion information. Researchers have tried to introduce frequency information into the polyp segmentation task, such as using the Laplacian pyramid to adaptively calibrate multi-frequency domain feature representations, or combining color features in the spatial domain and global structural features in the frequency domain. In addition, the discrete wavelet transform (DWT) has also been used to remove redundant frequency features and improve segmentation accuracy. These methods have improved the performance of the segmentation model by integrating frequency information. Besides frequency information, context information is also crucial for polyp segmentation. Context information reflects the spatial relationship between the polyp and its surrounding environment, helping the model better understand the whole polyp. Researchers have used methods such as dilated convolution, dense convolution, and feature fusion at different scales to expand the network receptive field and extract richer context information to improve segmentation performance. Although existing methods have made some progress, there are still some challenges. Future research needs to explore more effective methods for frequency information extraction and fusion, as well as more effective context information modeling methods to further improve the segmentation performance of the model. Summary of the Invention
[0005] To solve the above problems, the present invention provides a method for segmenting digestive tract polyps based on spatial-frequency feature interaction and boundary enhancement network. The method uses a new polyp segmentation network based on dual-domain feature interaction (DFINet) as the backbone network. Different from the current methods that mainly mine polyp features in the spatial domain, DFINet achieves segmentation by combining spatial and frequency domain features. To address the problems of irregular shapes and fuzzy boundaries, DFINet adopts a spatial and frequency feature interaction module SFFI, which explores shape perception information in the spatial and frequency domains. To reduce errors in boundary detection, DFINet applies a boundary enhancement module BE, which integrates cross-layer features to help the network focus on the boundary region. And deep supervision learning is used to gradually refine and finally obtain the prediction result.
[0006] The technical solution of the present invention is as follows:
[0007] A method for segmenting digestive tract polyps based on spatial-frequency feature interaction and boundary enhancement network, the method comprising:
[0008] Obtain a polyp segmentation dataset;
[0009] Input the dataset into the polyp segmentation network to extract multi-layer features, and fuse the multi-layer features to generate a rough segmentation map;
[0010] The generated rough segmentation map is input into the spatial and frequency feature interaction module to extract the output features of the fused spatial and frequency features;
[0011] The boundary enhancement module integrates the output features of all spatial and frequency feature interaction modules to generate the prediction result.
[0012] Furthermore, the polyp segmentation dataset is divided into a test set and a training set.
[0013] Furthermore, inputting the dataset into the polyp segmentation network to extract multi-layer features and fusing the multi-layer features to generate a rough segmentation map specifically includes:
[0014] The PVTv2-b2 module extracts features from the input polyp segmentation dataset to obtain the first-layer feature map, the second-layer feature map, the third-layer feature map, and the fourth-layer feature map;
[0015] The second-layer feature map, the third-layer feature map, and the fourth-layer feature map are input into the feature fusion module for convolution operations, and finally the three convolved feature maps are fused to generate a rough segmentation map.
[0016] Furthermore, the feature fusion module includes three parallel branches and an aggregation unit; among them, each branch includes a 1×1 convolution and two 3×3 convolutions.
[0017] Furthermore, inputting the generated rough segmentation map into the spatial and frequency feature interaction module to extract the output features of the fused spatial and frequency features specifically includes:
[0018] The first-layer feature map, the second-layer feature map, and the third-layer feature map are input into the spatial and frequency feature interaction module for channel compression; the compressed first-layer feature map, the second-layer feature map, and the third-layer feature map are obtained; then the compressed first-layer feature map is convolved with a 3×3 convolution to obtain spatial features; the compressed second-layer feature map and the third-layer feature map are respectively subjected to the Haar discrete wavelet transform to obtain low-frequency features and high-frequency features. The spatial features, low-frequency features, and high-frequency features are concatenated, and then passed through a spatial attention unit and a channel attention unit in sequence to remove irrelevant information. Finally, the obtained features are further convolved with a 1×1 convolution to adjust the number of channels to 32 to obtain the output features of the fused spatial and frequency features.
[0019] Furthermore, the formula for obtaining the output features of the fused spatial and frequency features is:
[0020]
[0021] Among them, and φ s (·) respectively represent the spatial attention unit and the channel attention unit, Fi s is a spatial feature, F i L is a low-frequency feature, F i H is a high-frequency feature.
[0022] Furthermore, the boundary enhancement module integrates the output features of all spatial and frequency feature interaction modules to generate the prediction result, specifically:
[0023] The boundary enhancement module integrates the output features of the spatial and frequency feature interaction modules of the current layer, the output features of the adjacent higher-layer spatial and frequency feature interaction modules, and the prediction results of the adjacent higher layer to generate the prediction result of the current layer.
[0024] Furthermore, it also includes:
[0025] Using the weighted binary cross-entropy loss function L seg to evaluate the quality of each prediction result, and the total loss function L o The result is represented as the sum of the L of all prediction results; the calculation formula of the total loss function is as follows: seg
[0026]
[0027] where, S i represents the prediction result of the i-th layer, and G represents the true segmentation region.
[0028] Compared with the prior art, the present invention has the following advantages:
[0029] The present invention can effectively extract spatial and frequency features, and deepen the interaction between layers, so as to more accurately identify the fuzzy boundaries of polyps, effectively overcome the challenges of traditional methods in dealing with irregular polyp shapes and fuzzy boundaries, and show strong performance in efficiently and quickly processing polyp segmentation. In addition, this method performs well on different datasets, has stronger learning ability and generalization ability, and can be applied to various image analysis systems, with important application value and market potential. Brief Description of the Drawings
[0030] The drawings generally illustrate various embodiments by way of example rather than limitation, and are used together with the description and the claims to explain the embodiments of the invention. Where appropriate, the same reference numerals are used in all the drawings to refer to the same or similar parts. Such embodiments are illustrative and are not intended to be an exhaustive or exclusive embodiment of the device or method.
[0031] Figure 1 Shows a schematic diagram of the polyp segmentation network framework of the present invention;
[0032] Figure 2 Shows a schematic diagram of the spatial and frequency feature interaction module of the present invention;
[0033] Figure 3 Shows a schematic diagram of the boundary enhancement module of the present invention. Detailed implementation manners
[0034] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.
[0035] An embodiment of the present invention proposes a digestive tract polyp segmentation method based on spatial-frequency feature interaction and boundary enhancement network. The specific technical solution includes the following steps:
[0036] 1. Dataset acquisition
[0037] The present invention uses five publicly available polyp segmentation datasets, including CVC-ClinicDB, Kvasir, ETIS, CVC-ColonDB, and CVC-300, to comprehensively evaluate the performance of endoscopic image segmentation algorithms. Specifically, 550 samples from CVC-ClinicDB and 900 samples from Kvasir are used as the training set, and the remaining CVC-ClinicDB and Kvasir samples, as well as all samples of the ETIS, CVC-ColonDB, and CVC-300 datasets, constitute the test set. The above dataset segmentation strategy divides the test process into two parts: in-domain testing and out-of-domain testing. In-domain testing (i.e., testing on CVC-ClinicDB and Kvasir) aims to evaluate the performance of the algorithm under similar data distributions and directly reflects the effectiveness of the algorithm. While out-of-domain testing (i.e., testing on ETIS, CVC-ColonDB, and CVC-300) is used to measure the generalization ability of the algorithm when facing different data distributions. Table 1 details the sample numbers of each dataset and their specific uses in this study.
[0038] 2. Model design
[0039] DFINet uses PVTv2-b2 as the backbone to extract multi-level features, and uses the feature fusion module FAB to fuse some features to generate a rough prediction map. At the same time, the spatial and frequency feature interaction module SFFI processes the extracted multi-level features to extract shape-aware features in both the spatial domain and the frequency domain. To further enhance the network's attention to the polyp boundary, the boundary enhancement module BE concatenates features of different levels and the prediction map of the adjacent higher level to generate the final segmentation map. The BE module can locate the polyp area from coarse to fine through a deep supervision mechanism. Next, the network framework, related modules, and implementation details will be introduced in detail.
[0040] Step 1: In the feature encoding stage, DFINet first uses PVTb2-v2 to extract feature maps with different spatial resolutions at multiple layers. Among them, C i , and represent the number of channels, height, and width of the i-th layer feature F i (i ∈ {1, 2, 3, 4}), respectively.
[0041] Step 2: The features of the 2nd, 3rd, and 4th layers are fused to generate a rough segmentation map. Specifically, F2, F3, and F4 are input into the FAB module, which contains three parallel branches, and each branch contains a 1×1 convolution and two 3×3 convolutions. Finally, the outputs of the three branches are aggregated into an aggregation unit AGG derived from PraNet to finally generate the rough segmentation map S4. To improve the model efficiency, the present invention does not use the features of the 1st layer because the information volume of this layer is limited and it will increase the network parameters.
[0042] Step 3: As Figure 2 shown, the spatial and frequency feature interaction module SFFI consists of three parallel branches. At the beginning of each branch, a 1×1 convolution and a 3×3 convolution are used to compress the input features in channels, and F i1 , F i2 and F i3 are obtained respectively, where F iτ = Conv 3×3 (Conv 1×1 (F i )), τ ∈ {1, 2, 3}. In the formula, Conv k×k (·) represents the convolution operation with a convolution kernel of k. Next, F i1 passes through a 3×3 convolution to obtain the spatial feature F i s . In the second and third branches, the Haar discrete wavelet transform (DWT) is used to decompose F i2 and F i3 into low-frequency components and high-frequency components. Taking the second branch as an example, F i2 is processed using low-pass and high-pass filters respectively to extract the low-frequency sub-component F i2 L and the high-frequency sub-component F i2 H . Then, F i2 L and F i2 H are respectively processed again with the same low-pass and high-pass filtering to obtain wavelet components with different frequencies F i2LL , F i2 LH , F i2 HL and F i2 HH . Next, perform a 3×3 convolution operation and inverse DWT (Inverse DWT, IDWT) on F i2 LL to extract the low-frequency feature F i L . Similarly, in the third branch, F i3 also undergoes a DWT operation to obtain the high-frequency components (F i3 LH , F i3 HL and F i3 HH ), and these three high-frequency components are concatenated. Similarly, after a 3×3 convolution and IDWT, the high-frequency feature F i H is finally obtained. The above operations can be expressed as:
[0043]
[0044] where represents the feature concatenation operation.
[0045] Next, concatenate the spatial feature F i s and the frequency feature F i L and F i H , and successively pass through a spatial attention (SA) unit and a channel attention (CA) unit to remove irrelevant information. Finally, the obtained feature passes through a 1×1 convolution to adjust the number of channels to 32, and the output of the current SFFI layer is obtained. Where in the formula and φ s (·) represent the CA and SA units respectively. This module directly adopts the widely used CBAM module, which includes a CA and an SA unit.
[0046] Step 4: The boundary enhancement module BE is used to integrate the output of the current layer SFFI module the output of the adjacent higher-layer SFFI module and the prediction result S of the adjacent higher layer i to generate the prediction result S of the current layer i+1. The BE module helps the network obtain richer semantic information while retaining detailed information by fusing features at different levels, thereby improving the accuracy of segmentation.
[0047] Figure 3 Shows the structural framework of the boundary enhancement module BE. First, perform and S i+1 upsampling to match the size. Then, perform a 3×3 convolution operation on the upsampled and respectively, multiply the obtained results, and finally obtain , where In the formula represents the dot product operation, and ↑2(·) represents the upsampling operation with a sampling rate of 2. Similarly, perform the same operation on S i+1 to obtain S i ', where Through the above operations, the BE module promotes the interaction between the current and adjacent high-level information, which helps to better capture polyp perception features.
[0048] Next, perform an average pooling operation with a convolution kernel size of 3 on S i '. Then, subtract the pooled result from the original S i ' to obtain the boundary information B. To highlight the boundary region, concatenate B with S i ', and the concatenated features are then passed through a 3×3 convolution. The above process can be summarized as follows:
[0049]
[0050] where Avgpool(·) represents the average pooling operation.
[0051] Finally, add and together, and then successively pass through a 3×3 convolution and a 1×1 convolution to obtain the prediction result S i of the current layer, where Since there is no higher layer, the BE module is not set at the top layer.
[0052] Step 5: Supervise the output prediction maps of each layer of the BE module to achieve rough-to-fine polyp localization, and use the average value of S1 and S2 as the final prediction result.
[0053] 3. Loss function
[0054] Use the weighted binary cross-entropy loss function L segTo evaluate the quality of each prediction result. To achieve deep supervision, the total loss function L o The result is represented as the L of all prediction results seg The sum. The calculation formula of the total loss function is as follows: where S i represents the prediction result of the i-th layer, and G represents the true segmentation region. Since the resolutions of S i and G may be different, before calculating the loss, adjust S i to the same resolution as G.
[0055] 4. Network Training
[0056] During the training of the model, the loss between each prediction result and the true segmentation region is calculated through the above loss function, and the network parameters are updated using the backpropagation algorithm. According to the gradient information calculated by the loss function, the extraction descent algorithm is used to adjust the parameters to minimize the loss. The training weights of the model are saved every 5 epochs.
[0057] 5. Network Testing
[0058] First, uniformly crop the image to be segmented to a size of 352×352. Then, use the previously trained weights to test the cropped image to obtain the prediction results S i (i ∈ {1, 2, 3, 4}). Take the average of S1 and S2 as the final prediction map.
[0059] Example 1
[0060] The framework is as Figure 1 shown. This method first obtains the colorectal polyp dataset of endoscopic images and uniformly crops all images to a size of 352×352. Subsequently, the images in the training set are input into the designed polyp segmentation model for training. The model learns the image features and generates corresponding prediction results. Then, the error between the prediction result and the true segmentation region is calculated through the loss function, and the model parameters are adjusted using the gradient descent method to minimize the loss value. The model weights are saved every 5 epochs. After the model training is completed, the images in the test set are input into the fixed model to obtain the final predicted polyp segmentation result. Finally, the prediction result is compared with the labeled true segmentation region, and relevant metrics are calculated to verify that the model proposed by the present invention can conveniently, quickly, and accurately identify polyps.
[0061] (1) Select the test platform
[0062] To evaluate the effectiveness of the model, the present invention conducted experiments on a platform based on the PyTorch deep learning framework. The experimental environment was equipped with an NVIDIA GeForce RTX 3090 GPU and an Intel(R) Xeon(R) Silver 4210R CPU. The popular PVTv2-b2 was adopted as the backbone, and all fully connected layers were removed. The backbone was pre-trained on the ImageNet dataset, while other layers were randomly initialized. The DFINet model was trained in an end-to-end manner, with the batch size set to 26, and the training was stopped at the 80th epoch. The network was optimized using the AdamW optimizer, and the initial learning rate was set to 1×10 i and adjusted to 1×10 r after the 50th epoch. -4 -5
[0063] (2) Dataset Acquisition
[0064] The present invention adopted five publicly available polyp segmentation datasets, including CVC-ClinicDB, Kvasir, ETIS, CVC-ColonDB, and CVC-300, to comprehensively evaluate the performance of endoscopic image segmentation algorithms. Specifically, 550 samples from CVC-ClinicDB and 900 samples from Kvasir were used as the training set, and the remaining CVC-ClinicDB and Kvasir samples, as well as all samples from the ETIS, CVC-ColonDB, and CVC-300 datasets, constituted the test set. The above dataset splitting strategy divided the test process into two parts: in-domain testing and out-of-domain testing. In-domain testing (i.e., testing on CVC-ClinicDB and Kvasir) aimed to evaluate the performance of the algorithm under similar data distributions and directly reflected the effectiveness of the algorithm. Out-of-domain testing (i.e., testing on ETIS, CVC-ColonDB, and CVC-300) was used to measure the generalization ability of the algorithm when facing different data distributions. Table 1 details the sample numbers of each dataset and their specific uses in this study. In addition, a multi-scale input strategy (i.e., randomly resizing the image size to 0.75, 1.00, and 1.25) was applied as data augmentation.
[0065] Table 1 Summary of the five polyp segmentation datasets used in this study
[0066] dataset resolution training sample test sample usage CVC-ClincDB 384×288 550 62 training + testing Kvasir 332×487~1920×1072 900 100 training + testing ETIS 1225×966 n / a 196 out-of-domain testing CVC-ColonDB 574×500 n / a 380 out-of-domain testing CVC-300 574×500 n / a 60 out-of-domain testing
[0067] (3) Model Design
[0068] As Figure 1 As shown in the figure. DFINet uses PVTv2-b2 as the backbone to extract multi-level features, and fuses some features through the feature fusion module FAB to generate a rough prediction map. At the same time, the spatial and frequency feature interaction module SFFI processes the extracted multi-level features to extract shape perception features in both the spatial domain and the frequency domain. To enhance the network's attention to the polyp boundary, the boundary enhancement module BE concatenates features of different levels and the prediction map of the adjacent higher level to generate the final segmentation map. The BE module locates the polyp area from rough to fine through the deep supervision mechanism.
[0069] (4) Loss function
[0070] The present invention uses a weighted binary cross-entropy loss function L seg to evaluate the quality of each prediction result. To achieve deep supervision, the total loss function L o The result is expressed as the sum of L seg for all prediction results. The calculation formula of the total loss function is as follows: where S i represents the prediction result of the i-th layer, and G represents the true segmentation region. Since the resolutions of S i and G may be different, before calculating the loss, S i is adjusted to the same resolution as G.
[0071] (5) Network training
[0072] During the training of the model, the loss between each prediction result and the true segmentation region is calculated through the above loss function, and the network parameters are updated using the backpropagation algorithm. According to the gradient information calculated by the loss function, the extraction descent algorithm is used to adjust the parameters to minimize the loss. Every 5 epochs of training, the model saves the training weights once.
[0073] (6) Network testing
[0074] First, the image to be segmented is uniformly cropped to a size of 352×352. Then, the cropped image is tested using the previously trained weights to obtain the prediction results S i (i∈{1,2,3,4}). The average of S1 and S2 is taken as the final prediction map.
[0075] (7) Method performance
[0076] To quantify and verify the performance of the proposed method, the network model of the present invention uses the average Dice similarity coefficient, average IoU similarity coefficient, S-measure, E-measure, weighted F-measure, and mean absolute error as image segmentation metric indicators.
[0077] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, shall be covered by the protection scope of the present invention.
Claims
1. A digestive tract polyp segmentation method based on space-frequency feature interaction and boundary enhancement network, characterized in that: The method comprises: Get the polyp segmentation dataset; Inputting the data set into a polyp segmentation network to extract multi-layer features, and fusing the multi-layer features to generate a rough segmentation map; The extracted multi-layer features are input into the spatial and frequency feature interaction module to extract the output features of the fusion of spatial features and frequency features; The output features of all spatial and frequency feature interaction modules are integrated through the boundary enhancement module to generate prediction results; The extracted multi-layer features are input into the spatial and frequency feature interaction module to extract the output features of the fusion of spatial features and frequency features. Specifically: The extracted multi-layer features are input into the spatial and frequency feature interaction module, which consists of three parallel branches. Each branch starts with a 1×1 convolution and a 3×3 convolution to compress the channels of the multi-layer features, and obtains the compressed feature maps F i1 , feature map F i2 and feature map F i3 ; Then the feature map F i1 Perform 3×3 convolution to obtain spatial features; feature map F i2 and feature map F i3 Haar discrete wavelet transform is used to obtain low-frequency features and high-frequency features respectively, and the spatial features, low-frequency features and high-frequency features are concatenated and sequentially passed through a spatial attention unit and a channel attention unit to remove irrelevant information. Finally, the obtained features are subjected to a 1×1 convolution to adjust the number of channels to 32, and the output features of the fusion of spatial features and frequency features are obtained.
2. The digestive tract polyp segmentation method based on space-frequency feature interaction and boundary enhancement network according to claim 1 is characterized in that: The polyp segmentation dataset is divided into a test set and a training set.
3. The digestive tract polyp segmentation method based on space-frequency feature interaction and boundary enhancement network according to claim 1, characterized in that: The data set is input into the polyp segmentation network to extract multi-layer features, and the multi-layer features are fused to generate a rough segmentation map as follows: The PVTv2-b2 module performs feature extraction on the input polyp segmentation dataset to obtain the first-layer feature map, the second-layer feature map, the third-layer feature map, and the fourth-layer feature map; The second-layer feature map, the third-layer feature map, and the fourth-layer feature map are input into the feature fusion module for convolution operation, and finally the three-layer feature maps after convolution are fused to generate a rough segmentation map.
4. The digestive tract polyp segmentation method based on space-frequency feature interaction and boundary enhancement network according to claim 1, characterized in that: The feature fusion module consists of three parallel branches and aggregation units; each branch contains one 1×1 convolution and two 3×3 convolutions.
5. The digestive tract polyp segmentation method based on space-frequency feature interaction and boundary enhancement network according to claim 1, characterized in that: The output feature formula of the fusion of spatial features and frequency features is: in, and φ s (·) represent the spatial attention unit and channel attention unit respectively, and F i s is the spatial feature, F i L is the low frequency feature, F i H is a high-frequency feature, Conv 1×1 To represent the convolution operation, Represents a feature concatenation operation.
6. The digestive tract polyp segmentation method based on space-frequency feature interaction and boundary enhancement network according to claim 1, characterized in that: The output features of all spatial and frequency feature interaction modules are integrated through the boundary enhancement module to generate the prediction results: The boundary enhancement module integrates the output features of the spatial and frequency feature interaction module of the current layer, the output features of the spatial and frequency feature interaction module of the adjacent high layer, and the prediction results of the adjacent high layer to generate the prediction results of the current layer.
7. The digestive tract polyp segmentation method based on space-frequency feature interaction and boundary enhancement network according to claim 1, characterized in that: Also includes: Use the weighted binary cross entropy loss function L seg To evaluate the quality of each prediction result, the total loss function L o The result is expressed as the L of all predicted results seg The total loss function is calculated as follows: Among them, S i represents the prediction result of the i-th layer, and G represents the actual segmentation area.