Medical image segmentation method based on edge feature capture and multi-scale feature fusion

By adopting neural network methods based on edge feature capture and multi-scale feature fusion in medical image segmentation, the problems of unclear edges and poor contrast of polyp images are solved, and higher precision edge positioning and segmentation are achieved, improving the efficiency of medical image processing.

CN120070477AActive Publication Date: 2025-05-30QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510542267.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

In medical image segmentation, especially the problems of unclear edges of polyp images, dim light in the intestines, and poor contrast between polyp tissue and surrounding tissues, it makes it difficult for the prior art to achieve high-precision edge positioning and segmentation.

Method used

The neural network method based on edge feature capture and multi-scale feature fusion is adopted to enhance edge feature information through edge feature capture module, and the multi-scale feature fusion module is used to fuse features of different scales and levels, combining the dual attention mechanism to reduce the impact of background noise and image artifacts.

Benefits of technology

It improves the accuracy of edge positioning of polyp images, solves the problem of background noise during feature fusion, reduces the impact of image artifacts such as halo and reflection, and significantly improves the efficiency and accuracy of medical image segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070477A_ABST
    Figure CN120070477A_ABST
Patent Text Reader

Abstract

The invention provides a medical image segmentation method based on edge feature capture and multi-scale feature fusion, and belongs to the field of medical image processing, and the method comprises the steps: carrying out the preprocessing of a to-be-segmented gastrointestinal polyp endoscope image; reading the preprocessed gastrointestinal polyp endoscope image, and obtaining digital information of the image; inputting the obtained image digital information into an edge feature capture and multi-scale feature fusion neural network to obtain edge features and multi-scale features of the input image, and forming a feature map; on the basis of multi-scale features, features of different scales and levels are fused together, post-processing operation is carried out on the basis of segmentation result features, and a final gastrointestinal polyp segmentation result is generated. On the basis of a deep learning method, the labor cost is reduced, the image segmentation time is shortened, and the efficiency is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image processing, and particularly to a medical image segmentation method based on edge feature capture and multi-scale feature fusion. Background Art

[0002] As a classic medical image segmentation model, the U-Net network structure has now been widely used in the field of medical images. Its U-shaped structure solves the deficiency of the FCN in losing non-local feature information. In recent years, researchers have continuously improved the U-Net network to enhance the medical image segmentation effect. Swin-Unet, Transformer-Unet, and UTNet all improve the segmentation performance of the model based on the U-Net structure. Although there are numerous related improvement methods emerging, once faced with tasks with extremely high precision requirements, especially when the tasks involve complex spatial features and fuzzy edge details, these methods still cannot meet the needs. In fact, more accurate image segmentation results often rely on context information and edge features, so many studies focus on exploring related methods. For example, CO-Net and T-Net obtain regionally significant information by fusing cross-layer features. Among various deep learning architectures, the PraNet model shows significant advantages in fusing high-level semantic features and extracting edge information. BDGNet provides a new perspective for boundary modeling by introducing the concept of a boundary distribution map. The MISSFormer model has made further achievements in capturing global dependencies and local context information. DANet strengthens the capture of extensive global context information by introducing a dual attention mechanism. The enhanced deep neural network, which fuses the shape reconstruction (SR) and spatial constraint (SC) mechanisms, effectively handles complex problems in segmentation. Another study shows that introducing star convex constraints can ensure accurate segmentation while preserving the topological structure.

[0003] Although existing studies have proposed polyp segmentation network models with significant effects, the field of polyp segmentation still faces many challenges. Such as Figure 1 shown, in the figure, they respectively represent (a) the polyp edge is not clear; (b) the polyp is not obvious due to dim light in the intestine; (c) the contrast between the polyp tissue and the surrounding tissue is poor. The challenges faced in locating the polyp boundary in medical images are multi-faceted. This article addresses these challenges through the following key contributions based on previous studies.

[0004] In gastrointestinal polyp medical images, the polyp lesion area and the surrounding normal tissues usually have low contrast in endoscopic images. The color of the polyp, the texture of the surrounding environment, and the mucosal color are similar, resulting in very unclear boundaries. In the case of manual annotation, it is very easy to miss detections or misdiagnose, which will have a bad impact on the treatment effect of patients and even endanger the lives of patients. However, the technology of medical image segmentation has overcome this problem. The emergence of medical image segmentation technology can clearly divide the lesion area. For doctors, it helps to improve the accuracy of doctor diagnosis and can easily distinguish the lesion area from the surrounding normal physiological tissues through the segmented results. Traditional image segmentation methods mainly rely on manual operations by doctors. This method is not only inefficient but also prone to errors. When faced with a large amount of medical image data, doctors need to spend a lot of time and energy to process them one by one, resulting in an increased work burden. At the same time, the results of manual segmentation may affect the accuracy of diagnosis due to subjective factors. Therefore, it is urgent to explore more efficient and automated image segmentation technologies to improve the efficiency of medical image processing, reduce human errors, and thus improve the overall quality of medical services. Summary of the Invention

[0005] The purpose of the present invention is to provide a medical image segmentation method based on edge feature capture and multi-scale feature fusion, which not only improves the accuracy of polyp image edge localization, but also solves the background noise problem that appears in the feature fusion process and reduces the influence of image artifacts such as light halos and reflections.

[0006] To achieve the above object, the present invention is realized through the following technical solutions: A medical image segmentation method based on edge feature capture and multi-scale feature fusion includes the steps of: (1) Preprocess the endoscopic image of gastrointestinal polyps to be segmented; (2) Read the preprocessed endoscopic image of gastrointestinal polyps to obtain the digital information of the image; (3) Input the obtained digital image information into a neural network for edge feature capture and multi-scale feature fusion to obtain the edge features and multi-scale features of the input image, and form a feature map; (4) Based on the multi-scale features, fuse the features of different scales and levels together, and based on the segmentation result features, perform post-processing operations to generate the final segmentation result of gastrointestinal polyps.

[0007] Further, the steps for the neural network of edge feature capture and multi-scale feature fusion to form a feature map include: Divide the input image into blocks and divide it into multiple stages, and perform downsampling before the start of each stage; Pass the downsampled input image through the backbone network of the Swim Transformer to obtain a feature map with halved height and width, generating multi-scale features; Enhance the edge feature information between the encoder and the decoder through an edge feature capture module; Adopt a multi-scale feature fusion module to fuse multiple layers and output a prediction map.

[0008] Furthermore, the steps of the multi-scale feature fusion module for fusing an image include: Use convolution operations to unify the number of feature channels; Enhance and then fuse the image features through a preposed dual attention enhancement module.

[0009] Furthermore, the dual attention enhancement module includes a channel attention mechanism and a spatial attention mechanism. Enhancing and then fusing includes the steps of: Dynamically learn the importance of each channel feature through the channel attention mechanism; Focus on the features of key spatial regions through the spatial attention mechanism to achieve feature enhancement; Adopt a Cat operation for feature splicing to achieve feature fusion and dimensionality reduction.

[0010] Furthermore, the edge feature capture module includes erosion, dilation operations, element addition operations for edge features, and element multiplication operations for retaining local features.

[0011] Furthermore, the steps for the edge feature capture module to enhance the edge feature information include: Divide the input encoder feature map X1 into four branches; The first branch includes Softmax binarization, Maxpooling max pooling, and Tanh activation. X1 is binarized through the Softmax function to obtain , and then erosion and dilation are achieved through Maxpooling max pooling to obtain , and is obtained through the Tanh activation function ; The second branch includes Softmax binarization, Maxpooling max pooling, and Sigmoid activation. Maxpooling max pooling achieves erosion and dilation to obtain , and is obtained through the Sigmoid activation function ; The third branch includes a linear layer and Sigmoid activation. Add the original feature map X1 to the eroded feature map to obtain and input it into the linear layer. The output of the linear layer is obtained through the Sigmoid activation function ; The fourth branch retains X1; The final single output is 。

[0012] Furthermore, the post - processing operation includes the steps of: Evaluating the quality of the result of image segmentation; Selecting a post - processing method according to the quality evaluation of the segmentation result; Using Gaussian filtering for denoising, calling a two - dimensional discrete Gaussian function to remove noise and retain edge details.

[0013] Furthermore, the training process of the neural network for edge feature capture and multi - scale feature fusion includes: (1) Pre - processing the endoscopic image of gastrointestinal polyps to be segmented; (2) Reading the pre - processed endoscopic image of gastrointestinal polyps to obtain the digital information of the image; (3) Inputting the obtained digital image information into the neural network for edge feature capture and multi - scale feature fusion to obtain the edge features and multi - scale features of the input image, and forming a feature map; (4) Based on the multi - scale features, fusing features of different scales and levels together, and performing post - processing operations based on the segmentation result features; (5) Optimizing the feature extraction and segmentation module of the model based on the judgment result and the real result of the model; Repeat steps (3) - (5) until the result is optimal; To evaluate the segmentation effect of the network model on gastroscope polyp images, it can be judged by evaluating the change trend of the accuracy rate when using the validation set and other evaluation indicators during testing.

[0014] The evaluation indicators used are Intersection - Over - Union and weighted binary cross - entropy, and IoU is used as the evaluation indicator to measure the similarity between the segmentation result and the true label.

[0015] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. It is characterized in that when the processor executes the computer program, the steps of the above - mentioned medical image segmentation method based on edge feature capture and multi - scale feature fusion are implemented.

[0016] A computer - readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above - mentioned medical image segmentation method based on edge feature capture and multi - scale feature fusion are implemented.

[0017] The advantages of the present invention are: The deep learning framework established by the present invention is characterized by edge feature capture and multi-scale feature fusion. The edge feature capture module enhances the edge features, thereby improving the accuracy of polyp image edge localization. The multi-scale feature fusion module first enhances and then fuses the features, aiming to precisely fuse the shallow features and deep features, which helps to establish the context information between the segmented target objects, solve the background noise problem that appears during the feature fusion process, and this module adds dual attention to strengthen the perception of detailed information (including texture and color), thereby reducing the impact of image artifacts such as halos and reflections. Using a neural network with edge feature capture and multi-scale feature fusion can make the features of the image edge more obvious during training and extract more scale feature information.

[0018] Based on the method of deep learning, the present invention greatly reduces the labor cost. Relying on the design and construction of the model framework, the computer can learn the standard for segmenting gastroscopy images by itself during multiple training and learning processes. Thus, the computer can segment the gastroscopy image objects by itself without manual intervention. Moreover, the present invention can shorten the time for segmenting images and greatly improve the efficiency. Brief Description of the Drawings

[0019] Figure 1 Shows the problems existing in polyp images in the prior art; Figure 2 Is the neural network structure diagram of edge feature capture and multi-scale feature fusion in Embodiment 1 of the present invention; Figure 3 Is the schematic diagram of the edge capture module in Embodiment 1 of the present invention; Figure 4 Is the schematic diagram of channel attention in Embodiment 1 of the present invention; Figure 5 Is the schematic diagram of spatial attention in Embodiment 1 of the present invention; Figure 6 Is the schematic diagram of the multi-scale feature fusion module structure in Embodiment 1 of the present invention; Figure 7 Is the schematic diagram of the segmentation result in Embodiment 1 of the present invention. Detailed Embodiment

[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0021] Embodiment 1 This embodiment discloses a medical image segmentation method based on edge feature capture and multi-scale feature fusion, including the steps of: preprocessing the endoscopic image of gastrointestinal polyps to be segmented; reading the preprocessed endoscopic image of gastrointestinal polyps to obtain the digital information of the image; inputting the obtained digital image information into a neural network for edge feature capture and multi-scale feature fusion to obtain the edge features and multi-scale features of the input image, forming a feature map; based on the multi-scale features, fusing features of different scales and levels together, and based on the segmentation result features, performing post-processing operations to generate the final segmentation result of gastrointestinal polyps.

[0022] Specifically, it includes the following steps.

[0023] S1. The image preprocessing module performs preprocessing operations on medical image data. Different from conventional data sets, the acquisition method of medical image data is relatively special. Its acquisition depends on hospital cooperation channels, and the annotation work is completed by professional physicians. During the image acquisition process, the manifestations of different types of medical images affected by noise interference are different. For example, the lesion area in the gastrointestinal endoscope image often has a blurred boundary with the surrounding tissues. To improve the segmentation accuracy, the image is optimized.

[0024] The threshold processing of brightness equalization is the core part of the image preprocessing module to eliminate the noise and interference in the image and improve the quality of the endoscopic image of gastrointestinal polyps. By comparing and analyzing the scoring results when the thresholds are 2, 4, 8, 16, and 32 respectively, this study finally determines the scheme with 16 as the threshold for subsequent experiments. This process effectively suppresses other noise interferences and enhances the features of the area to be segmented.

[0025] S2. By using a computer to read the image, the digital conversion of the image can be realized, that is, it is converted into a three-dimensional numerical form in the RGB color space. Given the computing characteristics of the computer itself and the artificially set standards, the value ranges of the three components (red, green, and blue) of the image in the RGB color space are all limited between 0 and 255.

[0026] To improve the efficiency and accuracy of subsequent data processing and calculation, during the process of reading the image, the RGB three-component values corresponding to each pixel point in the image are synchronously normalized. The specific normalization method is: divide the RGB component value of each pixel point by 127.5, and then subtract 1. After such processing, the numerical range of the image is successfully mapped to the interval [-1, 1], which is more convenient for subsequent analysis and calculation work.

[0027] S3. The edge feature capture module strengthens the edge feature information between the encoder and the decoder through the edge feature capture module to distinguish the target area from the background, so as to improve the performance of the segmentation model. For gastroscope polyp endoscopic images, due to the influence of factors such as the endoscopic shooting angle, illumination, and shadow, a large amount of detailed information will be ignored when extracting image features at a single scale. Therefore, the main function of the multi-scale feature fusion module is to fuse local features and global features at different levels, significantly improving the feature quality. This solution designs a feature extraction module mainly based on edge feature capture and multi-scale feature fusion.

[0028] The overall details of the image feature extraction module are as follows: The backbone network based on Swin Transformer extracts features from the input image. First, the input image is divided into blocks, and every 4*4 adjacent pixels form a patch. After flattening, the dimension is changed to a value acceptable by the Transformer through a linear embedding layer, and then through four stages. Before each stage, downsampling is performed. The input image passes through the backbone network of Swin Transformer to obtain a feature map with halved height and width, generating multi-scale features. The edge feature capture module strengthens the edge feature information between the encoder and the decoder to distinguish the target area from the background, so as to improve the performance of the segmentation model. The lower part of the model is the upsampling decoder. During the upsampling process, a prediction map is output for each layer. In the last three layers of output, a multi-scale feature fusion module is added, and the outputs of the last three layers are fused by the method of first enhancing and then fusing to output a prediction map.

[0029] In order to obtain information from high-level features, an edge feature capture module and a multi-scale feature fusion module are used to retain multi-scale features and more image detail information, simplify the model, and retain more fine features.

[0030] Please refer to the edge feature capture module Figure 3 , in the U-Net network architecture, the input image is passed layer by layer through the encoder-decoder structure, and its feature map is directly passed from the encoder to the decoder through skip connections, so as to achieve multi-level feature fusion and restore the original image information during the reconstruction process. Since polyp medical images contain complex edge details and spatial characteristics, it is difficult to accurately segment the lesion area only relying on traditional skip connections. Therefore, this paper proposes an optimized module specifically for extracting edge features. This module is single-input and single-output, and is divided into four branches, namely erosion, dilation operation, element addition operation of edge features, and element multiplication operation of retaining local features.

[0031] First, the encoder feature map X1 is binarized by the Softmax function. The Softmax function can map each pixel value in the feature map to a probability value between 0 and 1, thereby enhancing the contrast of the features. The first branch X1 is binarized by the Softmax function to obtain .

[0032] Then, erosion and dilation are achieved through Maxpooling max - pooling. MaxPooling (max - pooling) is a commonly used pooling operation in convolutional neural networks (CNNs). It retains the most prominent features by taking the maximum value in the local region of the input feature map. This module uses a MaxPooling operation with a kernel size k of 7 to implement the erosion and dilation process. During the dilation step, for the normalized feature map using the MaxPooling operation can capture the maximum value within the local region, thereby obtaining the expanded foreground region.

[0033] During the erosion step, for the inverted and normalized feature map using the MaxPooling operation can capture the minimum value within the local region, thereby obtaining the expanded background region. In this paper, we use a MaxPooling operation with a stride s of 1 and a padding p of 3 to ensure that the size of the feature map remains unchanged. The first and second branches are obtained through the following formulas respectively: The Tanh and Sigmoid activation functions are respectively applied to the feature maps after erosion and dilation. The Tanh function can map the pixel values of the feature map to a range between - 1 and 1, while the Sigmoid function maps the pixel values to a range between 0 and 1.

[0034] The first branch passes through the Tanh activation function to obtain , which is the data obtained by using the MaxPooling operation on the normalized feature map ; The second branch passes through the Sigmoid activation function to obtain , which is the data obtained by using the MaxPooling operation on the inverted and normalized feature map .

[0035] The activated feature maps are element - multiplied to obtain the similarity matrix V. The similarity matrix can reflect the similarity degree between different features, thereby helping the network better distinguish the target region and the background.

[0036] The third branch adds the original feature map X1 to the eroded feature map to obtain and inputs it into a linear layer. The output of the linear layer is obtained through the Sigmoid activation function .

[0037] The fourth branch retains X1, and finally obtains an enhanced similarity matrix: .

[0038] The details of the multi-scale feature fusion module are as follows: The DCFF (Deep Cross-Feature Fusion) multi-scale feature fusion module is introduced in the text. The results output by the last three layers in the network model are used as the input of the multi-scale feature fusion module. The input scales of the three-level features are 56×56, 112×112, and 224×224 respectively. Convolution operations are used to unify the number of feature channels, and then the image features are enhanced and fused through a preposed dual attention enhancement module, significantly improving the feature quality. First, the channel attention mechanism is used to dynamically learn the importance of each channel feature. The channel attention structure is as Figure 4 shown.

[0039] Then, the spatial attention mechanism is used to focus on the features in the key spatial regions to achieve feature enhancement. The spatial attention structure is as Figure 5 shown.

[0040] Next, feature concatenation is performed. The Cat operation is adopted, and feature fusion and dimensionality reduction are achieved through 3×3 convolution. This module adopts a lightweight design, significantly reducing the computational complexity. 1×1 convolution is used for compression, and finally, 1×1 convolution generates the prediction result. The structure of this module is as Figure 6 shown.

[0041] S4. Before post-processing, it is necessary to evaluate the quality of the image segmentation results. According to the quality evaluation of the segmentation results, the post-processing method is selected. Gaussian filtering is used for denoising. A two-dimensional discrete Gaussian function is called to remove noise, which can retain more edge details, making the image clearer and the smoothing effect softer, to help doctors determine the size, location, and shape of the polyp, so as to know the formulation of the surgical and treatment plans.

[0042] Example 2 This embodiment also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The feature is that when the processor executes the computer program, the steps of the above-mentioned medical image segmentation method based on edge feature capture and multi-scale feature fusion are implemented.

[0043] During the training process of the network model, it is recommended to use a GPU because the amount of computation required for network model training is large, the computation is relatively complex, and image processing is involved. Using a GPU can greatly accelerate the training process. However, it should be noted that when using a GPU, a corresponding operating framework needs to be selected, such as the GPU version of the TensorFlow framework.

[0044] Regarding the processing process of computer instructions, the type of processor is not limited. In this embodiment, the processor can be a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), or an application-specific integrated circuit (ASIC), etc. The memory can include a read-only memory (ROM) and a random access memory (RAM), and provide instructions and data to the processor. Software modules can be stored in mature storage media in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. These storage media are located in the memory, and the processor completes the steps of the above method by reading the information in the memory and combining its hardware. During implementation, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or instructions in software form.

[0045] Embodiment 3 This embodiment proposes a model training process, and the process is as follows: preprocess the medical images with labels, read the preprocessed images, extract the feature information of the images through different weights, segment different structures or tissues of the images, evaluate the accuracy of the image segmentation algorithm, calculate the result error, and update the weights according to the error, that is, optimize the feature extraction and segmentation module. Through continuous iteration, the accuracy of image feature extraction and segmentation is continuously optimized until the optimal result is achieved.

[0046] The following are the relevant settings of the experiment: (1) Experimental environment All models are implemented in Pytorch with two NVIDIA TITAN RTX 24GB GPUs.

[0047] (2) Optimization and iteration The Adam optimizer is used, and the initial learning rate of all models is 0.0001. The dice loss and cross-entropy are used as the loss functions, and the dice similarity coefficient (DSC) and intersection over union (IoU) are used as the performance metrics.

[0048] To train the model, deep supervision is performed on the output feature map. Before calculating the loss, the output is upsampled to the same size as the ground truth label. The total loss can be expressed as: Among them, G is the ground truth label, Si is the output on the i side, Sg is the global feature, L is the loss function, and i is the network layer.

[0049] In the dataset used, CVC-ClinicDB is a colonoscopy image dataset, including a total of 612 annotated images. Kvasir-SEG is a polyp segmentation dataset with 1000 annotated images. Subsequently, 80% of the CVC-ClinicDB and Kvasir-Seg datasets were randomly divided into a training dataset and 20% for the test dataset. The image sizes of all datasets were resized to 224 × 224. For the CVC-ClinicDB and Kvasir-Seg datasets, the batch was set to 4.

[0050] Denosing the divided training set and validation set images is the core part of the image preprocessing module to eliminate noise and interference in the images and improve the quality of gastrointestinal polyp endoscopic images. By comparing and analyzing the scoring results when the thresholds are 2, 4, 8, 16, and 32 respectively, this embodiment finally determines the scheme with 16 as the threshold for subsequent experiments. This process effectively suppresses other noise interferences and enhances the features of the area to be segmented.

[0051] After the model architecture is confirmed, a graphics processor is used to train the weights of the model. The hyperparameters of the model can be adjusted through the training accuracy change curve and validation accuracy change curve, as well as the training loss change curve and validation loss change curve (usually already tuned to the optimal by the experimenter). After multiple iterative trainings, the obtained weight values are most suitable for medical image segmentation.

[0052] The model testing process is as follows: load the trained weights and use the test set to perform the performance test of the model. The entire testing process is roughly the same as the training process, except that the testing process directly obtains the segmentation result without iteration. Finally, the effectiveness of the model is confirmed through comparison and relevant evaluation metrics.

[0053] In the experiment, to evaluate the segmentation effect of the network model on gastroscope polyp images, it can be judged by evaluating the change trend of the accuracy rate when using the validation set and other evaluation metrics during testing.

[0054] The evaluation metrics used are the Intersection-Over-Union and weighted binary cross-entropy. In image segmentation, the MIoU is usually used as an evaluation metric to measure the similarity between the segmentation result and the ground truth label. The calculation formula is: The calculation formula for the Dice coefficient is: Among them, k is the number of categories, that is, the number of target categories to be distinguished in the dataset. MIoU represents the mean intersection over union, true positive (TP), false positive (FP), true negative (TN), and false negative (FN). TP means making an accurate prediction for the target area that is actually an organ or lesion; FP indicates misjudging the area that originally belongs to the background as the target category; TN represents the successful and correct recognition of the background area; FN refers to misjudging the area that should actually be classified as an organ or lesion as the background.

[0055] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A medical image segmentation method based on edge feature capture and multi-scale feature fusion, characterized in that: Includes steps: (1) Preprocessing the segmented gastrointestinal polyp endoscopic images; (2) reading the pre-processed endoscopic image of gastrointestinal polyps and obtaining digital information of the image; (3) Input the obtained image digitization information into the neural network for edge feature capture and multi-scale feature fusion to obtain the edge features and multi-scale features of the input image and form a feature map; (4) Based on multi-scale features, features of different scales and levels are fused together, and post-processing operations are performed based on the segmentation result features to generate the final gastrointestinal polyp segmentation results.

2. The medical image segmentation method based on edge feature capture and multi-scale feature fusion according to claim 1, characterized in that: The step of forming a feature map by a neural network of edge feature capture and multi-scale feature fusion includes: The input image is divided into blocks and multiple stages, and downsampling is performed before each stage starts; The downsampled input image is passed through the backbone network of Swim Transformer to obtain a feature map with half the height and width, generating multi-scale features; The edge feature information is enhanced through the edge feature capture module between the encoder and the decoder; A multi-scale feature fusion module is used to fuse multiple layers to output a prediction map.

3. The medical image segmentation method based on edge feature capture and multi-scale feature fusion according to claim 2, characterized in that: The step of fusing images by the multi-scale feature fusion module comprises: Use convolution operation to unify the number of feature channels; The image features are first enhanced and then fused through the front-end dual attention enhancement module.

4. The medical image segmentation method based on edge feature capture and multi-scale feature fusion according to claim 3 is characterized in that: The dual attention enhancement module includes a channel attention mechanism and a spatial attention mechanism, and the enhancement and fusion steps include: Dynamically learn the importance of each channel feature through the channel attention mechanism; Focusing on the features of key spatial regions through the spatial attention mechanism to achieve feature enhancement; Cat operation is used for feature splicing to achieve feature fusion and dimensionality reduction.

5. The medical image segmentation method based on edge feature capture and multi-scale feature fusion according to claim 2, characterized in that: The edge feature capture module includes erosion and expansion operations, element addition operations of edge features, and element multiplication operations that retain local features.

6. The medical image segmentation method based on edge feature capture and multi-scale feature fusion according to claim 5, characterized in that: The edge feature capture module strengthens the edge feature information including the steps of: Divide the input encoder feature map X1 into four branches; The first branch includes Sofmax binarization, Maxpooling, and Tanh activation. X1 is binarized by the Softmax function. , and then use Maxpooling to achieve erosion and expansion, and get , obtained through the Tanh activation function ; The second branch includes Softmax binarization, Maxpooling, and Sigmoid activation. Maxpooling implements erosion and expansion to obtain , obtained through the Sigmoid activation function ; The third branch includes a linear layer and a Sigmoid activation layer, which combines the original feature map X1 with the eroded feature map Add together to get And input to the linear layer, the output of the linear layer Obtained through Sigmoid activation function ; The fourth branch retains X1; The final single output is .

7. The medical image segmentation method based on edge feature capture and multi-scale feature fusion according to claim 1, characterized in that: The post-processing operation comprises the steps of: Perform quality assessment on image segmentation results; Select post-processing methods based on the quality assessment of the segmentation results; Use Gaussian filtering to denoise, call a two-dimensional discrete Gaussian function to remove noise and retain edge details.

8. The medical image segmentation method based on edge feature capture and multi-scale feature fusion according to claim 1, characterized in that: The training process of the neural network for edge feature capture and multi-scale feature fusion includes: (1) Preprocessing the segmented gastrointestinal polyp endoscopic images; (2) reading the pre-processed endoscopic image of gastrointestinal polyps and obtaining digital information of the image; (3) Input the obtained image digitization information into the neural network for edge feature capture and multi-scale feature fusion to obtain the edge features and multi-scale features of the input image and form a feature map; (4) Based on multi-scale features, features of different scales and levels are integrated together, and post-processing operations are performed based on the segmentation result features; (5) Based on the model’s judgment results and actual results, optimize the model’s feature extraction and segmentation module; Repeat steps (3) to (5) until the result is optimal; The segmentation effect of the network model on gastrointestinal polyp images can be evaluated by evaluating the accuracy trend when using the validation set and other evaluation indicators during testing; The evaluation indicators used are Intersection-Over-Union and weighted binary cross entropy. IoU is used as the evaluation indicator to measure the similarity between the segmentation result and the true label.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the medical image segmentation method based on edge feature capture and multi-scale feature fusion as described in any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the medical image segmentation method based on edge feature capture and multi-scale feature fusion as described in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Polyp segmentation method and equipment based on colonoscope image and medium

    CN117197470A

  • Arophic gastritis area segmentation method based on multi-scale boundary refinement and fusion

    CN118115490A

  • Brain tumor image segmentation method based on multi-scale convolution and Mama structure

    CN118447244A

  • Colored drawing cultural relic true color image fusion method based on multi-branch fusion strategy convolution auto-encoder

    CN118710514A

  • Kidney tumor CT image segmentation method based on edge information enhancement

    CN118710672A