Eye fundus image lesion segmentation method and device, computer equipment and medium

A deep learning-based method for diabetic retinopathy lesion segmentation in retinal images addresses the limitations of existing methods by using a model with global-local attention and multi-scale features to automate accurate lesion detection, enhancing precision and reducing physician workload.

CN120318249AActive Publication Date: 2025-07-15QUFU NORMAL UNIV
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510388139.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-15
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The existing diabetic retinopathy lesions based on fundus color photos are difficult to achieve simultaneous segmentation of multiple types of lesions, and the accuracy and robustness are insufficient, and rely on high-cost expert manual labeling, making it difficult to popularize on a large scale.

Method used

Deep learning technology is used to build a fundus image lesion segmentation model, including an encoder, global local attention module and multi-scale feature capture module. Features are extracted through multi-head self-attention mechanism, deep separation convolution and pooling operations, combined with global local attention and multi-scale features, and Dice loss function is used to optimize the model to achieve automated segmentation of lesions.

Benefits of technology

High-precision automated segmentation of diabetic retinopathy areas is achieved, which reduces the workload of doctors, improves the accuracy of lesion detection, and enhances the generalization ability of the model to adapt to images of different diabetic retinopathy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318249A_ABST
    Figure CN120318249A_ABST
Patent Text Reader

Abstract

The invention provides a fundus image focus segmentation method and device, computer equipment and a medium, and belongs to the technical field of medical image processing. According to the method, the generalization ability of a model is improved through data preprocessing, and accurate segmentation of four kinds of retinopathy lesions in an eye fundus image is realized by using an eye fundus image lesion segmentation model comprising an encoder, a global local attention module, a multi-scale feature capture module and a decoder. The encoder adopts a three-branch structure to extract multi-scale features, the global and local attention module fuses global and local attention features, the multi-scale feature capture module extracts multi-scale lesion features through different convolution kernels, and the decoder generates a final segmentation result through channel splicing and up-sampling. According to the fundus image lesion segmentation method and device, the computer equipment and the medium, four kinds of sugar net disease lesions in the fundus image can be segmented at the same time, calculation parameters are few, and good accuracy and robustness are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and in particular to a method, device, computer device and medium for segmenting lesions in fundus images. Background Art

[0002] Diabetic Retinopathy is the most common microvascular complication of diabetes. With the development of social economy and the change of people's lifestyle, the number of diabetic patients is increasing day by day, and the incidence of Diabetic Retinopathy is rising year by year. Diabetic Retinopathy has no obvious symptoms in the early stage. However, with the development of the disease, symptoms such as blurred vision become more obvious, and blindness may occur in severe cases. Early detection of Diabetic Retinopathy is of great significance for timely treatment and alleviating the pain of patients.

[0003] The main pathological features of Diabetic Retinopathy include Microaneurysms (MA), Hemorrhages (HE), Soft Exudates (SE) and Hard Exudates (EX). Different types of lesions will appear in different stages of Diabetic Retinopathy. Usually, ophthalmologists manually mark the type, size, distribution and quantity of lesions to judge the progress of the disease. However, this method highly depends on expert experience, is costly and limited by medical resources, and it is difficult to achieve large-scale popularization.

[0004] Therefore, it is necessary to design an automatic and accurate method for segmenting diabetic retinopathy lesions to relieve the pressure on doctors. However, the existing methods for segmenting diabetic retinopathy lesions based on fundus color photos are difficult to achieve simultaneous segmentation of multiple types of lesions, and the accuracy and robustness need to be improved. Summary of the Invention

[0005] The purpose of the present invention is to provide a method, device, computer device and medium for segmenting lesions in fundus images, which can realize automatic and accurate segmentation of Diabetic Retinopathy regions through deep learning technology, effectively improve the accuracy of lesion detection, and reduce the workload of doctors.

[0006] To achieve the above purpose, the present invention provides a method for segmenting lesions in fundus images, including the following steps:

[0007] Preprocess the original fundus image to obtain a preprocessed image;

[0008] Construct a fundus image lesion segmentation model, which includes an encoder, a global-local attention module, a multi-scale feature capture module and a decoder;

[0009] The fundus image lesion segmentation model is used to extract features and segment the preprocessed image, including:

[0010] The encoder extracts the features of the input image through multi-head self-attention mechanism, depth-separable convolution and pooling operations to obtain encoder features;

[0011] The global-local attention module combines global attention and local attention features to enhance feature representation and obtain global-local attention features;

[0012] The multi-scale feature capture module extracts multi-scale lesion features through convolution operations of convolution kernels of different scales to obtain multi-scale features;

[0013] The decoder concatenates and upsamples the encoder features, global local attention features, and multi-scale features to generate the final lesion segmentation result;

[0014] The Dice loss function is used to guide model training, and the model weights are optimized through back propagation.

[0015] Preferably, the encoder comprises a left branch, a middle branch and a right branch, wherein:

[0016] The left branch uses linear layer, depthwise separable convolution, linear layer and maximum pooling to extract high-frequency information and obtain the left branch features;

[0017] The middle branch extracts the global context information of the image through average pooling, multi-head attention and upsampling to obtain the middle branch features;

[0018] The right branch extracts low-frequency information through maximum pooling and linear operations to obtain the right branch features;

[0019] The left branch features, middle branch features, and right branch features are concatenated in the channel direction and downsampled layer by layer to obtain the encoder features.

[0020] Preferably, the global local attention module includes:

[0021] Global attention module: It consists of a global channel attention submodule and a global spatial attention submodule;

[0022] The global channel attention submodule is used to capture the global contextual information of the image and generate weighted global channel attention features;

[0023] The global spatial attention submodule is used to generate weighted global spatial attention features by combining channel information and spatial information;

[0024] Multi-scale local attention module: extracts multi-scale features through convolution kernels of different scales and generates three weighted local attention features;

[0025] The weighted global channel attention feature and the weighted global spatial attention feature are added element-wise to obtain the global attention feature;

[0026] The three weighted local attention features are fused by addition to generate the multi-scale local attention feature.

[0027] Preferably, during the weighted fusion process, the global attention feature, the multi-scale local attention feature and the encoder feature are added with weights through skip connections to obtain the global-local attention feature.

[0028] Preferably, the multi-scale feature capture module includes a first branch, a second branch and a third branch, where:

[0029] The first branch extracts multi-scale lesion features through 1×1 convolution, 3×3 convolution and 1×1 convolution to obtain the first branch feature;

[0030] The second branch extracts multi-scale lesion features through 1×1 convolution, 5×5 convolution and 1×1 convolution to obtain the second branch feature;

[0031] The third branch extracts multi-scale lesion features through 1×1 convolution, 7×7 convolution and 1×1 convolution to obtain the third branch feature;

[0032] The first branch feature, the second branch feature and the third branch feature are added and fused element-wise to obtain the multi-scale feature.

[0033] Preferably, the decoder includes:

[0034] Channel concatenation: concatenate the encoder feature, the global-local attention feature and the multi-scale feature;

[0035] Layer-by-layer upsampling: gradually restore the resolution of the feature map to the size of the preprocessed image through multiple layers of convolution;

[0036] Output the segmentation result: output the final lesion segmentation result through a 1×1 convolutional layer.

[0037] The present invention also provides a fundus image lesion segmentation device, including:

[0038] A data preprocessing module for preprocessing the original fundus image to obtain the preprocessed image;

[0039] A model construction module for constructing a fundus image lesion segmentation model, where the fundus image lesion segmentation model includes an encoder, a global-local attention module, a multi-scale feature capture module and a decoder;

[0040] A feature extraction and segmentation module for extracting features and segmenting the preprocessed image by using the fundus image lesion segmentation model, including:

[0041] The encoder extracts the features of the input image through the multi-head self-attention mechanism, depthwise separable convolution, and pooling operations to obtain encoder features;

[0042] The global-local attention module combines global attention and local attention features to enhance the feature representation and obtain global-local attention features;

[0043] The multi-scale feature capture module extracts multi-scale lesion features through convolutional operations with different-scale convolutional kernels to obtain multi-scale features;

[0044] The decoder performs channel concatenation and upsampling on the encoder features, global-local attention features, and multi-scale features to generate the final lesion segmentation result;

[0045] The model training module is used to guide the model training using the Dice loss function and optimize the model weights through backpropagation.

[0046] The present invention also provides a computer device, including a memory and a processor. The memory is used to store instructions, and the processor is used to execute the instructions to implement the fundus image lesion segmentation method as described above.

[0047] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the fundus image lesion segmentation method as described above is implemented.

[0048] Therefore, the present invention adopts the above-mentioned fundus image lesion segmentation method, device, computer device, and medium, and the beneficial technical effects are as follows:

[0049] (1) High-precision lesion segmentation: Channel fusion is performed on multi-scale lesion features and global-local attention features at different levels, and key attention is paid to the channel information related to lesion segmentation, which helps to accurately identify and segment the diabetic retinopathy area.

[0050] (2) Enhanced model generalization ability: The training data set is expanded through data augmentation techniques (such as rotation, scaling, translation, etc.) to improve the adaptability of the model to different diabetic retinopathy images and avoid overfitting.

[0051] (3) Automated segmentation and high efficiency: The present invention realizes the automated segmentation of diabetic retinopathy, reduces manual intervention, improves the work efficiency of doctors, and has broad clinical application potential. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 is a flowchart of a fundus image lesion segmentation method of the present invention;

[0053] Figure 2It is the structural diagram of the fundus image lesion segmentation model;

[0054] Figure 3 It is the structural diagram of the encoder;

[0055] Figure 4 It is the structural diagram of the global-local attention module;

[0056] Figure 5 It is the structural diagram of the multi-scale feature capture module;

[0057] Figure 6 It is the structural diagram of the decoder;

[0058] Figure 7 They are the segmentation results of four kinds of lesions, namely hard exudates, soft exudates, microaneurysms and hemorrhages, on the IDRiD dataset; among them, Figure 7 (a) in it is the original fundus image; Figure 7 (b) in it is the fundus image lesion label; Figure 7 (c) in it is the lesion segmentation result obtained by the present invention. Detailed implementation manners

[0059] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0060] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the field to which the present invention belongs.

[0061] Embodiment 1

[0062] As Figure 1 shown, it is the flowchart of a method for segmenting fundus image lesions according to the present invention, which specifically includes the following steps:

[0063] Step 1, preprocess the original fundus image to enhance the image quality.

[0064] 1) Use the histogram equalization technology with enhanced contrast to improve the contrast of the image, making the details in the image more obvious, which helps the subsequent model to detect the lesion area.

[0065] 2) Apply the non-local means denoising method to remove the noise irrelevant to the lesion segmentation and retain the important structural features, thereby reducing the interference of the adverse noise on the accurate segmentation.

[0066] 3) Crop the entire contrast-enhanced and denoised fundus image into block images with a resolution of 224×224 to meet the network input requirements, and ensure that each image block can contain sufficient lesion information to provide high-quality input data for model training.

[0067] As Figure 2As shown in the figure, it is a structural diagram of the fundus image lesion segmentation model, and its specific process is as follows:

[0068] Step 2: Input the preprocessed image X into the encoder to extract features.

[0069] like Figure 3 As shown, first, the original fundus image is divided into three single-channel images of the same size and respectively input into the encoder (the encoder is a three-branch encoder) to extract features.

[0070] The left branch uses linear layer, depthwise separable convolution, linear layer and maximum pooling to extract high-frequency information and obtain the left branch features; the right branch uses maximum pooling and linear operations to calculate low-frequency information and obtain the right branch features. The specific calculation process is as follows:

[0071] Y,Z,J=split(X);

[0072] G=MaxPool(Linear(DWconv(Linear(Y))));

[0073] B = Linear(MaxPool(J));

[0074] Among them, Y, Z, and J represent the red, green, and blue channel images extracted from the original fundus image, respectively, G represents the left branch feature, B represents the right branch feature, split represents the channel partition function, Linear represents the linear layer, DWconv represents the depthwise separable convolution, and MaxPool represents the maximum pooling.

[0075] The middle branch performs average pooling, multi-head attention, and upsampling in turn to extract the global context information of the image and obtain the middle branch features. The specific calculation process is as follows:

[0076] M=AvgPool(MSA(Upsample(Z)));

[0077]

[0078] Among them, M represents the intermediate branch feature, AvgPool represents average pooling, MSA represents the multi-head attention mechanism, Upsample represents the upsampling operation, Q, K, and V represent the query vector, key vector, and value vector respectively, Softmax represents probabilistic, and d k represents the dimension of multi-head attention and T represents transposition.

[0079] Finally, the left branch features, the middle branch features, and the right branch features are concatenated in the channel direction and restored to the original image size, and then downsampled layer by layer.

[0080] Fi = Concat(G, B, M);

[0081] F i+1 = Downsample(F i );

[0082] Where F i represents the encoder feature of the current level, and F i+1 represents the encoder feature of the next level. Downsample represents the downsampling operation, Concat represents the channel concatenation operation, and i represents the index variable, where i ∈ {1, 2, 3, 4, 5}.

[0083] Step 3: As Figure 4 shown, the global-local attention module uses the global attention module and the multi-scale local attention module to capture high-frequency local edge detail information and low-frequency global information respectively, and performs weighted residual connection with the encoder feature to obtain the global-local attention feature.

[0084] Step 4: The global attention module consists of a global channel attention sub-module and a global spatial attention sub-module, which interact the channel and spatial information to generate the global channel attention map GC ∈ R C×C and the global spatial attention map GS ∈ R HW×HW , where R represents the real number field.

[0085] Multiply the global channel attention map GC ∈ R C×C element-wise with the encoder feature F i and then add it to the encoder feature F i to obtain the weighted global channel attention feature GCF. The calculation process of the global channel attention sub-module is shown as follows:

[0086]

[0087] Where Transpose represents the matrix transpose operation, σ represents the sigmoid activation function, represents the element-wise multiplication operation, represents the element-wise addition operation, and H, W, and C represent the height, width, and number of channels of the feature map respectively.

[0088] Next, perform a linear operation on the weighted global channel attention feature GCF to obtain the feature matrices Q, K, V ∈ R C×HW . Then transpose the query matrix Q and multiply it with the key matrix K and probabilize it to obtain the global spatial attention feature GS ∈ R HW×HW , and multiply the global spatial attention feature with the value matrix V to obtain the weighted global spatial attention feature GSF ∈ R C×HWFinally, the weighted global channel attention feature GCF and the weighted global spatial attention feature GSF are added element-wise to obtain the global attention feature GF g ∈R C×HW 。

[0089] Q, K, V = Linear(GCF);

[0090]

[0091] Step 6: The multi-scale local attention module first uses 3×3, 5×5, and 7×7 convolutional kernels to adjust different receptive fields to extract the multi-scale lesion features L i 1, L i 2, L i 3 ∈ R C×W×H , and adds the multi-scale features to obtain the multi-scale fusion feature FL i 。 Then, the global average pooling operation is used to compress FL i into Z i ∈ R C×1×1 , and the channel attention feature Z i ∈ R C×1×1 is linearly operated and mapped into the fine-grained channel attention feature Z' i ∈ R C'×1×1 , where C' < C, and C' represents the number of channels of the feature map. Then, Z' i ∈ R C'×1×1 is transformed into the local attention feature Then, the multi-scale lesion features L i 1, L i 2, L i 3 ∈ R C×W×H are multiplied element-wise with the local attention feature to obtain the weighted local attention feature Finally, the three weighted local attention features are additively fused to generate the multi-scale local attention feature LF l ∈ R C×HW 。 The calculation process of the multi-scale local attention module can be expressed as:

[0092] L i 1 = Conv 3×3 (F i ), L i 2 = Conv 5×5 (F i ), L i 3 = Conv 7×7 (F i );

[0093] FL i = L i1 + L i 2 + L i 3;

[0094] Z i = AvgPool(FL i );

[0095] Z i ' = Linear(Z i );

[0096]

[0097] Among them, AvgPool represents global average pooling, Reshape represents matrix shape reshaping, Conv represents convolution operation, and the subscript of Conv represents the convolution kernel size.

[0098] Step 7: Weightedly add the global attention feature, multi-scale local attention feature, and encoder feature through skip connections to obtain the global-local attention feature GLF i ∈R C×W×H , and the specific calculation process is as follows:

[0099]

[0100] Among them, μ e , μ g , and μ l all represent trainable weight coefficients, and μ e , μ g , and μ l are all set to 1.

[0101] Step 8: As Figure 5 shown, the multi-scale feature capture module extracts multi-scale lesion features MO i using different convolution operations and performs element-wise addition and fusion.

[0102] First, each input global-local attention feature GLF i goes through three-branch convolution operations respectively, namely the first branch, the second branch, and the third branch.

[0103] The first branch extracts multi-scale lesion features through 1×1 convolution, 3×3 convolution, and 1×1 convolution to obtain the first branch feature;

[0104] The second branch extracts multi-scale lesion features through 1×1 convolution, 5×5 convolution, and 1×1 convolution to obtain the second branch feature;

[0105] The third branch extracts multi-scale lesion features through 1×1 convolution, 7×7 convolution, and 1×1 convolution to obtain the third branch feature;

[0106] Convolution operations at different scales can capture multi-level features, which helps to promote the effective fusion of low-level edge details and high-level semantic information. The detailed calculation process of this module is as follows:

[0107] GLF1, GLF2, GLF3 = split(GLF i );

[0108]

[0109] Among them, GFL1, GFL2, and GFL3 respectively represent the feature maps of the red, green, and blue channels extracted from the global-local feature map, O1, O2, and O3 respectively represent the features extracted from the first, second, and third branches, and MO i represents the output result of the multi-scale feature capture module, that is, the multi-scale lesion feature, and ReLU represents the rectified linear activation function.

[0110] Step 9, as Figure 6 shown, input the decoder features of the previous stage and the multi-scale lesion feature MO i into the decoder for channel concatenation and upsampling to obtain the final lesion segmentation result. The decoder uses average pooling, linear layer, activation layer, and probability operation to calculate the attention weights of each channel, multiplies the channel attention weights element-wise with the fusion result of the encoder features and the multi-scale lesion features, focuses on the feature channels related to lesion segmentation, and suppresses the information of irrelevant feature channels. Then, the feature map is upsampled layer by layer, gradually restoring the low-resolution feature map to a higher spatial resolution, making the details in the image clearer. Finally, the decoder feature map passes through a 1×1 convolutional layer to output the diabetic retinopathy lesion segmentation result, accurately segmenting the lesion area.

[0111] D = Sigmoid(Linear(Activate(LN(AvgPool(Concat(F i , MO i ));

[0112] E i = Conv 3×3 (ReLU(Conv 3×3 (ReLU(Conv 3×3 (Concat(F i , MO i ))))));

[0113] Output = Conv 1×1 (E5);

[0114] Among them, D represents the channel attention weight, Activate represents the activation layer, LN represents layer normalization, E i represents the output of each layer of the decoder, and Output represents the final output segmentation result.

[0115] Step 10: Use the Dice loss as the loss function to guide the model training. The calculation formula is as follows:

[0116]

[0117] Among them, U and Y respectively represent the lesion label and the predicted segmentation result obtained by the model. |U| and |Y| respectively represent the number of pixels in the lesion label and the segmentation prediction result. N represents the total number of pixels in the fundus image. y j and respectively represent the label value and the predicted value of pixel j.

[0118] This embodiment conducts experimental verification on the internationally public IDRiD dataset. The model designed by the present invention is trained on the IDRiD training set and the best weights are saved, and the segmentation performance is verified on the test set. Table 1 shows the comparison results of the four lesion segmentations of the present invention, Unet, ResUnet, and Unet++ on the IDRiD dataset. Figure 7 are the four lesion segmentation results obtained by the present invention on the IDRiD dataset. As Figure 7 shown, the present invention can accurately segment the four lesions of diabetic retinopathy, and the segmentation result is close to the manual annotation.

[0119] Table 1 Comparison Results

[0120]

[0121] Embodiment 2

[0122] A fundus image lesion segmentation device, comprising:

[0123] A data preprocessing module for preprocessing the original fundus image to obtain a preprocessed image;

[0124] A model construction module for constructing a fundus image lesion segmentation model, which includes an encoder, a global-local attention module, a multi-scale feature capture module, and a decoder;

[0125] A feature extraction and segmentation module for extracting features and segmenting the preprocessed image by using the fundus image lesion segmentation model, including:

[0126] The encoder extracts the features of the input image through the multi-head self-attention mechanism, depthwise separable convolution, and pooling operations to obtain encoder features;

[0127] The global-local attention module combines global attention and local attention features to enhance feature representation and obtain global-local attention features;

[0128] The multi-scale feature capture module extracts multi-scale lesion features through convolutional operations with convolutional kernels of different scales to obtain multi-scale image features;

[0129] The decoder performs channel concatenation and upsampling on the encoder features, global-local attention features, and multi-scale features to generate the final lesion segmentation result;

[0130] The model training module is used to guide the model training using the Dice loss function and optimize the model weights through backpropagation.

[0131] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, and other media that can store program codes.

[0132] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0133] More specific examples (nonexhaustive list) of computer-readable media include the following: electrical connections (electronic devices) having one or more wirings, portable computer diskettes (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber devices, and portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.

[0134] It should be noted that the content not elaborated in detail in the present invention is all prior art and well-known to those skilled in the art.

[0135] Therefore, by adopting the above method, device, computer equipment and medium for fundus image lesion segmentation, the present invention can segment four diabetic retinopathy lesions in fundus images simultaneously with fewer calculation parameters, and has good accuracy and robustness.

[0136] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for segmenting lesions in fundus images, characterized in that, It includes the following steps: Preprocess the original fundus image to obtain the preprocessed image; Construct a fundus image lesion segmentation model, which includes an encoder, a global-local attention module, a multi-scale feature capture module, and a decoder; Use the fundus image lesion segmentation model to extract features and segment the preprocessed image, including: The encoder extracts the features of the input image through the multi-head self-attention mechanism, depthwise separable convolution, and pooling operations to obtain the encoder features; The global-local attention module combines global attention and local attention features to enhance the feature representation and obtain the global-local attention features; The multi-scale feature capture module extracts multi-scale lesion features through convolution operations with different scale convolutional kernels to obtain multi-scale features; The decoder performs channel concatenation and upsampling on the encoder features, global-local attention features, and multi-scale features to generate the final lesion segmentation result; Use the Dice loss function to guide the model training and optimize the model weights through backpropagation.

2. The method for segmenting fundus image lesions according to claim 1, wherein The encoder includes a left branch, a middle branch, and a right branch, where: The left branch uses a linear layer, depthwise separable convolution, a linear layer, and max pooling to extract high-frequency information and obtain the left branch features; The middle branch extracts the global context information of the image through average pooling, multi-head attention, and upsampling to obtain the middle branch features; The right branch extracts low-frequency information through max pooling and linear operations to obtain the right branch features; The left branch features, middle branch features, and right branch features are concatenated in the channel direction and downsampled layer by layer to obtain the encoder features.

3. The fundus image lesion segmentation method according to claim 2, characterized in that The global-local attention module includes: The global attention module: consists of a global channel attention sub-module and a global spatial attention sub-module; The global channel attention sub-module is used to capture the global context information of the image and generate weighted global channel attention features; The global spatial attention sub-module is used to combine channel information and spatial information to generate weighted global spatial attention features; The multi-scale local attention module: extracts multi-scale features through different scale convolutional kernels and generates three weighted local attention features; The weighted global channel attention features and the weighted global spatial attention features are element-wise added to obtain the global attention features; The three weighted local attention features are added and fused to generate multi-scale local attention features.

4. The fundus image lesion segmentation method according to claim 3, characterized in that, During the weighted fusion process, the global attention features, multi-scale local attention features, and encoder features are weighted and added through skip connections to obtain the global-local attention features.

5. A method for segmenting fundus image lesions according to claim 4, wherein The multi-scale feature capture module includes a first branch, a second branch, and a third branch, where: The first branch extracts multi-scale lesion features through 1×1 convolution, 3×3 convolution, and 1×1 convolution to obtain the first branch features; The second branch extracts multi-scale lesion features through 1×1 convolution, 5×5 convolution, and 1×1 convolution to obtain the second branch features; The third branch extracts multi-scale lesion features through 1×1 convolution, 7×7 convolution, and 1×1 convolution to obtain the third branch features; The first branch features, second branch features, and third branch features are element-wise added and fused to obtain multi-scale features.

6. The method for segmenting fundus image lesions according to claim 5, wherein The decoder includes: Channel concatenation: Concatenate the encoder features, global-local attention features, and multi-scale features in channels; Layer-by-layer upsampling: Gradually restore the resolution of the feature map to the size of the preprocessed image through multiple layers of convolution; Output the segmentation result: Output the final lesion segmentation result through a 1×1 convolutional layer.

7. An apparatus for segmenting lesions in fundus images, characterized in that, It includes: A data preprocessing module for preprocessing the original fundus image to obtain the preprocessed image; A model construction module for constructing a fundus image lesion segmentation model, where the fundus image lesion segmentation model includes an encoder, a global-local attention module, a multi-scale feature capture module, and a decoder; A feature extraction and segmentation module for using the fundus image lesion segmentation model to extract features and segment the preprocessed image, including: The encoder extracts the features of the input image through the multi-head self-attention mechanism, depthwise separable convolution, and pooling operations to obtain encoder features; The global-local attention module combines global attention and local attention features to enhance the feature representation and obtain global-local attention features; The multi-scale feature capture module extracts multi-scale lesion features through convolution operations with different-scale convolutional kernels to obtain multi-scale features; The decoder concatenates and upsamples the encoder features, global-local attention features, and multi-scale features to generate the final lesion segmentation result; A model training module for guiding the model training using the Dice loss function and optimizing the model weights through backpropagation.

8. A computer device, characterized in that, It includes a memory and a processor, where the memory is used to store instructions, and the processor is used to execute the instructions to implement the fundus image lesion segmentation method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the fundus image lesion segmentation method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Mine image super-resolution reconstruction method and system based on multi-scale residual network

    CN113592718A

  • Deep learning-based skin lesion image segmentation method and device, and storage medium

    CN114066904A

  • Intracranial aneurysm image detection method, system, equipment and medium

    CN117392137A

  • Multi-modal target detection method and device based on feature enhancement and collaborative interaction

    CN117423007A

  • Heart MRI image segmentation method based on edge feature enhancement

    CN117635942A