Retinal lesion segmentation method and device, electronic device and storage medium
Through the method of multi-stage encoding and decoding combined with feature fusion and attention processing, the problem of inaccurate retinal lesion segmentation in OCT images is solved, more efficient retinal lesion segmentation is achieved, and the accuracy and completeness of lesion detection are improved.
Patent Information
- Application Number
- CN202311310160.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-10
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2043-10-10
AI Technical Summary
In the existing technology, the segmentation of retinal lesions in OCT images is inaccurate and it is easy to miss tiny lesion structures. In addition, the automatic segmentation method is incomplete, making it difficult to achieve efficient retinal lesion segmentation.
A multi-stage encoding and decoding method is adopted, combined with feature fusion and attention processing. Multi-scale features are extracted through multi-stage encoding of retinal images, and feature splicing and attention processing are performed to improve the accuracy of lesion segmentation.
It improves the accuracy and completeness of retinal lesion segmentation, can effectively detect multi-scale lesions, reduce omissions, and improve the effect of automatic segmentation.
Smart Images

Figure CN117893745B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of medical image processing, and in particular to a retinal lesion segmentation method and device, an electronic device, and a storage medium. Background Art
[0002] Optical Coherence Tomography (OCT) is a non-contact imaging technology with micron-level high resolution that has been widely used in the diagnosis of retinal diseases. In related technologies, retinal lesions are manually segmented, or computer-assisted methods are used to automatically segment retinal lesions. However, manual segmentation methods are not only time-consuming and labor-intensive, but also easily introduce subjective factors, resulting in errors in the diagnosis of retinal diseases. OCT images have problems such as a lot of noise, large differences in lesion size, and complex lesion shape contours, which make the automatically segmented lesions incomplete and inaccurate, and it is easy to miss tiny lesion structures. How to automatically and accurately segment retinal lesions has become an urgent problem to be solved. Summary of the Invention
[0003] The main purpose of the embodiments of the present application is to provide a retinal lesion segmentation method and device, electronic device and storage medium, aiming to improve the accuracy of retinal lesion segmentation.
[0004] To achieve the above objectives, a first aspect of an embodiment of the present application provides a retinal lesion segmentation method, the method comprising:
[0005] Acquire retinal images;
[0006] Performing multi-stage encoding on the retinal image to obtain preliminary retinal coding features; the preliminary retinal coding features include first-stage coding features and second-stage coding features; the first-stage coding features are coding features of a first preset coding stage; the second-stage coding features are coding features of a second preset coding stage; the first preset stage and the second preset stage are adjacent;
[0007] Performing feature fusion on the first-stage coding features and the second-stage coding features to obtain target retinal coding features;
[0008] performing attention processing on the target retinal encoding feature to obtain a target retinal attention feature;
[0009] Performing multi-stage decoding on the second-stage encoding features to obtain preliminary retinal decoding features; the preliminary retinal decoding features include target-stage decoding features of a target decoding stage, and the target decoding stage matches the first preset encoding stage;
[0010] Performing feature concatenation on the target retinal attention feature and the target stage decoding feature from a channel dimension to obtain a target retinal decoding feature;
[0011] The retinal image is segmented for lesions according to the target retinal decoding features to obtain a retinal lesion category and a retinal lesion location.
[0012] In some embodiments, performing multi-stage encoding on the retinal image to obtain preliminary retinal coding features includes:
[0013] Performing first-stage encoding on the retinal image to obtain first-stage encoding features;
[0014] Performing second-stage encoding on the first-stage encoding features to obtain the second-stage encoding features.
[0015] In some embodiments, performing first-stage encoding on the retinal image to obtain the first-stage encoding features includes:
[0016] performing a first feature extraction on the retinal image to obtain a first retinal feature map;
[0017] performing a second feature extraction on the first retinal feature map to obtain a second retinal feature map;
[0018] performing feature fusion on the first retinal feature map and the second retinal feature map to obtain a fused retinal feature map;
[0019] performing regional perception attention processing on the fused retinal feature map to obtain regional perception attention;
[0020] Window attention processing is performed on the regional perception attention to obtain the first-stage encoding features.
[0021] In some embodiments, performing regional-aware attention processing on the fused retinal feature map to obtain regional-aware attention includes:
[0022] Acquire a first channel feature and a second channel feature of the fused retinal feature map; the first channel feature is a channel maximum value of the fused retinal feature map; the second channel feature is a channel average value of the fused retinal feature map;
[0023] Concatenate the first channel feature and the second channel feature from the channel dimension to obtain a target channel feature;
[0024] The regional perception attention is transformed by performing a regional perception attention transformation on the fused retinal feature map according to a preset regional attention matrix and the target channel characteristics to obtain the regional perception attention.
[0025] In some embodiments, performing window attention processing on the region-aware attention to obtain the first-stage encoding features includes:
[0026] Dividing the regional perception attention into windows to obtain multiple preliminary window attentions;
[0027] Performing multi-head attention processing on each of the preliminary window attentions to obtain candidate window attentions corresponding to each of the preliminary window attentions;
[0028] The attention of each candidate window is concatenated to obtain the encoding features of the first stage.
[0029] In some embodiments, performing attention processing on the target retinal encoding feature to obtain the target retinal attention feature includes:
[0030] Performing average pooling processing on the target retinal coding feature to obtain a first pooling feature;
[0031] Performing maximum pooling processing on the target retinal coding feature to obtain a second pooling feature;
[0032] Performing attention calculation based on the first pooled features and the second pooled features to obtain a target attention matrix;
[0033] Performing attention processing on the target retinal encoding feature according to the target attention matrix to obtain a preliminary retinal attention feature;
[0034] The preliminary retinal attention feature and the first-stage encoding feature are feature concatenated from the channel dimension to obtain the target retinal attention feature.
[0035] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application provides a retinal lesion segmentation device, comprising:
[0036] An acquisition module, used for acquiring retinal images;
[0037] a multi-stage encoding module, configured to perform multi-stage encoding on the retinal image to obtain preliminary retinal coding features; the preliminary retinal coding features include first-stage coding features and second-stage coding features; the first-stage coding features are coding features of a first preset coding stage; the second-stage coding features are coding features of a second preset coding stage; the first preset stage and the second preset stage are adjacent;
[0038] a feature fusion module, configured to fuse the first-stage coding features and the second-stage coding features to obtain target retinal coding features;
[0039] an attention processing module, configured to perform attention processing on the target retinal encoding feature to obtain a target retinal attention feature;
[0040] a multi-stage decoding module, configured to perform multi-stage decoding on the second-stage encoding features to obtain preliminary retinal decoding features; the preliminary retinal decoding features include target-stage decoding features of a target decoding stage, the target decoding stage matching the first preset encoding stage;
[0041] a feature splicing module, configured to perform feature splicing on the target retinal attention feature and the target stage decoding feature from a channel dimension to obtain a target retinal decoding feature;
[0042] The lesion segmentation module is used to perform lesion segmentation on the retinal image according to the target retinal decoding features to obtain the retinal lesion category and retinal lesion location.
[0043] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the retinal lesion segmentation method described in the first aspect.
[0044] To achieve the above objectives, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the retinal lesion segmentation method described in the first aspect.
[0045] The retinal lesion segmentation method, retinal lesion segmentation device, electronic device and storage medium proposed in this application obtain retinal images, perform multi-stage encoding on the retinal images, and obtain preliminary retinal coding features. Multi-stage encoding can obtain retinal features of different scales in multiple coding stages to comprehensively extract the features of the retinal image, improve the segmentation ability of multi-scale lesions, and avoid missing tiny lesion structures. By fusing the first-stage coding features and the second-stage coding features, it is possible to establish a connection between adjacent scale features, fuse multi-scale features, and retain the contextual information of the features to obtain target retinal coding features. In order to extract features important for lesion segmentation from the target retinal coding features obtained by fusion of adjacent stages, attention processing is performed on the target retinal coding features to obtain target retinal attention features. By performing multi-stage decoding on the second-stage coding features, high-resolution multi-scale features can be extracted to obtain preliminary retinal decoding features. By performing feature splicing on the target retinal attention features and the target stage decoding features from the channel dimension, it is possible to combine high-resolution features to supplement the extraction of local detail information and obtain target retinal decoding features. The retinal image is segmented for lesions according to the target retinal decoding features to obtain the retinal lesion category and retinal lesion location, which improves the segmentation capability of multi-scale lesions and thus improves the accuracy of retinal lesion segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is a flow chart of the retinal lesion segmentation method provided in an embodiment of the present application;
[0047] Figure 2 yes Figure 1 Flowchart of step S120 in FIG.
[0048] Figure 3 This is the network structure diagram of the retinal lesion segmentation model;
[0049] Figure 4 yes Figure 2 Flowchart of step S210 in FIG.
[0050] Figure 5 yes Figure 4 Flowchart of step S440 in FIG.
[0051] Figure 6 yes Figure 5 Flowchart of step S530 in FIG.
[0052] Figure 7 yes Figure 4 Flowchart of step S450 in FIG.
[0053] Figure 8 yes Figure 1 Flowchart of step S140 in FIG.
[0054] Figure 9 is a schematic diagram of feature fusion provided in an embodiment of the present application;
[0055] Figure 10a is a retinal image provided by an embodiment of the present application;
[0056] Figure 10b The embodiment of this application provides Figure 10a The result of retinal lesion segmentation on the retinal image;
[0057] Figure 11 is a schematic structural diagram of a retinal lesion segmentation device provided in an embodiment of the present application;
[0058] Figure 12 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0059] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0060] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.
[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0062] In recent years, artificial intelligence technology has developed rapidly and has been widely used in daily production and life, especially in the field of medicine, where many computer-assisted diagnosis methods and means have emerged.
[0063] The macula, located in the center of the retina, is responsible for vision and color perception. Macular edema is caused by fluid accumulation from damaged retinal blood vessels, leading to swelling of parts of the retina. This condition is often caused by retinal diseases such as age-related macular degeneration (AMD), retinal vein occlusion (RVO), or diabetic macular edema (DME).
[0064] Optical coherence tomography (OCT) is a non-contact imaging technology with micron-level high resolution. OCT has been widely used in the diagnosis of retinal diseases. Accurate segmentation and quantitative analysis of macular retinal fluid are necessary to more accurately diagnose retinal diseases, develop personalized treatment plans for patients, and evaluate treatment efficacy. Retinal fluid in the macular region primarily includes intraretinal fluid (IRF), subretinal fluid (SRF), and pigment epithelial detachment (PED). Manual segmentation of retinal fluid is not only time-consuming and labor-intensive, but also prone to subjective factors and errors. Therefore, the development of computer-assisted automatic segmentation methods for retinal lesions is particularly important.
[0065] OCT images contain excessive noise, large variations in lesion size, and complex lesion shapes and contours. This makes automatic segmentation of lesions incomplete and inaccurate, and can easily miss tiny lesion structures, making computer-assisted automatic segmentation difficult. Automatically and accurately segmenting retinal lesions has become a pressing issue.
[0066] Based on this, the embodiments of the present application provide a retinal lesion segmentation method, a retinal lesion segmentation device, an electronic device and a computer-readable storage medium, aiming to improve the accuracy of retinal lesion segmentation.
[0067] The retinal lesion segmentation method, retinal lesion segmentation device, electronic device and computer-readable storage medium provided in the embodiments of the present application are specifically illustrated by the following embodiments. First, the retinal lesion segmentation method in the embodiments of the present application is described.
[0068] The retinal lesion segmentation method provided in the embodiment of the present application relates to the field of medical image processing technology. The retinal lesion segmentation method provided in the embodiment of the present application can be applied to a terminal, can also be applied to a server side, and can also be software running in a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, and can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the retinal lesion segmentation method, etc., but is not limited to the above forms.
[0069] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0070] Figure 1 This is an optional flowchart of the retinal lesion segmentation method provided in the embodiment of the present application. Figure 1 The method may include but is not limited to steps S110 to S170.
[0071] Step S110, acquiring a retinal image;
[0072] Step S120, performing multi-stage encoding on the retinal image to obtain preliminary retinal coding features; the preliminary retinal coding features include first-stage coding features and second-stage coding features; the first-stage coding features are coding features of the first preset coding stage; the second-stage coding features are coding features of the second preset coding stage; the first preset stage and the second preset stage are adjacent;
[0073] Step S130, performing feature fusion on the first-stage coding features and the second-stage coding features to obtain target retinal coding features;
[0074] Step S140, performing attention processing on the target retinal coding feature to obtain the target retinal attention feature;
[0075] Step S150, performing multi-stage decoding on the second-stage coding features to obtain preliminary retinal decoding features; the preliminary retinal decoding features include target-stage decoding features of a target decoding stage, and the target decoding stage matches the first preset coding stage;
[0076] Step S160, performing feature concatenation on the target retinal attention feature and the target stage decoding feature from the channel dimension to obtain the target retinal decoding feature;
[0077] Step S170 , performing lesion segmentation on the retinal image according to the target retinal decoding features to obtain the retinal lesion category and retinal lesion location.
[0078] In step S110 of some embodiments, a retinal image is obtained from an OCT retinal lesion segmentation dataset. Specifically, an OCT retinal lesion segmentation dataset is selected and preprocessed, the OCT retinal lesion segmentation dataset includes multiple retinal images and multiple mask images, the retinal image is matched with the mask image, the mask image and the retinal image are made to correspond one to one, and they are synchronously divided into training data and test data. The mask image contains segmentation information of multiple categories of retinal fluid lesions, and the retinal fluid lesions include three categories: intraretinal fluid, subretinal fluid and pigment epithelial detachment. The retinal fluid lesion segmentation information includes the location, size and area information of the three lesions. The OCT retinal lesion segmentation dataset can be a public dataset RETOUCH, which includes 6732 OCT retinal images and their corresponding mask images, and the segmentation information of the retinal fluid lesions is marked in the mask image. It is understandable that not every mask image is marked with three lesions. Based on the presence of lesions in the retinal image, the mask image can be divided into four categories. The first category contains no lesions, indicating that the image represents a healthy retina. The second category contains one lesion, the third category contains two lesions, and the fourth category contains all three lesions. To enable the retinal lesion segmentation model to simultaneously segment multiple retinal fluid lesions, the mask image uses three pixel values to annotate the location, size, and area of specific fluid lesions. Intraretinal fluid is represented by pixel 1, subretinal fluid by pixel 2, pigment epithelial detachment by pixel 3, and background by pixel 0. The RETOUCH dataset is randomly partitioned according to a certain ratio to generate training and test data. For example, 80% of the data can be used as training data and 20% as testing data, resulting in 5386 images as training data and 1346 images as testing data.
[0079] Before model training and testing, the OCT retinal lesion segmentation dataset requires data processing and augmentation. Specifically, the retinal image and mask image are scaled from their original size to 512x512. During the training phase, data augmentation operations such as random horizontal flipping and random rotation can be performed on the retinal image.
[0080] See also Figure 2 In some embodiments, step S120 may include but is not limited to steps S210 to S220:
[0081] Step S210, performing first-stage encoding on the retinal image to obtain first-stage encoding features;
[0082] Step S220 , performing second-stage encoding on the first-stage encoding features to obtain second-stage encoding features.
[0083] In step S210 of some embodiments, a retinal lesion segmentation model is constructed for the OCT image. Figure 3 As shown in the figure, the retinal lesion segmentation model includes an encoder, a decoder, a bottleneck layer, and an attention module (Focused Attention, FA). The retinal lesion segmentation model is a U-shaped network model. The encoder is divided into four stages, each of which is a combination of a CSWAB module (Convnext Spatial and Window-based self-Attention Block) and a downsampling operation. The decoder is divided into four stages, each of which is a combination of a WAB module (Window-based self-Attention Block) and an upsampling operation. The WAB module is a window-based self-attention module. Each stage of the encoder is skip-connected to the corresponding stage of the decoder. The skip connection is introduced after the CSWAB module in each stage, and the skip connection introduced in each stage must be processed by the FA module.
[0084] The encoder performs multi-stage encoding on the retinal image to obtain preliminary retinal coding features. The preliminary retinal coding features include first-stage coding features and second-stage coding features. The first-stage coding features are coding features of the first preset coding stage, which is the first stage of the encoder. The first-stage coding features are obtained by performing first-stage encoding on the retinal image using the first-stage CSWAB module.
[0085] In step S220 of some embodiments, in order to extract deeper features and improve the retinal lesion segmentation model's ability to segment multi-scale lesions, the first-stage coded features are downsampled, and the downsampled first-stage coded features are subjected to second-stage encoding by the second-stage CSWAB module to obtain second-stage coded features. The second-stage coded features are coded features of the second preset encoding stage, where the second preset encoding stage is the second stage of the encoder, and the first preset stage and the second preset stage are adjacent. The second-stage encoding is encoded in the same manner as the first-stage encoding.
[0086] In some embodiments, the preliminary retinal coding features include third-stage coding features and fourth-stage coding features. The third-stage coding features are coding features of the third preset coding stage, i.e., coding features output by the third CSWAB module. The third preset coding stage is the third stage of the encoder. The second preset stage is adjacent to the first and third preset stages. The fourth-stage coding features are coding features of the fourth preset coding stage, i.e., coding features output by the fourth CSWAB module. The fourth preset coding stage is the fourth stage of the encoder. The third preset stage is adjacent to the second and fourth preset stages.
[0087] Downsampling is performed after the CSWAB module in each encoding stage. This downsampling is implemented using a convolutional layer with a kernel size of 2 and a stride of 2, followed by layer normalization. This halves the spatial size of the input feature map and doubles the number of channels. The number of channels in the first, second, third, and fourth stage encoded features is 96, 192, 384, and 768, respectively.
[0088] In the above steps S210 to S220, features of different scales can be extracted through multi-stage encoding to ensure the comprehensiveness of feature extraction and improve the segmentation capability of multi-scale effusion lesions.
[0089] See also Figure 4 In some embodiments, step S210 may include, but is not limited to, steps S410 to S450:
[0090] Step S410, performing first feature extraction on the retinal image to obtain a first retinal feature map;
[0091] Step S420, performing second feature extraction on the first retinal feature map to obtain a second retinal feature map;
[0092] Step S430, performing feature fusion on the first retinal feature map and the second retinal feature map to obtain a fused retinal feature map;
[0093] Step S440, performing regional perception attention processing on the fused retinal feature map to obtain regional perception attention;
[0094] Step S450: Perform window attention processing on the region-aware attention to obtain the first-stage coding features.
[0095] In step S410 of some embodiments, convolution processing is performed on the retinal image to obtain a first retinal feature map. The first retinal feature map may be features such as edges, textures, and lesion outlines of the retinal image.
[0096] In step S420 of some embodiments, the CSWAB module includes convnext, a region-aware attention module, and a window-based self-attention module. Convnext is designed with large convolution kernels and residual connections. It is a convolutional neural network with a large convolution kernel. There is a residual connection between the input and output of convnext. Convnext sequentially includes a convolution layer with a convolution kernel size of 7×7, layer normalization, a convolution layer with a convolution kernel size of 1×1, a GELU activation function, a global response normalization, and a convolution layer with a convolution kernel size of 1×1. The first retinal feature map is sequentially subjected to second feature extraction through each layer of convnext to obtain a second retinal feature map. Through convnext, the inductive bias can be effectively utilized and long-range relationships of features can be established to avoid the loss of contextual information.
[0097] In step S430 of some embodiments, the first retinal feature map and the second retinal feature map are added element-by-element through a residual connection to obtain a fused retinal feature map.
[0098] In step S440 of some embodiments, regional awareness attention is processed on the fused retinal feature map by a regional awareness-based attention module to obtain regional awareness attention.
[0099] In step S450 of some embodiments, window attention processing is performed on the region-aware attention through a window-based self-attention module to obtain the first-stage encoding features.
[0100] In the above steps S410 to S450, the model can focus on the relevant features of the lesion area through regional perception attention processing, and can extract long-distance dependencies through window attention processing, thereby improving the model's feature extraction ability and the ability to detect minor lesions.
[0101] See also Figure 5 In some embodiments, step S440 may include but is not limited to steps S510 to S530:
[0102] Step S510, obtaining a first channel feature and a second channel feature of the fused retinal feature map; the first channel feature is the channel maximum value of the fused retinal feature map; the second channel feature is the channel average value of the fused retinal feature map;
[0103] Step S520, concatenating the first channel feature and the second channel feature from the channel dimension to obtain a target channel feature;
[0104] Step S530 , performing regional perception attention transformation on the fused retinal feature map according to the preset regional attention matrix and the target channel features to obtain regional perception attention.
[0105] In step S510 of some embodiments, the fused retinal feature map is a tensor input to the region-aware attention module. The tensor is processed to maintain the size of the spatial dimension unchanged. The average and maximum values of the fused retinal feature map are calculated in the channel dimension. The maximum value of the fused retinal feature map is used as the first channel feature, and the average value of the fused retinal feature map is used as the second channel feature. The channel dimension of both the first channel feature and the second channel feature tensors is 1.
[0106] In step S520 of some embodiments, the two tensors of the first channel feature and the second channel feature are concatenated along the channel dimension to obtain the target channel feature.
[0107] In step S530 of some embodiments, the target channel features are convolved to obtain convolution channel features, the convolution channel features are activated to obtain spatial attention, and the fused retinal feature map is transformed into regional perception attention according to the preset regional attention matrix and spatial attention to obtain regional perception attention.
[0108] Through the above steps S510 to S530, relevant features for segmentation of fluid effusion lesions in each area of the image can be automatically extracted, so that the segmented fluid effusion lesions are more complete and accurate.
[0109] See also Figure 6 In some embodiments, step S530 may include but is not limited to steps S610 to S640:
[0110] Step S610, performing convolution processing on the target channel feature to obtain a convolution channel feature;
[0111] Step S620, activating the convolution channel features to obtain spatial attention;
[0112] Step S630, dividing the preset regional attention matrix into regions to obtain a first preset regional attention, a second preset regional attention, and a third preset regional attention; the second preset regional attention is greater than the first preset regional attention and the third preset regional attention;
[0113] Step S640: performing regional perception attention transformation on the fused retinal feature map according to the first preset regional attention, the second preset regional attention, the third preset regional attention and the spatial attention to obtain regional perception attention.
[0114] In step S610 of some embodiments, convolution processing is performed on the target channel feature, and the channel dimension is reprocessed to 1 to obtain a convolution channel feature.
[0115] In step S620 of some embodiments, the convolution channel features are activated using a sigmoid activation function to obtain spatial attention. The shape of the spatial attention is the same as the spatial dimension of the fused retinal feature map. The value of the spatial attention ranges from [0, 1] and is used to represent the importance of each spatial location to the segmentation of the effusion lesion.
[0116] In step S630 of some embodiments, the preset regional attention matrix is a pre-created prior attention weight matrix of the same shape as the spatial attention matrix. The preset regional attention matrix is evenly divided into three regions: upper, middle, and lower, to obtain the first preset regional attention, the second preset regional attention, and the third preset regional attention. The first preset regional attention is the attention weight of the upper region, the second preset regional attention is the attention weight of the middle region, and the third preset regional attention is the attention weight of the lower region.
[0117] Retinal fluid lesions are primarily located in the central region of OCT retinal images. To make the model focus more on the lesion-related features in the central region, the second preset region attention needs to be greater than the first and third preset region attention. The second preset region attention can be set to 1, while the first and third preset region attention can be set to 0.5.
[0118] In step S640 of some embodiments, the spatial attention is multiplied element-wise by the a priori attention weight matrix having the first preset regional attention, the second preset regional attention, and the third preset regional attention to obtain the spatial attention under the weight matrix. The spatial attention under the weight matrix is multiplied element-wise by the fused retinal feature map to obtain the regional perceptual attention.
[0119] In the above steps S610 to S640, the regional perception attention takes into account the characteristics of the location of the lesion itself, effectively utilizes the prior knowledge of retinal images and retinal lesions, can extract lesion-related features, and improves the accuracy of fluid lesion segmentation.
[0120] See also Figure 7 In some embodiments, step S450 may include but is not limited to steps S710 to S730:
[0121] Step S710, dividing the regional perception attention into windows to obtain multiple preliminary window attentions;
[0122] Step S720, performing multi-head attention processing on each preliminary window attention to obtain candidate window attention corresponding to each preliminary window attention;
[0123] In step S730, the attention of each candidate window is concatenated to obtain the first-stage encoding features.
[0124] In step S710 of some embodiments, the window-based self-attention module includes layer normalization, window self-attention, layer normalization, and a feedforward network in sequence. There is a first residual connection between the input of the first layer normalization and the output of the window self-attention, and a second residual connection between the input of the second layer normalization and the output of the feedforward network. Layer normalization is performed on the region-aware attention, and a window self-attention operation is performed on the region-aware attention after layer normalization. The window self-attention operation is to divide the region-aware attention after layer normalization into non-overlapping windows of a fixed size, and the region-aware attention within the window is the preliminary window attention.
[0125] In step S720 of some embodiments, a multi-head attention operation is performed on the preliminary window attention of each window to obtain candidate window attention corresponding to the preliminary window attention of each window. Performing the multi-head attention operation within each window rather than on the entire image reduces the computational complexity of the attention calculation. Performing the multi-head attention calculation on offset windows through the window shift operation facilitates information exchange between adjacent windows and enhances the model's ability to capture long-range dependencies.
[0126] In step S730 of some embodiments, the candidate window attention of each window is feature concatenated to obtain a concatenated attention feature. The region-aware attention and the concatenated attention features are element-wise added to obtain a summed attention feature. The summed attention feature is layer-normalized, and the layer-normalized summed attention feature is feedforwarded through a feedforward network to obtain a feedforward feature. The summed attention feature and the feedforward feature are element-wise added to obtain a first-stage encoding feature.
[0127] The above steps S710 to S730 can effectively model long-distance dependencies through the window-based self-attention module, avoid the loss of context information, and improve the feature extraction capability of the model.
[0128] In step S130 of some embodiments, the features output by the CSWAB modules at each stage are used as skip connections, such as Figure 3 As shown in the figure, the FA module can automatically fuse multi-stage features, so that each jump connection is fused with the jump connection of its adjacent stage. For the jump connection of the first stage, it is fused with the jump connection of the second stage. For the jump connection of the fourth stage, it is fused with the jump connection of the third stage. For the jump connection of the second stage, it is fused with the jump connections of the first and third stages respectively. For the jump connection of the third stage, it is fused with the jump connections of the second and fourth stages respectively.
[0129] Specifically, for the first-stage coded features, the second-stage coded features are upsampled so that the second-stage coded features have the same spatial size and channel size as the first-stage coded features. The upsampled second-stage coded features and the first-stage coded features are added element-by-element to obtain the target retinal coded features.
[0130] For the second-stage encoded features, the first-stage encoded features are downsampled, and the downsampled first-stage encoded features and the second-stage encoded features are added element by element to obtain the first fused features. The third-stage encoded features are upsampled, and the upsampled third-stage encoded features and the second-stage encoded features are added element by element to obtain the second fused features.
[0131] For the third-stage coded features, the second-stage coded features are downsampled, and the downsampled second-stage coded features and the third-stage coded features are added element-by-element to obtain the third fused features. The fourth-stage coded features are upsampled so that the third-stage coded features have the same spatial size and channel size as the fourth-stage coded features, and the upsampled fourth-stage coded features and the third-stage coded features are added element-by-element to obtain the fourth fused features.
[0132] For the fourth stage coding features, the third stage coding features are downsampled, and the downsampled third stage coding features and the fourth stage coding features are added element by element to obtain the fifth fusion features.
[0133] See also Figure 8 In some embodiments, step S140 may include, but is not limited to, steps S810 to S850:
[0134] Step S810, performing average pooling processing on the target retinal coding feature to obtain a first pooled feature;
[0135] Step S820, performing maximum pooling processing on the target retinal coding feature to obtain a second pooled feature;
[0136] Step S830, performing attention calculation based on the first pooled features and the second pooled features to obtain a target attention matrix;
[0137] Step S840, performing attention processing on the target retinal encoding features according to the target attention matrix to obtain preliminary retinal attention features;
[0138] Step S850: perform feature concatenation on the preliminary retinal attention features and the first-stage encoding features from the channel dimension to obtain the target retinal attention features.
[0139] In steps S810 to S820 of some embodiments, all fused jump connections need to undergo attention processing to extract important features of adjacent stage fusion. Figure 9 As shown in the figure, the target retinal encoding features are average pooled and max pooled respectively through the pooling layer to obtain two two-dimensional tensors. The two-dimensional tensor obtained by average pooling is the first pooling feature. The two-dimensional tensor obtained by max pooling is the second pooling feature.
[0140] In step S830 of some embodiments, the first pooled feature and the second pooled feature are processed sequentially through a fully connected layer, a ReLU activation function, and a fully connected layer to obtain a first fully connected feature corresponding to the first pooled feature and a second fully connected feature corresponding to the second pooled feature. The first fully connected feature and the second fully connected feature are added to obtain a fused fully connected feature. The fused fully connected feature is activated using a sigmoid activation function to obtain the attention weight of each channel, i.e., the target attention matrix.
[0141] In step S840 of some embodiments, the target attention matrix is reshaped into the same shape as the input fusion feature, i.e., the target retinal coding feature. The reshaped target attention matrix is element-wise multiplied with the target retinal coding feature to apply the attention weight to the target retinal coding feature, thereby obtaining a preliminary retinal attention feature.
[0142] In step S850 of some embodiments, for the first and fourth stages, the fused features obtained through attention processing are concatenated with the original features of the stage in the channel dimension. For the second and third stages, the two features obtained through attention processing are concatenated in the channel dimension. A convolution operation is applied to the concatenated features of each stage, reducing the number of channels by half, and the final skip connection features are input to the corresponding decoder stage.
[0143] Specifically, for the first preset encoding stage, the preliminary retinal attention features and the first-stage encoding features are feature spliced from the channel dimension to obtain the target spliced encoding features, and the target spliced encoding features are convolved to obtain the target retinal attention features.
[0144] In the second preset encoding stage, attention processing is performed on the first fused feature to obtain the first attention, and attention processing is performed on the second fused feature to obtain the second attention. The first attention and the second attention are concatenated along the channel dimension, and the concatenated result is convolved. The calculation method for attention processing can be referred to steps S810 and S840 and will not be repeated here.
[0145] In the third preset encoding stage, attention processing is performed on the third fused feature to obtain a third attention, and attention processing is performed on the fourth fused feature to obtain a fourth attention. The third and fourth attention are concatenated along the channel dimension, and the concatenated result is convolved. The calculation method for attention processing can be referred to steps S810 and S840 and will not be repeated here.
[0146] For the fourth preset encoding stage, the fifth fusion feature is subjected to attention processing to obtain the fifth attention, and the fifth attention is concatenated with the fourth stage encoding feature from the channel dimension. The calculation method of the attention processing can be referred to steps S810 and S840 and will not be repeated here.
[0147] Through the above steps S810 to S850, important features of adjacent fusion stages can be extracted, ensuring the comprehensiveness of feature extraction.
[0148] In step S150 of some embodiments, the second-stage coded features are downsampled, and the downsampled second-stage coded features are subjected to a third encoding via a third CSWAB module to obtain third-stage coded features. The third-stage coded features are downsampled, and the downsampled third-stage coded features are subjected to a fourth encoding via a fourth CSWAB module to obtain fourth-stage coded features. The fourth-stage coded features are downsampled, and feature extraction is performed on the downsampled fourth-stage coded features via a bottleneck layer to obtain bottleneck features.
[0149] The bottleneck layer consists of a convolution kernel of size 7, which is used to process the features from the encoder and transmit the processed features to the decoder. The convnext of the bottleneck layer can be referred to as the convnext of the CSWAB module and will not be repeated here.
[0150] The bottleneck features are input into the decoder, and the decoder performs multi-stage decoding on the bottleneck features to obtain preliminary retinal decoding features. The preliminary retinal decoding features include the target stage decoding features, the first stage decoding features, and the second stage decoding features of the target decoding stage. The target decoding stage matches the first preset encoding stage. The target stage decoding features refer to the decoding features output by the third WAB module starting from the bottleneck layer. The first stage decoding features refer to the decoding features output by the first WAB module starting from the bottleneck layer. The second stage decoding features refer to the decoding features output by the second WAB module starting from the bottleneck layer. The WAB module has the same definition as the window-based self-attention module in the CSWAB module.
[0151] The upsampling operation in the decoder includes bilinear interpolation to double the spatial size of the feature map. The upsampling operation also includes a convolutional layer with a stride of 3 and a padding of 1 and layer normalization to make the number of channels of the feature map in the decoder stage become the required number of channels. The number of channels output by the WAB module in each stage of the decoder is 192, 384, 768, and 1536 respectively. The processing flow of the decoder is as follows: the features from the bottleneck layer or the decoder of the previous stage are received and spliced with the features transmitted from the encoder of the corresponding stage through the jump connection. The splicing operation is performed on the channel dimension. Figure 3 The concatenated features are fed into a windowed self-attention module for processing, followed by upsampling via bilinear interpolation and convolutional layers. The windowed self-attention module further extracts long-range image dependencies. By concatenating features with skip connections, high-resolution features can be combined to supplement the extraction of local image details, avoiding missing features of subtle lesion structures.
[0152] In step S160 of some embodiments, the target stage decoding features are upsampled, and the upsampled target stage decoding features and the target retinal attention features are feature concatenated from the channel dimension to obtain the target retinal decoding features.
[0153] In step S170 of some embodiments, the target retinal decoding features are decoded by the last WAB module, and the decoded target retinal decoding features are input to the segmentation head to perform lesion segmentation on the retinal image to obtain the retinal lesion category and retinal lesion location. Figure 10a As shown in the figure, the types and locations of retinal lesions are shown in the figure. Figure 10b As shown. Figure 10b Medium gray represents intraretinal fluid, and white represents subretinal fluid.
[0154] Construct a weighted combination loss function, which is composed of the weighted sum of the cross entropy loss function and the Dice loss function. The weighted combination loss function is used to calculate the loss value during segmentation model training, and backpropagates based on the loss value to update and optimize the model parameters. The expression of the cross entropy loss function is shown in formula (1):
[0155]
[0156] Where y is the true mask, p is the predicted segmentation result of the input OCT retinal image, and N is the number of lesion types.
[0157] The expression of the Dice loss function is shown in formula (2):
[0158]
[0159] Among them, |X| and |Y| represent the predicted result and the true mask respectively, and |X∩Y| represents the intersection between the predicted result and the true mask.
[0160] The expression of the weighted combination loss function is shown in formula (3):
[0161] L all =L ce +L dice Formula (3)
[0162] The retinal lesion segmentation model was trained by inputting OCT training data into the model to generate a predicted lesion segmentation map. This map was then compared with the mask image containing lesion segmentation information. The loss was calculated using a constructed weighted loss function, and the model parameters were iteratively updated. The segmentation model with the highest evaluation metric on the test data was obtained. The predicted lesion segmentation map generated by the retinal lesion segmentation model had the same height and width as the input OCT retinal image and the corresponding mask image. The segmentation map simultaneously included segmentation information for the three lesions, represented by different pixel values. The pixel types and representation methods of the predicted segmentation map and the mask image were consistent.
[0163] The evaluation index on the test data is the Dice index. The Dice index is used to evaluate the model on the test data. When the Dice index value of the model on the test data reaches the highest, it proves that the segmentation effect of the model is the best at this time, and the model is selected as the final segmentation model. The model is evaluated on the test data and the value of the Dice evaluation index is calculated. The expression of the Dice evaluation index is shown in formula (4):
[0164]
[0165] Among them, |X| and |Y| represent the predicted result and the true mask respectively, and |X∩Y| represents the intersection between the predicted result and the true mask. The Dice indicator measures the similarity between the predicted segmentation result and the true mask, and its range is between 0 and 1. When the value of the Dice indicator is the highest, it is also necessary to calculate the values of other evaluation indicators of the model on the test data to comprehensively evaluate the performance of the model. Therefore, the IoU indicator, AVD indicator and ACC indicator are also introduced, and their expressions are shown in formulas (5) to (7) respectively:
[0166]
[0167] AVD=abs(|X|-|Y|) Formula (6)
[0168]
[0169] Where |X∪Y| represents the union of the predicted result and the true mask. TP (True Positives) is the number of positive samples correctly classified as positive samples, TN (True Negatives) is the number of negative samples correctly classified as negative samples, FP (False Positives) is the number of negative samples incorrectly classified as positive samples, and FN (False Negatives) is the number of positive samples incorrectly classified as negative samples.
[0170] This application trains the retinal lesion segmentation model on a workstation equipped with an NVIDIA-A100 GPU. The programming language is Python 3.7 and the Pytorch 1.12 deep learning framework is used. The batch size of the training phase is set to 8, the optimizer is AdamW, and a total of 150 rounds of training are performed. The initial learning rate is set to 10 -4 , using the “Poly” strategy to adjust the learning rate.
[0171] Table 1 shows the lesion segmentation effect of the retinal lesion segmentation model on OCT images. The values of all indicators when the Dice index of the retinal lesion segmentation model on the test data is the highest are shown in Table 1:
[0172] Lesion type Intraretinal fluid Subretinal fluid Pigment epithelial detachment Dice 0.7649 0.8626 0.8569 Indicator (IoU) 0.6382 0.7764 0.7695 Indicator (AVD) 0.1872 0.1387 0.1304 Index (ACC) 0.9193 0.9659 0.9616
[0173] Table 1
[0174] Through experiments and statistics, the model has high performance indicators in lesion segmentation and has high segmentation accuracy.
[0175] The retinal lesion segmentation model proposed in this application reconstructs the model backbone network by applying the convolutional neural network convnext with large convolution kernels and the window self-attention module, and adds an adaptive attention module with prior knowledge. At the same time, the multi-scale feature fusion method is applied to enable the model to better learn lesion-related features, has a strong feature extraction capability, can accurately segment the retinal lesion area in the OCT image, improves the segmentation accuracy of the lesion, is suitable for assisting clinical disease diagnosis and treatment, and has high application value. The retinal lesion segmentation model proposed in this application improves the feature extraction capability and segmentation effect, the contour edge segmentation is more accurate, and the segmentation capability for tiny lesions is stronger, assisting doctors to discover tiny lesions that are difficult to observe, accurately quantify the patient's condition, formulate personalized treatment plans for patients, and improve the doctor's diagnostic efficiency. The method of this application improves the segmentation accuracy of retinal lesions in OCT images, solves the segmentation errors that may be caused by the doctor's lack of experience and subjective carelessness, is suitable for assisting clinical disease diagnosis and treatment, and has important clinical significance and application value.
[0176] See also Figure 11 The present application also provides a retinal lesion segmentation device that can implement the above-mentioned retinal lesion segmentation method. The device includes:
[0177] An acquisition module 1110 is configured to acquire a retinal image;
[0178] The multi-stage encoding module 1120 is configured to perform multi-stage encoding on the retinal image to obtain preliminary retinal coding features; the preliminary retinal coding features include first-stage coding features and second-stage coding features; the first-stage coding features are coding features of the first preset coding stage; the second-stage coding features are coding features of the second preset coding stage; the first preset stage and the second preset stage are adjacent;
[0179] A feature fusion module 1130 is used to fuse the first-stage coding features and the second-stage coding features to obtain target retinal coding features;
[0180] an attention processing module 1140 for performing attention processing on the target retinal encoding feature to obtain a target retinal attention feature;
[0181] A multi-stage decoding module 1150 is configured to perform multi-stage decoding on the second-stage coding features to obtain preliminary retinal decoding features; the preliminary retinal decoding features include target-stage decoding features of a target decoding stage, where the target decoding stage matches the first preset coding stage;
[0182] A feature concatenation module 1160 is configured to concatenate the target retinal attention feature and the target stage decoding feature from a channel dimension to obtain a target retinal decoding feature;
[0183] The lesion segmentation module 1170 is used to perform lesion segmentation on the retinal image according to the target retinal decoding features to obtain the retinal lesion category and retinal lesion location.
[0184] The specific implementation of the retinal lesion segmentation device is basically the same as the specific embodiment of the above-mentioned retinal lesion segmentation method, and will not be repeated here.
[0185] The present application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned retinal lesion segmentation method when executing the computer program. The electronic device can be any smart terminal including a tablet computer, an in-vehicle computer, or the like.
[0186] See also Figure 12 , Figure 12 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:
[0187] The processor 1210 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0188] The memory 1220 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1220 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program codes are stored in the memory 1220 and are called by the processor 1210 to execute the retinal lesion segmentation method of the embodiments of this application.
[0189] Input / output interface 1230, used for information input and output;
[0190] Communication interface 1240, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0191] bus 1250 , which transmits information between various components of the device (e.g., processor 1210 , memory 1220 , input / output interface 1230 , and communication interface 1240 );
[0192] The processor 1210 , the memory 1220 , the input / output interface 1230 , and the communication interface 1240 are communicatively connected to each other within the device via a bus 1250 .
[0193] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, which implements the above-mentioned retinal lesion segmentation method when executed by a processor.
[0194] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0195] The retinal lesion segmentation method, retinal lesion segmentation device, electronic device and computer-readable storage medium provided in the embodiments of the present application obtain a retinal image, perform multi-stage encoding on the retinal image, and obtain preliminary retinal coding features. Retinal features of different scales in multiple coding stages can be obtained through multi-stage encoding to comprehensively extract the features of the retinal image, improve the segmentation capability of multi-scale lesions, and avoid missing tiny lesion structures. By performing feature fusion on the first-stage coding features and the second-stage coding features, it is possible to establish a connection between adjacent scale features, fuse multi-scale features, and retain the contextual information of the features to obtain target retinal coding features. In order to extract features important for lesion segmentation from the target retinal coding features obtained by fusion of adjacent stages, attention processing is performed on the target retinal coding features to obtain target retinal attention features. By performing multi-stage decoding on the second-stage coding features, high-resolution multi-scale features can be extracted to obtain preliminary retinal decoding features. By performing feature splicing on the target retinal attention features and the target stage decoding features from the channel dimension, it is possible to supplement the extraction of local detail information in combination with high-resolution features to obtain target retinal decoding features. The retinal image is segmented for lesions according to the target retinal decoding features to obtain the retinal lesion category and retinal lesion location, which improves the segmentation capability of multi-scale lesions and thus improves the accuracy of retinal lesion segmentation.
[0196] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0197] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0198] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0199] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0200] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0201] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0202] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0203] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0204] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0205] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0206] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A retinal lesion segmentation method, characterized in that: The method comprises: Acquire retinal images; Performing multi-stage encoding on the retinal image to obtain preliminary retinal coding features; the preliminary retinal coding features include first-stage coding features and second-stage coding features; the first-stage coding features are coding features of a first preset coding stage; the second-stage coding features are coding features of a second preset coding stage; the first preset coding stage and the second preset coding stage are adjacent; Performing feature fusion on the first-stage coding features and the second-stage coding features to obtain target retinal coding features; performing attention processing on the target retinal encoding feature to obtain a target retinal attention feature; Performing multi-stage decoding on the second-stage encoding features to obtain preliminary retinal decoding features; the preliminary retinal decoding features include target-stage decoding features of a target decoding stage, and the target decoding stage matches the first preset encoding stage; Performing feature concatenation on the target retinal attention feature and the target stage decoding feature from a channel dimension to obtain a target retinal decoding feature; Performing lesion segmentation on the retinal image according to the target retinal decoding features to obtain a retinal lesion category and a retinal lesion location; The multi-stage encoding of the retinal image to obtain preliminary retinal coding features includes: Performing first-stage encoding on the retinal image to obtain first-stage encoding features; performing second-stage encoding on the first-stage encoding features to obtain second-stage encoding features; The performing first-stage encoding on the retinal image to obtain the first-stage encoding features includes: Perform a first feature extraction on the retinal image to obtain a first retinal feature map; perform a second feature extraction on the first retinal feature map to obtain a second retinal feature map; perform feature fusion on the first retinal feature map and the second retinal feature map to obtain a fused retinal feature map; perform regional perception attention processing on the fused retinal feature map to obtain regional perception attention; perform window attention processing on the regional perception attention to obtain the first-stage coding feature.
2. The retinal lesion segmentation method according to claim 1, wherein: The performing regional perception attention processing on the fused retinal feature map to obtain regional perception attention includes: Acquire a first channel feature and a second channel feature of the fused retinal feature map; the first channel feature is a channel maximum value of the fused retinal feature map; the second channel feature is a channel average value of the fused retinal feature map; Concatenate the first channel feature and the second channel feature from the channel dimension to obtain a target channel feature; The regional perception attention is transformed by performing a regional perception attention transformation on the fused retinal feature map according to a preset regional attention matrix and the target channel characteristics to obtain the regional perception attention.
3. The retinal lesion segmentation method according to claim 2, wherein: The performing regional perception attention transformation on the fused retinal feature map according to the preset regional attention matrix and the target channel feature to obtain the regional perception attention includes: Performing convolution processing on the target channel feature to obtain a convolution channel feature; Activate the convolution channel features to obtain spatial attention; Dividing the preset area attention matrix into regions to obtain a first preset area attention, a second preset area attention, and a third preset area attention; the second preset area attention is greater than the first preset area attention and the third preset area attention; The regional perception attention is obtained by performing a regional perception attention transformation on the fused retinal feature map according to the first preset regional attention, the second preset regional attention, the third preset regional attention and the spatial attention.
4. The retinal lesion segmentation method according to claim 1, wherein: The performing window attention processing on the region-aware attention to obtain the first-stage encoding features includes: Dividing the regional perception attention into windows to obtain multiple preliminary window attentions; Performing multi-head attention processing on each of the preliminary window attentions to obtain candidate window attentions corresponding to each of the preliminary window attentions; The attention of each candidate window is concatenated to obtain the encoding features of the first stage.
5. The retinal lesion segmentation method according to any one of claims 1 to 4, characterized in that: The performing attention processing on the target retinal encoding feature to obtain the target retinal attention feature includes: Performing average pooling processing on the target retinal coding feature to obtain a first pooling feature; Performing maximum pooling processing on the target retinal coding feature to obtain a second pooling feature; Performing attention calculation based on the first pooled features and the second pooled features to obtain a target attention matrix; Performing attention processing on the target retinal encoding feature according to the target attention matrix to obtain a preliminary retinal attention feature; The preliminary retinal attention feature and the first-stage encoding feature are feature concatenated from the channel dimension to obtain the target retinal attention feature.
6. A retinal lesion segmentation device, characterized in that: The device comprises: An acquisition module, used for acquiring retinal images; a multi-stage encoding module, configured to perform multi-stage encoding on the retinal image to obtain preliminary retinal coding features; the preliminary retinal coding features include first-stage coding features and second-stage coding features; the first-stage coding features are coding features of a first preset coding stage; the second-stage coding features are coding features of a second preset coding stage; the first preset coding stage and the second preset coding stage are adjacent; a feature fusion module, configured to fuse the first-stage coding features and the second-stage coding features to obtain target retinal coding features; an attention processing module, configured to perform attention processing on the target retinal encoding feature to obtain a target retinal attention feature; a multi-stage decoding module, configured to perform multi-stage decoding on the second-stage encoding features to obtain preliminary retinal decoding features; the preliminary retinal decoding features include target-stage decoding features of a target decoding stage, the target decoding stage matching the first preset encoding stage; a feature splicing module, configured to perform feature splicing on the target retinal attention feature and the target stage decoding feature from a channel dimension to obtain a target retinal decoding feature; a lesion segmentation module, configured to perform lesion segmentation on the retinal image according to the target retinal decoding features to obtain a retinal lesion category and a retinal lesion location; The retinal lesion segmentation device is also used for: Performing first-stage encoding on the retinal image to obtain first-stage encoding features; performing second-stage encoding on the first-stage encoding features to obtain second-stage encoding features; Perform a first feature extraction on the retinal image to obtain a first retinal feature map; perform a second feature extraction on the first retinal feature map to obtain a second retinal feature map; perform feature fusion on the first retinal feature map and the second retinal feature map to obtain a fused retinal feature map; perform regional perception attention processing on the fused retinal feature map to obtain regional perception attention; perform window attention processing on the regional perception attention to obtain the first-stage coding feature.
7. An electronic device, characterized in that The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the retinal lesion segmentation method according to any one of claims 1 to 5 when executing the computer program.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the retinal lesion segmentation method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Automatic segmentation method and system for rectal cancer image and chemoradiotherapy reaction prediction system
CN115170568A
Glaucoma image detection method and system, electronic equipment and storage medium
CN116342524A