A deep learning-based method and system for segmenting and identifying diabetic retinopathy lesions
By building an intelligent lesion segmentation identification system, using sample expansion and multi-scale feature integration technology, the problem of insufficient timeliness and accuracy of lesion segmentation in the existing technology is solved, and real-time fine segmentation of lesion lesions is achieved.
Patent Information
- Application Number
- CN202210575490.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-24
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2042-05-24
AI Technical Summary
In the prior art, the deep learning-based lesion segmentation method of sugar network image lesion segmentation has problems of low timeliness and insufficient accuracy, especially on mobile devices, it is difficult to achieve efficient lesion fine segmentation.
By building an intelligent lesion segmentation identification system, the pre-stored fundus image data set is called, sample expansion and feature matching is performed, multi-scale feature integration is used by encoder and decoder, local continuity features are generated, and the basic network set training module is optimized to realize real-time fine segmentation of different types of sugar network lesions.
Real-time fine segmentation of different types of sugar mesh lesions is achieved, the accuracy of lesion recognition and the performance of processing equipment is improved, and the problem of low timeliness and accuracy in the prior art is solved.
Smart Images

Figure CN114882054B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field related to image processing, and specifically to a method and system for segmenting and identifying diabetic retinopathy images based on deep learning. Background Art
[0002] Diabetic retinopathy (DR) refers to a series of lesions caused by pathological changes in retinal capillaries, arterioles, and venules, as well as leakage or blockage of these microvascular tissues.
[0003] Retinal fundus imaging is an important imaging tool for observing and diagnosing diabetic retinopathy. However, due to limitations such as the limited field of view of traditional optical lenses, high requirements for pupil and refractive media, electronic system noise, and suboptimal image acquisition environments, the acquired fundus image data remains inadequate for detecting and analyzing the location and morphology of diabetic retinopathy lesions.
[0004] In recent years, with the development of deep learning algorithms based on neural networks, deep learning models have been widely used in the field of medical imaging. However, due to the diversity of noise in fundus images and the limitations of mobile device performance, denoising and enhancing fundus images on mobile devices remains a challenging task. Current deep learning methods for diabetic retinopathy lesion segmentation still suffer from technical issues such as low timeliness and limited accuracy. Summary of the Invention
[0005] The present application provides a method and system for segmenting and identifying diabetic retinopathy images based on deep learning, which is used to solve the problems of insufficient fine analysis performance of fundus diabetic retinopathy images and performance limitations of processing equipment in the existing technology. The deep learning model still has problems of timeliness and low accuracy, and achieves the technical effect of real-time fine segmentation of different types of diabetic retinopathy lesions.
[0006] In view of the above problems, the present application provides a method and system for segmenting and identifying diabetic retinopathy images based on deep learning.
[0007] In the first aspect, an embodiment of the present application provides a method for segmenting and identifying diabetic retinopathy images based on deep learning, the method being applied to an intelligent lesion segmentation and identification system, the intelligent lesion segmentation and identification system being communicatively connected to a basic network set training module, the method comprising: calling a pre-stored fundus image dataset through the intelligent lesion segmentation and identification system, wherein the fundus image dataset comprises hemorrhage, microaneurysm, hard exudation, soft exudation, neovascular fibroproliferative membrane, and abnormal microvascular diabetic retinopathy lesions; and performing sample expansion of the fundus image dataset according to the lesion type to obtain an expanded fundus image dataset; inputting the expanded fundus image dataset into the encoder of the basic network set training module to obtain n overlapping image blocks of a predetermined size, wherein n is A positive integer greater than 1; performing feature matching on the n overlapping image blocks through the encoder to obtain segmentation multi-level features; merging the n overlapping image blocks through the basic network set training module and recording the merging parameters; performing feature matching based on the merging result of the n image blocks, and generating local continuity features through the merging parameters and the feature matching results; inputting the segmentation multi-level features and the local continuity features into the decoder of the basic network set training module for multi-scale feature integration, and verifying the multi-scale feature integration results based on the constraint results corresponding to the expanded fundus image data set. When the verification passes, the construction of the basic network set training module is completed; inputting the image data into the basic network set training module to obtain lesion recognition and segmentation identification results.
[0008] In a second aspect, an embodiment of the present application provides a deep learning-based diabetic retinopathy image lesion segmentation and identification system, the system comprising: an information calling unit, the information calling unit being used to call a pre-stored fundus image dataset through the intelligent lesion segmentation and identification system, wherein the fundus image dataset comprises hemorrhage, microaneurysm, hard exudation, soft exudation, neovascular fibroproliferative membrane, and microvascular abnormal diabetic retinopathy lesions; a sample expansion unit, the sample expansion unit being used to perform sample expansion of the fundus image dataset according to the lesion type to obtain an expanded fundus image dataset; an input unit, the input unit being used to input the expanded fundus image dataset into the encoder of the basic network set training module to obtain n overlapping image blocks of a predetermined size, wherein n is a positive integer greater than 1; a matching unit, the matching unit being used to obtain n overlapping image blocks of a predetermined size through the encoder Perform feature matching on the n overlapping image blocks to obtain segmentation multi-level features; a merging unit, the merging unit is used to merge the n overlapping image blocks through the basic network set training module and record the merging parameters; a feature matching unit, the feature matching unit performs feature matching according to the merging result of the n image blocks, and generates local continuity features through the merging parameters and the feature matching results; a construction unit, the construction unit is used to input the segmentation multi-level features and the local continuity features into the decoder for multi-scale feature integration, and verify the multi-scale feature integration results based on the constraint results corresponding to the expanded fundus image data set. When the verification is passed, the construction of the basic network set training module is completed; a processing unit, the processing unit is used to input image data into the basic network set training module to obtain lesion recognition and segmentation identification results.
[0009] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0010] The embodiment of the present application provides a method and system for building a modular simulation model of a typical control method, which calls a pre-stored fundus image data set through the intelligent lesion segmentation and identification system, wherein the fundus image data set includes hemorrhage, microaneurysm, hard exudation, soft exudation, neovascular fibroproliferative membrane, and microvascular abnormal diabetic retinopathy lesions; and expands the sample of the fundus image data set according to the lesion type to obtain an expanded fundus image data set; inputs the expanded fundus image data set into the encoder of the basic network set training module to obtain n overlapping image blocks of a predetermined size, wherein n is a positive integer greater than 1; and performs segmentation on the n overlapping image blocks through the encoder. Feature matching to obtain segmentation multi-level features; merge the n overlapping image blocks through the basic network set training module and record the merging parameters; perform feature matching based on the merging results of the n image blocks, and generate local continuity features through the merging parameters and feature matching results; input the segmentation multi-level features and the local continuity features into the decoder of the basic network set training module for multi-scale feature integration, and verify the multi-scale feature integration results based on the constraint results corresponding to the expanded fundus image data set. When the verification passes, the construction of the basic network set training module is completed; the image data is input into the basic network set training module to obtain the lesion recognition and segmentation identification results. This solves the problem of insufficient performance in fine analysis of fundus diabetic retinopathy images and performance limitations of processing equipment in the prior art, and the problem that deep learning models still have timeliness and low accuracy, and achieves the technical effect of real-time fine segmentation of different types of diabetic retinopathy lesions. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 A flowchart of a deep learning-based diabetic retinopathy image lesion segmentation and identification method is provided for this application;
[0012] Figure 2 This application provides a flowchart for obtaining an expanded fundus image dataset in a deep learning-based diabetic retinopathy image lesion segmentation and identification method;
[0013] Figure 3 This application provides a flowchart of merging overlapping image blocks in a deep learning-based diabetic retinopathy image lesion segmentation and identification method;
[0014] Figure 4 This application provides a structural diagram of a diabetic retinopathy image lesion segmentation and identification system based on deep learning;
[0015] Explanation of the accompanying drawings: information calling unit 100, sample expansion unit 200, input unit 300, matching unit 400, merging unit 500, feature matching unit 600, construction unit 700, processing unit 800. DETAILED DESCRIPTION
[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0017] This application provides a deep learning-based diabetic retinopathy image lesion segmentation and identification method and system to solve the problems of insufficient fine analysis performance of fundus diabetic retinopathy image lesions and performance limitations of processing equipment in the existing technology. Deep learning models still have problems of timeliness and low accuracy, and achieve the technical effect of real-time fine segmentation of different types of diabetic retinopathy lesions.
[0018] Example 1
[0019] like Figure 1 As shown, the present application provides a deep learning-based diabetic retinopathy image lesion segmentation and identification method, the method is applied to an intelligent lesion segmentation and identification system, the intelligent lesion segmentation and identification system is communicatively connected to a basic network set training module, and the method includes:
[0020] S100: calling a pre-stored fundus image dataset through the intelligent lesion segmentation and identification system, wherein the fundus image dataset includes hemorrhage, microaneurysm, hard exudate, soft exudate, neovascular fibroproliferative membrane, and microvascular abnormal diabetic retinopathy lesions;
[0021] Specifically, the earliest lesions of diabetic retinopathy include microaneurysms and small hemorrhages. Afterwards, vascular changes can progress to capillary non-perfusion, leading to hemorrhage, vitreous abnormalities and retinal microvascular abnormalities. Late lesions include occlusion of small arteries and veins, neovascularization of the optic disc, and formation and proliferation of neovascularization of the retina, iris and chamber angle.
[0022] The intelligent lesion segmentation and identification system records various types of fundus images processed in the past. In order to achieve real-time and fine segmentation of different types of diabetic retinopathy lesions, in an embodiment of the present application, a pre-stored fundus image dataset is called by the intelligent lesion segmentation and identification system, wherein the fundus image dataset includes hemorrhage, microaneurysm, hard exudation, soft exudation, neovascular fibroproliferative membrane, and microvascular abnormal diabetic retinopathy lesions. By obtaining different types of fundus images, on the one hand, the sample size of deep learning is increased, and on the other hand, the processing of different types of lesions is achieved.
[0023] S200: Expanding the sample of the fundus image dataset according to the lesion type to obtain an expanded fundus image dataset;
[0024] Specifically, in order to expand the number of fundus images and increase the number of training samples to obtain a high-precision diabetic retinopathy lesion segmentation model, in an embodiment of the present application, the fundus image dataset is sample expanded according to the lesion type to obtain an expanded fundus image dataset.
[0025] Alternatively, as Figure 2 As shown, an implementation of step S200 in the method provided in the embodiment of the present application includes:
[0026] S210: Constructing a generation model, wherein the generation model is a model for image generation;
[0027] S220: Inputting the fundus image dataset and random noise into the generation model to obtain a pre-expanded fundus image dataset;
[0028] S230: Identifying the pre-expanded fundus image dataset, and obtaining expanded data and feedback data according to the identification result;
[0029] S240: adding the expanded data to the expanded fundus image dataset, feeding the feedback data back to the generation model for model optimization, and generating a new expanded fundus image dataset using the optimized generation model;
[0030] S250: Repeat the process of feedback and expansion of the feedback data and the expansion data according to the new expanded fundus image dataset to obtain the expanded fundus image dataset.
[0031] Specifically, a generative model is constructed, and the generative model can be a mathematical model based on a neural network, and the generative model is a model for image generation, which is used to generate images that can be used to expand the fundus image data set; the random noise of the fundus image data set is input into the generative model to obtain a pre-expanded fundus image data set; the random noise is the noise caused by the inside of the electronic system or the noise caused by the image acquisition process, which causes unnecessary or redundant interference information in the fundus image data; the pre-expanded fundus image data set includes the fundus image data set and the fundus random image noise; the identifier can facilitate the identification of the data during the processing, and on the other hand, the type of data can be identified according to the identifier. In an embodiment of the present application, the pre-expanded fundus image data set is identified, and the expanded data and feedback data are obtained according to the identification result; the feedback The data includes a gap between the generated image and the real image, which has not yet met the requirements of the real picture; the expanded data is added to the expanded fundus image dataset, and the feedback data is fed back to the generative model for model optimization, and the probability of the generated image being a real image is judged by manual recognition. The generative model is continuously iterated interactively until the process converges, that is, the generative model can generate a real image, and the optimization of the generative model is realized. The new expanded fundus image dataset is generated by the optimized generative model; the feedback data and expanded data feedback and expansion process are repeated according to the new expanded fundus image dataset to obtain the expanded fundus image dataset, thereby expanding the sample size and providing rich data support for the subsequent construction of the basic network set training module, improving the output results of the basic network set training module, and thereby improving the accuracy of diabetic retinopathy image lesion segmentation.
[0032] S300: Inputting the expanded fundus image dataset into the encoder of the basic network set training module to obtain n overlapping image blocks of a predetermined size, where n is a positive integer greater than 1;
[0033] S400: performing feature matching on the n overlapping image blocks by the encoder to obtain segmentation multi-level features;
[0034] Specifically, the expanded fundus image dataset is fed into the encoder of the base network training module. The encoder then divides the input fundus image of a given size into n overlapping image blocks of a predetermined size, where n is a positive integer greater than 1. Dividing the fundus image of a given size into smaller blocks is beneficial for predicting fundus images with densely populated lesions. These blocks are then used as input to the encoder, which performs feature matching on the n overlapping blocks to generate multi-level segmentation features. These features can effectively improve the global performance of semantic segmentation.
[0035] S500: merging the n overlapping image blocks through the basic network set training module and recording merging parameters;
[0036] S600: performing feature matching according to the merging result of the n image blocks, and generating a local continuity feature by using the merging parameters and the feature matching result;
[0037] Further, such as Figure 3 As shown, an implementation of step S500 in the method provided in an embodiment of the present application includes:
[0038] S510: Obtaining overlapping image block size data according to the predetermined size;
[0039] S520: Setting the stride of adjacent blocks according to the overlapping image block size data to obtain a stride parameter;
[0040] S530: Calculate padding size data according to the stride parameter and the overlapping image block size data, and record the image block size data, stride parameter and padding size data as the merging parameter.
[0041] Specifically, in order to generate local continuity features, the n overlapping image blocks are merged through the basic network set training module, and the merging parameters are recorded; specifically, according to the predetermined size division rule of the input fundus image, the size data of the overlapping image blocks after division is obtained; the predetermined size can be determined according to demand or empirical value; the stride between adjacent overlapping image blocks is set to obtain the stride parameter, and the filling size data is calculated according to the stride parameter and the overlapping image block size data, and the image block size data, stride parameter and the filling size data are recorded as the merging parameter. Further, the local continuity feature is generated by the merging parameter and the feature matching result.
[0042] S700: Inputting the segmented multi-level features and the local continuity features into the decoder of the basic network set training module for multi-scale feature integration, and verifying the multi-scale feature integration result based on the constraint result corresponding to the expanded fundus image dataset. When the verification passes, the construction of the basic network set training module is completed;
[0043] S800: Inputting image data into the basic network set training module to obtain lesion recognition and segmentation results.
[0044] Specifically, the segmented multi-level features and local continuity features are used as training data to input into the decoder of the basic network set training module for multi-scale feature integration, and the multi-scale feature integration results are verified based on the constraint results corresponding to the expanded fundus image data set, that is, the constraint results corresponding to the expanded fundus image data set are used as supervision data, so that the basic network set training module reaches convergence. When the verification passes, the construction of the basic network set training module is completed, and the converged basic network set training module is used to perform lesion segmentation, and the image data is input into the basic network set training module to obtain the lesion recognition and segmentation identification results. The basic network set training module after deep learning is used to achieve the technical effect of real-time fine segmentation of different types of diabetic retinopathy lesions.
[0045] Furthermore, the method further comprises:
[0046] S900: Construct an attention constraint layer, wherein the attention constraint layer is calculated as:
[0047]
[0048] Among them, Q, K, and V are matrix vectors obtained by linear transformation of the embedding vector output of the previous layer; T is the matrix transpose; d h is the vector dimension of Q, K corresponding to each head of the self-attention mechanism, and the attention constraint layer is added to the basic network set training module.
[0049] Specifically, a large number of encoder calculations are concentrated on the attention constraint layer to construct the attention constraint layer. The attention constraint layer calculation is expressed as:
[0050]
[0051] Among them, Q, K, and V are matrix vectors obtained by linear transformation of the embedding vector output of the previous layer; T is the matrix transpose; d h is the vector dimension of Q, K corresponding to each head of the self-attention mechanism, and the attention constraint layer is added to the basic network set training module.
[0052] Furthermore, an implementation of step S300 in the method provided in the embodiment of the present application includes:
[0053] S310: performing resolution determination on the expanded fundus image dataset to obtain a determination result;
[0054] S320: When there is an image within a predetermined resolution range in the discrimination result, introducing a zoom ratio parameter;
[0055] S330: Optimize the processing sequence length according to the scaling parameter. The optimization formula is as follows:
[0056]
[0057] K1=Linear(C·R,C)(K′)
[0058] where K′ is the reshape sequence, K1 is the optimization sequence, R is the scaling parameter, C is the number of channels, N = H × W, H is the image width, and W is the image height. Linear(C·R,C) refers to a linear layer that takes a C·R-dimensional tensor as input and produces a C-dimensional tensor as output.
[0059] Specifically, the attention constraint layer calculation method established in the above steps has a high computational cost for images with higher resolutions. In order to effectively reduce the computational cost when processing images, the expanded fundus image dataset is subjected to resolution discrimination to obtain a discrimination result. When an image within a predetermined resolution range exists in the discrimination result, a scaling ratio parameter is introduced. The processing sequence is optimized according to the scaling ratio parameter to reduce the length of the processing sequence. The optimization formula is as follows:
[0060]
[0061] K1=Linear(C·R,C)(K′)
[0062] Here, K′ is the reshape sequence, K1 is the optimization sequence, R is the scaling parameter, C is the number of channels, N = H × W, where H is the image width and W is the image height. Linear(C·R,C) refers to a linear layer that takes a C·R-dimensional tensor as input and produces a C-dimensional tensor as output. The introduction of the scaling factor effectively reduces the computational overhead when processing images.
[0063] Furthermore, the basic network set training module also includes Mix-FFN, and Mix-FFN is:
[0064] X out =MLP(GELU(Conv 3×3 (MLP(X in ))))+X in
[0065] Among them, X in is the self-attention module feature, MLP is the multi-layer perceptron, X-out is the feature output by MLP, Conv 3×3 is a 3×3 convolution, and GELU is a Gaussian error linear unit.
[0066] Specifically, to optimize the effect of zero padding on leaked position information, a Mix-FFN is used in the network. By adding 3×3 convolutions to the feedforward network (FFN), the effect of zero padding on leaked position information is optimized. Mix-FFN combines 3×3 convolutions with MLP in each FFN. Mix-FFN can be expressed as:
[0067] X out =MLP(GELU(Conv 3×3 (MLP(X in ))))+X in
[0068] Among them, X in is the self-attention module feature, MLP is the multi-layer perceptron, X-out is the feature output by MLP, Conv 3×3 is a 3×3 convolution, and GELU is a Gaussian error linear unit.
[0069] MLP is a multi-layer perceptron, which is a feed-forward artificial neural network model. In addition to the input and output layers, there can be multiple hidden layers in the middle. The layers of MLP are fully connected. 3×3 The 3×3 convolution provides feature location information to the network while reducing the number of parameters and improving efficiency. GELU, a Gaussian Error Linear Unit, introduces the concept of random regularity into activations. Compared to Relus and ELUs, GELU activation transformations have the advantage of random dependency inputs.
[0070] In summary, the deep learning-based diabetic retinopathy image lesion segmentation and identification method provided in the embodiments of the present application has the following technical effects:
[0071] The embodiment of the present application provides a method for segmenting and identifying diabetic retinopathy images based on deep learning. The method calls a pre-stored fundus image dataset through the intelligent lesion segmentation and identification system, wherein the fundus image dataset includes hemorrhage, microaneurysm, hard exudation, soft exudation, neovascular fibroproliferative membrane, and microvascular abnormal diabetic retinopathy lesions; and performs sample expansion of the fundus image dataset according to the lesion type to obtain an expanded fundus image dataset; inputs the expanded fundus image dataset into the encoder of the basic network set training module to obtain n overlapping image blocks of a predetermined size, wherein n is a positive integer greater than 1; and performs special segmentation and identification on the n overlapping image blocks through the encoder. The method comprises the following steps: performing feature matching to obtain segmentation multi-level features; merging the n overlapping image blocks through the basic network set training module, and recording the merging parameters; performing feature matching based on the merging results of the n image blocks, and generating local continuity features through the merging parameters and feature matching results; inputting the segmentation multi-level features and the local continuity features into the decoder of the basic network set training module for multi-scale feature integration, and verifying the multi-scale feature integration results based on the constraint results corresponding to the expanded fundus image data set. When the verification passes, the construction of the basic network set training module is completed; inputting the image data into the basic network set training module to obtain the lesion recognition and segmentation identification results. The method solves the problems of insufficient performance in fine analysis of fundus diabetic retinopathy images and performance limitations of processing equipment in the prior art, and the problem that deep learning models still have timeliness and low accuracy, and achieves the technical effect of real-time fine segmentation of different types of diabetic retinopathy lesions.
[0072] In order to expand the number of fundus images and increase the number of training samples to obtain a high-precision diabetic retinopathy lesion segmentation model, in an embodiment of the present application, a generative model is constructed to achieve sample expansion of the fundus image dataset according to the lesion type to obtain an expanded fundus image dataset.
[0073] The embodiment of the present application generates overlapping image blocks by dividing the image into predetermined sizes, thereby meeting the prediction requirements for fundus images with dense lesions and satisfying the refined segmentation processing of various types of fundus images.
[0074] The embodiment of the present application effectively reduces the computational cost of image processing by introducing a scaling ratio parameter; by adopting Mix-FFN in the basic network set training module, the impact of zero filling on leakage position information is optimized, thereby achieving the technical effect of real-time fine segmentation of different types of diabetic retinopathy lesions.
[0075] Example 2
[0076] Based on the same inventive concept as the method for segmenting and identifying diabetic retinopathy images based on deep learning in the aforementioned embodiment, Figure 4As shown, the present application provides a diabetic retinopathy image lesion segmentation and identification system based on deep learning, wherein the system includes:
[0077] An information calling unit 100 is configured to call a pre-stored fundus image dataset through an intelligent lesion segmentation and identification system, wherein the fundus image dataset includes hemorrhage, microaneurysm, hard exudate, soft exudate, neovascular fibroproliferative membrane, and microvascular abnormal diabetic retinopathy lesions;
[0078] A sample expansion unit 200 is configured to expand the sample of the fundus image dataset according to the lesion type to obtain an expanded fundus image dataset;
[0079] An input unit 300 is configured to input the expanded fundus image dataset into an encoder of a basic network set training module to obtain n overlapping image blocks of a predetermined size, where n is a positive integer greater than 1;
[0080] A matching unit 400 is configured to perform feature matching on the n overlapping image blocks through the encoder to obtain segmentation multi-level features;
[0081] A merging unit 500 is configured to merge the n overlapping image blocks using the basic network set training module and record merging parameters;
[0082] A feature matching unit 600 is configured to perform feature matching based on the merging results of the n image blocks, and to generate a local continuity feature using the merging parameters and the feature matching results;
[0083] A construction unit 700 is configured to input the segmented multi-level features and the local continuity features into a decoder of the basic network set training module for multi-scale feature integration, and verify the multi-scale feature integration result based on the constraint result corresponding to the expanded fundus image dataset. When the verification passes, the construction of the basic network set training module is completed.
[0084] The processing unit 800 is used to input the image data into the basic network set training module to obtain lesion recognition and segmentation results.
[0085] Furthermore, the sample expansion unit in the system is further configured to:
[0086] Constructing a generative model, wherein the generative model is a model for image generation;
[0087] Inputting the fundus image dataset and random noise into the generation model to obtain a pre-expanded fundus image dataset;
[0088] Marking the pre-expanded fundus image data set, and obtaining expanded data and feedback data according to the marking result;
[0089] Adding the expanded data to the expanded fundus image dataset, feeding the feedback data back to the generation model for model optimization, and generating a new expanded fundus image dataset using the optimized generation model;
[0090] The process of feedback and expansion of the feedback data and the expansion data is repeated according to the new expanded fundus image dataset to obtain the expanded fundus image dataset.
[0091] Furthermore, the merging unit in the system is also used for:
[0092] Obtaining overlapping image block size data according to the predetermined size;
[0093] Setting the stride of adjacent blocks according to the overlapping image block size data to obtain a stride parameter;
[0094] The padding size data is calculated according to the stride parameter and the overlapping image block size data, and the image block size data, the stride parameter and the padding size data are recorded as the merging parameters.
[0095] Furthermore, the system further comprises:
[0096] A constraint layer construction unit is used to construct an attention constraint layer, wherein the attention constraint layer is calculated as:
[0097]
[0098] Among them, Q, K, and V are matrix vectors obtained by linear transformation of the embedding vector output of the previous layer; T is the matrix transpose; d h is the vector dimension of Q, K corresponding to each head of the self-attention mechanism, and the attention constraint layer is added to the basic network set training module.
[0099] Furthermore, the input unit in the system is also used for:
[0100] Performing resolution discrimination on the expanded fundus image data set to obtain a discrimination result;
[0101] When there is an image within a predetermined resolution range in the discrimination result, introducing a zoom ratio parameter;
[0102] The processing sequence length is optimized according to the scaling parameters, and the optimization formula is as follows:
[0103]
[0104] K1=Linear(C·R,C)(K′)
[0105] where K′ is the reshape sequence, K1 is the optimization sequence, R is the scaling parameter, C is the number of channels, N = H × W, H is the image width, and W is the image height. Linear(C·R,C) refers to a linear layer that takes a C·R-dimensional tensor as input and produces a C-dimensional tensor as output.
[0106] Furthermore, the basic network set training module in the system also includes Mix-FFN, and Mix-FFN is:
[0107] X out =MLP(GELU(Conv 3×3 (MLP(X in ))))+X in
[0108] Among them, X in is the self-attention module feature, MLP is the multi-layer perceptron, X-out is the feature output by MLP, Conv 3×3 is a 3×3 convolution, and GELU is a Gaussian error linear unit.
[0109] The specific working process of the modules disclosed in the above embodiments of this application can be found in the corresponding method embodiments and will not be repeated here.
[0110] The present application is capable of being implemented or used by those skilled in the art. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to be embodied in the widest possible manner consistent with the principles and novel features disclosed herein.
Claims
1. A deep learning-based diabetic retinopathy image lesion segmentation and identification method, characterized by: The method is applied to an intelligent lesion segmentation and identification system, wherein the intelligent lesion segmentation and identification system is communicatively connected to a basic network set training module, and the method comprises: The intelligent lesion segmentation and identification system is used to call a pre-stored fundus image dataset, wherein the fundus image dataset includes hemorrhage, microaneurysm, hard exudate, soft exudate, neovascular fibroproliferative membrane, and microvascular abnormal diabetic retinopathy lesions; and expanding the sample of the fundus image dataset according to the lesion type to obtain an expanded fundus image dataset; Inputting the expanded fundus image dataset into the encoder of the basic network set training module to obtain n overlapping image blocks of a predetermined size, wherein n is a positive integer greater than 1; Performing feature matching on the n overlapping image blocks by the encoder to obtain segmentation multi-level features; Merging the n overlapping image blocks through the basic network training set module and recording the merging parameters; Performing feature matching based on the merging results of the n image blocks, and generating local continuity features through the merging parameters and the feature matching results; Inputting the segmented multi-level features and the local continuity features into the decoder of the basic network set training module for multi-scale feature integration, and verifying the multi-scale feature integration result based on the constraint result corresponding to the expanded fundus image dataset. When the verification passes, the construction of the basic network set training module is completed; The image data is input into the basic network set training module to obtain the lesion recognition and segmentation identification results.
2. The method according to claim 1, wherein The method further comprises: Constructing a generative model, wherein the generative model is a model for image generation; Inputting the fundus image dataset and random noise into the generation model to obtain a pre-expanded fundus image dataset; Marking the pre-expanded fundus image data set, and obtaining expanded data and feedback data according to the marking result; Adding the expanded data to the expanded fundus image dataset, feeding the feedback data back to the generation model for model optimization, and generating a new expanded fundus image dataset using the optimized generation model; The process of feedback and expansion of the feedback data and the expansion data is repeated according to the new expanded fundus image dataset to obtain the expanded fundus image dataset.
3. The method according to claim 1, wherein The merging of the n overlapping image blocks by the basic network training set module and recording the merging parameters further includes: obtaining overlapping image block size data according to the predetermined size; Setting the stride of adjacent blocks according to the overlapping image block size data to obtain a stride parameter; The padding size data is calculated according to the stride parameter and the overlapping image block size data, and the image block size data, the stride parameter and the padding size data are recorded as the merging parameters.
4. The method according to claim 1, wherein The method further comprises: Construct an attention constraint layer, where the attention constraint layer calculation is expressed as: Among them, Q, K, and V are matrix vectors obtained by linear transformation of the embedding vector output of the previous layer; T is the matrix transpose; d h is the vector dimension of Q, K corresponding to each head of the self-attention mechanism, and the attention constraint layer is added to the basic network set training module.
5. The method according to claim 4, wherein The step of inputting the expanded fundus image data set into the encoder further includes: Performing resolution discrimination on the expanded fundus image data set to obtain a discrimination result; When there is an image within a predetermined resolution range in the discrimination result, a scaling parameter is introduced; The sequence length is optimized according to the scaling parameters, and the optimization formula is as follows: K1=Linear(C·R,C)(K′); where K′ is the reshape sequence, K1 is the optimization sequence, R is the scaling parameter, C is the number of channels, N = H × W, H is the image width, and W is the image height. Linear(C·R,C) refers to a linear layer that takes a C·R-dimensional tensor as input and produces a C-dimensional tensor as output.
6. The method according to claim 5, wherein The basic network set training module also includes Mix-FFN, and Mix-FFN is: X-out = MLP(GELU(Conv 3×3 (MLP(X in ))))+X in Among them, X in is the self-attention module feature, MLP is the multi-layer perceptron, X-out is the feature output by MLP, Conv 3×3 is a 3×3 convolution, and GELU is a Gaussian error linear unit.
7. A deep learning-based diabetic retinopathy image lesion segmentation and identification system, characterized by: The system comprises: An information calling unit, the information calling unit being used to call a pre-stored fundus image dataset through an intelligent lesion segmentation and identification system, wherein the fundus image dataset includes hemorrhage, microaneurysm, hard exudate, soft exudate, neovascular fibroproliferative membrane, and microvascular abnormal diabetic retinopathy lesions; a sample expansion unit, configured to expand the sample of the fundus image dataset according to the lesion type to obtain an expanded fundus image dataset; An input unit, configured to input the expanded fundus image dataset into an encoder of a basic network set training module to obtain n overlapping image blocks of a predetermined size, where n is a positive integer greater than 1; a matching unit, configured to perform feature matching on the n overlapping image blocks through the encoder to obtain segmentation multi-level features; a merging unit, configured to merge the n overlapping image blocks through the basic network set training module and record merging parameters; a feature matching unit, wherein the feature matching unit performs feature matching based on a merging result of the n image blocks, and generates a local continuity feature through the merging parameter and the feature matching result; A construction unit, the construction unit being configured to input the segmented multi-level features and the local continuity features into a decoder of the basic network set training module for multi-scale feature integration, and verify the multi-scale feature integration result based on a constraint result corresponding to the expanded fundus image dataset, and when the verification passes, completing the construction of the basic network set training module; A processing unit is used to input image data into the basic network set training module to obtain lesion recognition and segmentation identification results.