Method, system, medium, terminal and program product for identifying fundus disease types based on visual transformer model
Through the fundus disease type identification method based on the vision transformer model, the problem of inefficiency and limited application of the prior art in identifying fundus diseases in communities and primary care clinics is solved, and efficient identification and classification of multiple fundus diseases is achieved.
Patent Information
- Application Number
- CN202411876980.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-12-19
AI Technical Summary
The prior art is poor when used in community screening and primary health clinics for initial screening of fundus diseases, and the model has limited application scope and cannot effectively identify multiple fundus diseases.
The fundus disease type recognition method based on the vision transformer model is adopted to obtain and preprocess the fundus image of the unmyodilated fundus, and deploy it in the community or primary care clinic to realize real-time identification of multiple fundus diseases.
Effectively identify and classify multiple fundus diseases, reduce the life impact of primary screeners, and improve the efficiency of fundus disease screening in communities and primary care clinics.
Smart Images

Figure CN119339168B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method, system, medium, terminal and program product for identifying fundus disease types based on a visual transformer model. Background Art
[0002] In recent years, artificial intelligence (AI) models have reached or even exceeded the diagnostic performance of doctors in automatically detecting diabetic retinopathy (DR), age-related macular degeneration (AMD), optic nerve abnormalities, and a variety of eye diseases based on fundus photographs. However, the advantages of AI in detecting eye diseases are often not fully demonstrated in real-world settings.
[0003] Current AI models are mainly trained on high-quality images captured by strict research protocols, which are usually obtained after the pupil is dilated. However, when these models are applied to community screening and primary care clinics, the performance of these models will drop significantly because the photos obtained in these places are fundus photos without pupil dilation. In addition, current AI models are often targeted at a specific type of fundus disease, which reduces the scope of applicability of the models.
[0004] Therefore, it is necessary to provide a method, system, medium, terminal and program product for identifying fundus disease types based on a visual transformer model to solve the above-mentioned problems existing in the prior art. Summary of the invention
[0005] In view of the shortcomings of the prior art mentioned above, the purpose of the present application is to provide a method, system, medium, terminal and program product for identifying fundus disease types based on a visual transformer model, so as to solve the technical problem that the prior art cannot be applied to preliminary screening in scenarios such as community screening and primary care clinics.
[0006] To achieve the above-mentioned purpose and other related purposes, the first aspect of the present application provides a method for identifying fundus disease types based on a visual transformer model, comprising:
[0007] Obtaining fundus disease images screened by the image quality screening model, and preprocessing the fundus disease images to form a fundus disease image data set; each fundus disease image in the fundus disease image data set has been labeled with a corresponding fundus disease;
[0008] Inputting the fundus disease image data set into a visual transformer model for training to construct a fundus disease type recognition model;
[0009] Deploy the fundus disease type recognition model to identify the fundus disease in real time according to the currently input fundus disease image;
[0010] The training process of the image quality screening model includes: acquiring a non-dilated fundus image, and annotating the fundus image with a quality label to form a fundus image quality data set; the quality label includes a field of view centered on the macula and the optic nerve head, macula visibility, optic nerve head visibility, eyelid artifacts, camera artifacts, media opacity, blood vessel visibility, and any one or more combinations of illumination; inputting the fundus image quality data set into an image classification model for training to obtain a trained image quality screening model; the image quality screening model is used to perform quality screening on the currently input fundus image; and the image quality screening model is connected to the fundus disease type recognition model to input the screened fundus disease image into the fundus disease type recognition model.
[0011] In some embodiments of the first aspect of the present application, the training process of the fundus disease type recognition model includes: transforming each fundus disease image in the fundus disease image data set based on the embedding layer of the visual transformer model, and adding a classification mark sequence and a position code to obtain an embedding sequence of fundus disease image blocks; inputting the embedding sequence of fundus disease image blocks into the encoder of the visual transformer model for feature extraction of the fundus disease image blocks to obtain a feature sequence; after removing the features of the classification mark position from the feature sequence, performing global pooling to generate a global feature vector, and inputting the global feature vector into the multi-layer perceptron of the visual transformer model to output the fundus disease type recognition result.
[0012] In some embodiments of the first aspect of the present application, the embedding layer based on the visual transformer model transforms each fundus disease image in the fundus disease image dataset, and adds a classification mark sequence and a position code to obtain a fundus disease image block embedding sequence, which includes: dividing the fundus disease image into a number of image blocks based on a convolutional neural network, and flattening each image block to form an image block embedding sequence; inserting a classification mark sequence before the image block embedding sequence to obtain an image block embedding sequence containing classification information and image features; adding a position code to each embedding vector in the image block embedding sequence containing classification information and image features, and performing a discarding operation to obtain a fundus disease image block embedding sequence.
[0013] In some embodiments of the first aspect of the present application, the fundus disease image block embedding sequence is input into the encoder of the visual transformer model for feature extraction of the fundus disease image block to obtain a feature sequence, which includes: performing a layer normalization operation on the fundus disease image block embedding sequence to obtain a standardized embedding sequence; inputting the standardized embedding sequence into a multi-head attention mechanism to obtain an attention feature sequence and performing a discarding operation; adding the attention feature sequence to the fundus disease image block embedding sequence through a residual connection to obtain a first residual feature sequence; performing a layer normalization operation on the first residual feature sequence to obtain a standardized first residual feature sequence; inputting the standardized first residual feature sequence into a feedforward fully connected network to obtain a feedforward feature sequence; performing a discarding operation on the feedforward feature sequence and then adding it to the first residual feature sequence through a residual connection to obtain a second residual feature sequence; repeating the above-mentioned layer normalization, multi-head attention mechanism, residual connection, layer normalization, feedforward fully connected network, and residual connection operations multiple times to gradually extract features and obtain a feature sequence.
[0014] In some embodiments of the first aspect of the present application, the fundus disease labels annotated for the fundus disease images include: age-related macular degeneration, diabetic retinopathy, myopic maculopathy, suspected glaucoma, cataract, epiretinal membrane, macular hole, retinal vein occlusion, retinal artery occlusion, retinitis pigmentosa, retinal detachment, vitreous degeneration, and any one or more combinations of myelinated nerve fibers.
[0015] To achieve the above-mentioned purpose and other related purposes, the second aspect of the present application provides a fundus disease type recognition system based on a visual transformer model, comprising:
[0016] A data acquisition module, used to acquire fundus disease images screened by the image quality screening model, and pre-process the fundus disease images to form a fundus disease image data set; each fundus disease image in the fundus disease image data set has been labeled with a corresponding fundus disease;
[0017] A model building module, used for inputting the fundus disease image data set into a visual transformer model for training to build a fundus disease type recognition model;
[0018] A model deployment module, used to deploy the fundus disease type recognition model to perform real-time recognition of the fundus disease type according to the currently input fundus disease image;
[0019] The training process of the image quality screening model includes: acquiring a non-dilated fundus image, and annotating the fundus image with a quality label to form a fundus image quality data set; the quality label includes a field of view centered on the macula and the optic nerve head, macula visibility, optic nerve head visibility, eyelid artifacts, camera artifacts, media opacity, blood vessel visibility, and any one or more combinations of illumination; inputting the fundus image quality data set into an image classification model for training to obtain a trained image quality screening model; the image quality screening model is used to perform quality screening on the currently input fundus image; and the image quality screening model is connected to the fundus disease type recognition model to input the screened fundus disease image into the fundus disease type recognition model.
[0020] To achieve the above-mentioned purpose and other related purposes, the third aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, and the computer program implements the method when executed by a processor.
[0021] To achieve the above-mentioned purpose and other related purposes, the fourth aspect of the present application provides a computer program product, which includes a computer program code. When the computer program code runs on a computer, the computer implements the method.
[0022] To achieve the above-mentioned purpose and other related purposes, the fifth aspect of the present application provides an electronic terminal, including a memory, a processor and a computer program stored in the memory; the processor executes the computer program to implement the method.
[0023] As described above, the method, system, medium, terminal and program product for identifying fundus disease types based on the visual transformer model of the present application have the following beneficial effects:
[0024] By obtaining fundus disease images screened by an image quality screening model and preprocessing the fundus disease images, each fundus disease image is labeled with the corresponding fundus disease to form a fundus disease image dataset; the fundus disease image dataset is input into a visual transformer model for training, thereby constructing a fundus disease type recognition model; the trained fundus disease recognition model is deployed in specific application scenarios, such as communities or primary care clinics, and the deployed fundus disease type recognition model recognizes the fundus disease corresponding to the image based on the currently input fundus disease image. This application uses non-dilated fundus images to reduce the impact on the lives of initial screeners, and the fundus disease type recognition model obtained by training with the visual transformer model can effectively recognize a variety of labeled fundus disease types. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1Shown is a flowchart of a method for identifying fundus disease types based on a visual transformer model in one embodiment of the present application.
[0026] Figure 2 Shown is a flowchart of the training process of the fundus disease type recognition model in one embodiment of the present application.
[0027] Figure 3 Shown is a schematic diagram of a framework for identifying fundus disease types based on a visual transformer model in one embodiment of the present application.
[0028] Figure 4 Shown is a schematic diagram of the encoder framework of the visual transformer model in one embodiment of the present application.
[0029] Figure 5 Shown is a flowchart of the training process of the image quality screening model in one embodiment of the present application.
[0030] Figure 6 Shown is a structural schematic diagram of a fundus disease type identification system based on a visual transformer model in one embodiment of the present application.
[0031] Figure 7 Shown is a schematic diagram of the structure of an electronic terminal in one embodiment of the present application. DETAILED DESCRIPTION
[0032] The following describes the embodiments of the present application through specific examples, and those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.
[0033] In the embodiments of the present application, words such as "first" and "second" are used to distinguish the same or similar items with substantially the same functions and effects. For example, the first XX and the second XX are only used to distinguish different XXs, and do not limit their order. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit them to be different.
[0034] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" represent examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0035] In the embodiments of the present application, "at least one" refers to one or more, and "plurality" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can represent: a, b, c, ab, ac, bc or abc, where a, b, c can be single or multiple.
[0036] Before further describing the present invention in detail, the nouns and terms involved in the embodiments of the present invention are explained. The nouns and terms involved in the embodiments of the present invention are applicable to the following interpretations:
[0037] <1> Convolutional Neural Networks (CNN): A type of feedforward neural network that includes convolutional calculations and has a deep structure. It is a deep learning algorithm that is particularly effective in the fields of image processing and computer vision. CNN simulates the mechanism of the human visual system and can automatically learn and extract useful features from images to complete tasks such as image classification, target detection, and image segmentation.
[0038] To facilitate understanding of the embodiments of the present application, first Figure 1 Detailed description. Figure 1 The flowchart of a method for identifying fundus disease types based on a visual transformer model in an embodiment of the present invention is shown. The method for identifying fundus disease types based on a visual transformer model in this embodiment mainly includes the following steps:
[0039] Step S11: Acquire fundus disease images screened by the image quality screening model, and preprocess the fundus disease images to form a fundus disease image dataset; each fundus disease image in the fundus disease image dataset has been labeled with a corresponding fundus disease.
[0040] In some embodiments of the present application, the method of preprocessing the fundus disease images with qualified quality screened out by the image quality screening model includes: uniformly processing the collected fundus disease images so that the number of image channels is fixed to three RGB channels to ensure that different images have the same color information format, which is convenient for subsequent feature extraction and comparative analysis; at the same time, adjusting the image size to 520×520 pixels so that the fundus disease image is displayed in the center of the square frame to eliminate the image size differences caused by different acquisition devices or shooting conditions, ensure the consistency of input data, and contribute to the effective training of the model; professional ophthalmologists manually label each fundus disease image, and by providing label annotations for each fundus disease image, the model can clearly learn the characteristics of the image, and then make effective judgments in the classification task. After preprocessing the fundus disease images, a fundus disease image dataset is formed.
[0041] In some embodiments of the present application, the fundus disease labels annotated to the fundus disease images include: age-related macular degeneration, diabetic retinopathy, myopic maculopathy, suspected glaucoma, cataract, epiretinal membrane, macular hole, retinal vein occlusion, retinal artery occlusion, retinitis pigmentosa, retinal detachment, vitreous degeneration, myelinated nerve fibers, any one or more of the combination. By annotating the fundus disease images with the above 13 fundus diseases, the model can clearly learn the characteristics of the above 13 fundus diseases, and can accurately identify and classify these 13 fundus diseases in fundus disease recognition work.
[0042] In some embodiments of the present application, the fundus disease image data set is divided into a training set, a validation set, and a test set according to a certain ratio. Among them, the training set is used for model learning and parameter adjustment; the validation set is used to monitor and tune the overfitting during model training; the test set is used to independently evaluate the final performance of the model. By dividing the fundus disease image data set, it can be ensured that the model obtains sufficient learning data during the training stage. At the same time, through the allocation of validation sets and test sets, the generalization ability of the model and the robustness to new data are ensured, thereby achieving accurate classification and evaluation of fundus disease images. It should be noted that the fundus disease image data set is divided into a training set, a validation set, and a test set, for example, in a ratio of 7:1:2, and the specific ratio is not limited to this.
[0043] The image quality screening model in this embodiment identifies, distinguishes and describes various image artifacts. The specific training process will be described below and will not be repeated here.
[0044] Step S12: inputting the fundus disease image dataset into the visual transformer model for training to construct a fundus disease type recognition model.
[0045] It should be understood that the Vision Transformer (ViT) model is a deep learning model based on the Transformer architecture, which aims to apply the Transformer to computer vision tasks. It divides the image into multiple fixed-size image patches and regards these patches as sequence data, using the self-attention mechanism to capture the global and local information in the image.
[0046] In some embodiments of the present application, Figure 2 As shown, the training process of the fundus disease type recognition model includes:
[0047] Step S21: transforming each fundus disease image in the fundus disease image dataset based on the embedding layer of the visual transformer model, and adding a classification mark sequence and a position code to obtain an embedding sequence of fundus disease image blocks.
[0048] In some embodiments of the present application, step S21 includes: dividing the fundus disease image into a number of image blocks based on a convolutional neural network, and flattening each image block to form an image block embedding sequence; inserting a classification mark sequence before the image block embedding sequence to obtain an image block embedding sequence containing classification information and image features; adding a position code to each embedding vector in the image block embedding sequence containing classification information and image features, and performing a discarding operation to obtain a fundus disease image block embedding sequence.
[0049] Specifically, Figure 3 As shown in the figure, the size of the preprocessed fundus disease image is adjusted to 224×224 pixels to ensure that the image size meets the input requirements of the model, maintain data consistency and reduce errors caused by resolution differences. Then the adjusted fundus disease image is input into the convolutional neural network model, and the fundus disease image is convolved using a convolution kernel of size 16×16 and a step size of 16. This convolution is equivalent to dividing the entire image into 14×14 fixed image blocks of the same size of 16×16 pixels. Through this division method, the local spatial information in the image is retained, so that each image block can capture the detailed features of the image. Each 16×16 pixel image block is flattened, and all pixels in each image block are arranged in order of their positions and flattened into a one-dimensional vector. After flattening, each image block forms a one-dimensional embedding vector containing all pixel values. These one-dimensional embedding vectors are arranged in sequence to form a complete image block embedding sequence. This processing method can convert the information of local image blocks into feature vectors, providing input for feature extraction of subsequent models.
[0050] An additional classification tag sequence is inserted at the front of the image block embedding sequence to capture the feature representation of the global image and the image classification information, thereby obtaining an image block embedding sequence containing classification information and image features. This classification tag sequence will be used to learn the overall features of the entire image during the model training process in order to perform the final classification task.
[0051] Subsequently, a learnable position encoding vector corresponding to its position is added to each embedding vector in the image block embedding sequence containing classification information and image features, thereby obtaining the input fundus disease image block embedding sequence. The addition of position encoding enables the model to understand the relative position relationship of image blocks, thereby capturing the spatial structural information of the image, overcoming the shortcomings of the fully connected network that lacks spatial position perception. Position encoding is learnable, which means that during the training process, the model can further optimize the parameters of position encoding according to the characteristics of the data to better describe the position relationship of image blocks. In order to improve the generalization ability of the model and reduce overfitting, a dropout operation is performed. The dropout operation randomly sets the elements in the embedding vector to zero with a certain probability, forcing the model to make effective predictions in the absence of certain features, thereby improving the robustness of the model.
[0052] Step S22: inputting the fundus disease image block embedding sequence into the encoder of the visual transformer model to extract features of the fundus disease image block to obtain a feature sequence.
[0053] In some embodiments of the present application, step S22 includes: performing a layer normalization operation on the embedding sequence of the fundus disease image blocks to obtain a standardized embedding sequence; inputting the standardized embedding sequence into a multi-head attention mechanism to obtain an attention feature sequence and performing a discarding operation; adding the attention feature sequence to the embedding sequence of the fundus disease image blocks through a residual connection to obtain a first residual feature sequence; performing a layer normalization operation on the first residual feature sequence to obtain a standardized first residual feature sequence; inputting the standardized first residual feature sequence into a feedforward fully connected network to obtain a feedforward feature sequence; performing a discarding operation on the feedforward feature sequence and then adding it to the first residual feature sequence through a residual connection to obtain a second residual feature sequence; repeating the above-mentioned layer normalization, multi-head attention mechanism, residual connection, layer normalization, feedforward fully connected network, and residual connection operations multiple times to gradually extract features and obtain a feature sequence.
[0054] Specifically, Figure 3 and Figure 4As shown, the embedding sequence of fundus disease image blocks is subjected to layer normalization operation to obtain a standardized embedding sequence. The layer normalization operation normalizes each feature vector of the embedding sequence of fundus disease image blocks so that the mean of all features is zero and the variance is one. In this way, it helps to alleviate the numerical instability problem of the network model during training, prevent gradient vanishing or gradient explosion, thereby accelerating convergence and improving the stability of training. Especially in the visual transformer model, the layer normalization operation can make the features between different levels consistent in scale, which is conducive to the effective training of the visual transformer model.
[0055] The standardized embedding sequence is input into the multi-head attention mechanism to calculate the global dependencies between the embedding vectors in the sequence. The standardized embedding sequence is divided into multiple attention heads according to the set number of heads, and each head is mapped into a query (Q) vector, a key (K) vector, and a value (V) vector through a linear transformation matrix. Among them, the query vector is used to calculate the similarity with other key vectors to measure the degree of attention of each input to other inputs; the key vector and the query vector work together to calculate the similarity measure; the value vector represents the information that needs to be paid attention to, and finally the attention output is obtained through weighted aggregation. The formula for obtaining the query vector, key vector, and value vector is as follows:
[0056] ;Formula (1)
[0057] Where X is the input matrix; is the query weight matrix; is the key weight matrix; is the value weight matrix; , , are all learnable weight matrices, R represents a set of real numbers, D is the dimension of the embedding vector, and d k is the dimension of each attention head, which is equal to D / h, where h is the number of attention heads.
[0058] The attention weights are calculated using the query vector, key vector, and value vector. For each attention head, the similarity between the query and the key is calculated. The query vector Q and the transpose of the key vector Perform a dot product operation to obtain the similarity score between each query and each key, that is, This similarity score indicates how much attention each element in the embedding sequence pays to other elements. Divide the similarity score by , to scale it, which is used to prevent the result value of the dot product from being too large in high-dimensional situations, so that the output probability after the softmax function will tend to be extreme, affecting the learning effect. The scaled similarity score is subjected to a softmax operation to obtain a normalized attention weight matrix. Each weight represents the degree of dependence of an element on other elements in the input embedding sequence. Multiply the attention weight matrix with the value (V) vector to obtain the feature representation of each attention head for output. This output represents the information extracted by each element based on paying attention to other elements. The basic formula is:
[0059] ;Formula (2)
[0060] It should be understood that the softmax function is an activation function widely used in machine learning and deep learning, suitable for the output layer of multi-classification problems. It converts a vector containing arbitrary real numbers into a probability distribution, that is, each element in the vector is mapped between 0 and 1, and the sum of all elements is 1.
[0061] Concatenate the feature representations of all attention heads to obtain the overall output of the multi-head attention, that is, the output attention feature sequence. The basic formula is:
[0062] ;Formula (3)
[0063] in, Represents the output of a single attention head; the concatenation operation concatenates the output vectors of each attention head to form a dimension of The matrix of represents the length of the sequence, is the output dimension of each attention head, Represents the dimension after concatenation; then through the linear projection matrix Mapped to the attention feature sequence of the original input dimension D. Finally, the elements in the embedding vector are randomly set to zero after the discard operation.
[0064] The attention feature sequence is added to the original input fundus disease image block embedding sequence through residual connection to form the first residual feature sequence. This design helps to alleviate the problem of gradient vanishing in deep neural networks and ensures that the gradient can be effectively propagated, enabling it to capture more complex features.
[0065] A layer normalization operation is performed on the first residual feature sequence to obtain a standardized first residual feature sequence.
[0066] The standardized first residual feature sequence is input into the feedforward fully connected network to obtain the feedforward feature sequence. Specifically, the standardized first residual feature sequence is linearly transformed through the first fully connected layer, and the feature dimension is mapped from the input size to the hidden layer size to obtain the feedforward feature. This operation expands the feature dimension to a higher-dimensional feature space for feature extraction, so that the model can capture more complex features and improve the model's expressiveness; then the GELU activation function is applied to the feedforward feature, and the feature is nonlinearly activated to obtain the activated feedforward feature, and the discard operation is performed. In this step, the GELU is a smooth nonlinear activation function that can make the gradient change of the network more continuous and prevent numerical instability and gradient explosion and disappearance problems; then the activated feedforward feature is linearly transformed through the second fully connected layer, and the feature dimension is mapped back to the input size, and finally the feedforward feature sequence is obtained. The purpose of mapping back to the input size in this step is to facilitate the subsequent layers to receive inputs with the same feature dimension and maintain the consistency of the network structure. It should be understood that GELU (Gaussian Error Linear Unit) is an activation function whose basic concept is to achieve nonlinearity by multiplying the input by a random Bernoulli distribution, whose value probability is associated with the input value, thereby simulating the behavior of natural neurons.
[0067] After the above feedforward feature sequence is discarded, it is added to the first residual feature sequence through residual connection to obtain the second residual feature sequence. The features after the feedforward fully connected network will be richer, but the features originally extracted from the self-attention mechanism may be lost. Through the second residual connection, the global features of the self-attention module can be fused with the deep features of the feedforward fully connected network, thereby retaining more original global information and adding the nonlinear features extracted by the feedforward fully connected network.
[0068] The above steps of layer normalization, multi-head attention mechanism, residual connection, layer normalization, feedforward fully connected network, and residual connection are stacked multiple times as needed to extract deeper features layer by layer, and finally obtain a feature sequence containing global and local information. By continuously stacking these modules, the model can abstract the input features layer by layer, from shallow low-level features such as edges and textures to high-level patterns, structures, and global information, gradually capturing the multi-level feature expression of the image. The feature extraction of each layer is performed on the basis of the previous layer. As the network deepens, the dimension and content of the features are also richer, thus forming a more representative and discriminative feature representation, and finally obtaining a feature sequence containing global and local information.
[0069] Step S23: After removing the features of the classification mark position from the feature sequence, global pooling is performed to generate a global feature vector, and the global feature vector is input into the multi-layer perceptron of the visual transformer model to output the fundus disease type recognition result.
[0070] Specifically, Figure 3 As shown in the figure, firstly, the feature sequence is subjected to layer normalization to obtain a standardized feature sequence. The classification mark feature is removed from the standardized feature sequence, and the remaining image block features are retained. The classification mark captures the global features of the entire input image. The classification mark is added at the front of the input feature sequence to represent the input feature expression. The classification mark is input into the classification head to obtain the classification result, which can achieve a good classification effect. However, the classification mark relies on the multi-layer self-attention mechanism of the Transformer to capture information. It does not represent the information of any local image block itself, but the result of interaction with all image blocks. The final result may tend to those places with strong attention mechanisms, and the attention to insignificant features may not be fully captured. Removing the classification mark and performing global pooling on the remaining image block features can better utilize the features of all image blocks, taking the average of the features of all blocks to generate a comprehensive global feature vector. This global feature vector integrates local information well, can avoid ignoring or losing detail information in the attention mechanism, helps to retain more image details, and improves the accuracy of classification. The features of each image block in the fundus disease image may contain information from different regions, such as the macula region, the optic nerve head region, etc. By global pooling, all image block features are averaged and aggregated, which can more evenly integrate the information of each image block, thereby more comprehensively reflecting the characteristics of the entire image. However, the classification tag depends on the update of the multi-head self-attention mechanism, and may not be able to capture all local feature information evenly. In other classifications, you can choose to use classification tags or global pooling as required. In this embodiment, the global feature vector obtained after global pooling is input into a multi-layer perceptron, and the feature dimension is mapped from the embedding dimension to the number of categories. Such an output can obtain a score for each category, thereby outputting the result of fundus disease type recognition.
[0071] Step S13: deploying the fundus disease type recognition model to perform real-time recognition of fundus diseases according to the currently input fundus disease image.
[0072] Specifically, the trained fundus disease type recognition model is deployed in the community or primary care clinics for real-time recognition of fundus diseases based on the currently input fundus disease images, thereby identifying and classifying 13 fundus diseases including age-related macular degeneration, diabetic retinopathy, myopic maculopathy, suspected glaucoma, cataracts, epiretinal membrane, macular hole, retinal vein occlusion, retinal artery occlusion, retinitis pigmentosa, retinal detachment, vitreous degeneration, and myelinated nerve fibers.
[0073] In some embodiments of the present application, the training process of the image quality screening model includes: step S51: obtaining a non-dilated fundus image, and annotating the fundus image with a quality label to form a fundus image quality data set; the quality label includes a combination of any one or more of the field of view centered on the macula and the optic nerve head, macula visibility, optic nerve head visibility, eyelid artifacts, camera artifacts, media opacity, blood vessel visibility, and illumination; step S52: inputting the fundus image quality data set into an image classification model for training to obtain a trained image quality screening model; the image quality screening model is used to perform quality screening on the currently input fundus image; step S53: connecting the image quality screening model to the fundus disease type recognition model to input the screened fundus disease image into the fundus disease type recognition model.
[0074] It should be understood that in the community or primary care clinic, only fundus images without mydriasis are collected for the initial screening. If fundus images with mydriasis are collected for the initial screening, it will affect the daily life of the initial screening and is not suitable for the initial screening. Therefore, the fundus images collected in this application are fundus images without mydriasis, so as to conduct preliminary screening of fundus diseases in the community or primary care clinic.
[0075] Specifically, low-quality fundus images were collected in 240 community screenings in a large city, that is, these low-quality fundus images were fundus images without pupil dilation and mydriasis. The acquired fundus images were preprocessed and labeled with quality labels, so that qualified fundus disease images were screened out according to the quality labels as input to the fundus disease type recognition model.
[0076] The image classification model in this embodiment may also select a visual transformer model, and the specific training process of the image quality screening model is as described above, which will not be repeated here.
[0077] Figure 6 is a schematic block diagram of a fundus disease type identification system based on a visual transformer model provided in an embodiment of the present application. Figure 6 As shown, the fundus disease type recognition system 600 based on the visual transformer model includes:
[0078] The data acquisition module 601 is used to acquire the fundus disease images screened by the image quality screening model, and pre-process the fundus disease images to form a fundus disease image data set; each fundus disease image in the fundus disease image data set has been labeled with the corresponding fundus disease;
[0079] A model building module 602 is used to input the fundus disease image data set into a visual transformer model for training to build a fundus disease type recognition model;
[0080] The model deployment module 603 is used to deploy the fundus disease type recognition model to perform real-time recognition of the fundus disease type according to the currently input fundus disease image.
[0081] Through the cooperation of the data acquisition module 601, the model construction module 602 and the model deployment module 603, effective identification and classification of fundus disease types are achieved. Specifically, the fundus disease images screened by the image quality model are obtained and comprehensively preprocessed, including size unification, channel number standardization and label annotation, so as to ensure the standardization and consistency of the input data, and greatly improve the efficiency and stability of model training. Each fundus disease image is divided into several image blocks and converted into a one-dimensional vector, so that the features of each image block are easier to be processed by the model, and classification tags and position codes are added. The classification tags introduce global information, which helps to improve the classification accuracy; the position code retains the spatial relationship of the image blocks, ensuring that the model does not lose position information when capturing local and global features. The layer normalization operation can make the features between different levels consistent in scale. The discard operation randomly sets the elements in the embedded vector to zero, which improves the robustness of the model. The multi-head attention mechanism captures local and global information, supports efficient parallel computing, and can improve the performance and generalization ability of the model. Residual connections help alleviate the problem of vanishing gradients in deep neural networks, ensuring that gradients can be effectively propagated so that they can capture more complex features. Feedforward fully connected networks can better capture complex feature relationships in the input by increasing the dimension, and can help compress redundant information and simplify feature representation by reducing the dimension. Global pooling aggregates the features of all image blocks by the mean, effectively integrating local features, avoiding over-reliance on features in a specific area, and ensuring that the model maintains a global perception of the entire image when processing.
[0082] It should be understood that the specific process of each module executing the above corresponding steps has been described in detail in the above method embodiment, and for the sake of brevity, it will not be repeated here.
[0083] It should also be understood that the division of modules in the embodiments of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, each functional module in each embodiment of the present application may be integrated into a processor, or may exist physically separately, or two or more modules may be integrated into one module. The above-mentioned integrated modules may be implemented in the form of hardware or in the form of software functional modules.
[0084] Figure 7 is a schematic block diagram of an electronic terminal provided in an embodiment of the present application. Figure 7 As shown, the electronic terminal 700 includes: at least one processor 701, a memory 702, at least one network interface 703 and a user interface 705. The various components in the electronic terminal 700 are coupled together through a bus system 704. It can be understood that the bus system 704 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 704 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 704 is not used in the following description. Figure 7 In the specification, various buses are labeled as bus systems.
[0085] The user interface 705 may include a display, a keyboard, a mouse, a trackball, a click gun, keys, buttons, a touch pad or a touch screen.
[0086] It is understood that the memory 702 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), which is used as an external cache. By way of exemplary but not limiting explanation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM). The memory described in the embodiments of the present invention is intended to include but is not limited to these and any other suitable categories of memory.
[0087] The memory 702 in the embodiment of the present invention is used to store various categories of data to support the operation of the electronic terminal 700. Examples of these data include: any executable program for operating on the electronic terminal 700, such as an operating system 7021 and an application 7022; the operating system 7021 includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application 7022 may include various applications, such as a media player (Media Player), a browser (Browser), etc., for implementing various application services. The method for identifying the type of fundus disease based on the visual transformer model provided in the embodiment of the present invention may be included in the application 7022.
[0088] The method disclosed in the above embodiment of the present invention can be applied to the processor 701, or implemented by the processor 701. The processor 701 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit in the processor 701 or the instruction in the form of software. The above processor 701 may be a general processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor 701 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiment of the present invention. The general processor 701 may be a microprocessor or any conventional processor, etc. In combination with the steps of the accessory optimization method provided in the embodiment of the present invention, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in a memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.
[0089] In an exemplary embodiment, the electronic terminal 700 may be implemented by one or more application specific integrated circuits (ASIC), DSP, programmable logic device (PLD), complex programmable logic device (CPLD) to execute the aforementioned method.
[0090] According to the method provided in the embodiment of the present application, the present application also provides a computer program product, which includes: a computer program code, when the computer program code is run on a computer, the computer executes Figures 1 to 5A method according to any one of the embodiments shown.
[0091] According to the method provided in the embodiment of the present application, the present application also provides a computer-readable storage medium, which stores a program code, and when the program code is run on a computer, the computer executes Figures 1 to 5 A method according to any one of the embodiments shown.
[0092] The terms "component", "module", "system", etc. used in this specification are used to represent computer-related entities, hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process, a processor, an object, an executable file, an execution thread, a program and / or a computer running on a processor. By way of illustration, both applications and computing devices running on a computing device can be components. One or more components may reside in a process and / or an execution thread, and a component may be located on a computer and / or distributed between two or more computers. In addition, these components may be executed from various computer-readable media having various data structures stored thereon. Components may, for example, communicate through local and / or remote processes according to signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system and / or a network, such as the Internet interacting with other systems through signals).
[0093] Those of ordinary skill in the art will appreciate that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0094] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0095] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0096] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0097] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0098] In the above embodiments, the functions of each functional unit can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated. Available media may be magnetic media (e.g., floppy disks, hard disks, tapes), optical media (e.g., high-density digital video discs (DVDs), or semiconductor media (e.g., solid state disks (SSDs)).
[0099] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.
[0100] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0101] In summary, the present application provides a method, system, medium, terminal and program product for identifying fundus disease types based on a visual transformer model, by obtaining fundus disease images screened by an image quality screening model, and preprocessing the fundus disease images, labeling each fundus disease image with the corresponding fundus disease, and forming a fundus disease image data set; inputting the fundus disease image data set into the visual transformer model for training, thereby constructing a fundus disease type recognition model; deploying the trained fundus disease recognition model in specific application scenarios, such as communities or primary care clinics, etc., and the deployed fundus disease type recognition model recognizes the fundus disease corresponding to the image according to the currently input fundus disease image. The present application uses fundus images without mydriasis to reduce the impact on the lives of the initial screeners, and the fundus disease type recognition model obtained by training the visual transformer model can effectively identify a variety of fundus disease types. Therefore, the present application effectively overcomes the various shortcomings in the prior art and has a high industrial utilization value.
[0102] The above embodiments are merely illustrative of the principles and effects of the present application and are not intended to limit the present application. Anyone familiar with the technology may modify or change the above embodiments without violating the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by a person of ordinary skill in the art without departing from the spirit and technical ideas disclosed in the present application shall still be covered by the claims of the present application.
Claims
1. A method for identifying fundus disease types based on a visual transformer model, characterized in that: include: Acquire fundus disease images screened by the image quality screening model, and preprocess the fundus disease images to form a fundus disease image data set; Each fundus disease image in the fundus disease image dataset has been labeled with a corresponding fundus disease; The fundus disease labels annotated for the fundus disease image include: age-related macular degeneration, diabetic retinopathy, myopic maculopathy, suspected glaucoma, cataract, epiretinal membrane, macular hole, retinal vein occlusion, retinal artery occlusion, retinitis pigmentosa, retinal detachment, vitreous degeneration, myelinated nerve fibers, any one or more combinations thereof; Inputting the fundus disease image data set into a visual transformer model for training to construct a fundus disease type recognition model; Deploy the fundus disease type recognition model to identify the fundus disease in real time according to the currently input fundus disease image; The training process of the image quality screening model includes: Acquire a non-mydriatic fundus image, and annotate the fundus image with a quality label to form a fundus image quality data set; the quality label includes any one or more combinations of a field of view centered on the macula and the optic nerve head, macula visibility, optic nerve head visibility, eyelid artifacts, camera artifacts, media opacity, blood vessel visibility, and illumination; Inputting the fundus image quality data set into an image classification model for training to obtain a trained image quality screening model; the image quality screening model is used to perform quality screening on the currently input fundus image; The image quality screening model is connected to the fundus disease type recognition model to input the screened fundus disease images into the fundus disease type recognition model.
2. The method for identifying fundus disease types based on a visual transformer model according to claim 1, characterized in that: The training process of the fundus disease type recognition model includes: The embedding layer based on the visual transformer model transforms each fundus disease image in the fundus disease image data set, and adds a classification mark sequence and a position code to obtain an embedding sequence of fundus disease image blocks; Inputting the fundus disease image block embedding sequence into the encoder of the visual transformer model to extract features of the fundus disease image block to obtain a feature sequence; After removing the features of the classification mark position from the feature sequence, global pooling is performed to generate a global feature vector, and the global feature vector is input into the multi-layer perceptron of the visual transformer model to output the fundus disease type recognition result.
3. The method for identifying fundus disease types based on a visual transformer model according to claim 2, characterized in that: The embedding layer based on the visual transformer model transforms each fundus disease image in the fundus disease image data set, and adds a classification mark sequence and a position code to obtain an embedding sequence of fundus disease image blocks, which includes: Dividing the fundus disease image into a plurality of image blocks based on a convolutional neural network, and performing a flattening operation on each image block to form an image block embedding sequence; Inserting a classification tag sequence before the image block embedding sequence to obtain an image block embedding sequence containing classification information and image features; A positional encoding is added to each embedding vector in the image block embedding sequence containing classification information and image features, and a discard operation is performed to obtain an embedding sequence of fundus disease image blocks.
4. The method for identifying fundus disease types based on a visual transformer model according to claim 2, characterized in that: The embedding sequence of the fundus disease image blocks is input into an encoder of a visual transformer model, and is used to extract features from the fundus disease image blocks to obtain a feature sequence, which includes: Performing a layer normalization operation on the fundus disease image block embedding sequence to obtain a standardized embedding sequence; Inputting the standardized embedding sequence into a multi-head attention mechanism to obtain an attention feature sequence and perform a discard operation; Adding the attention feature sequence and the fundus disease image block embedding sequence through a residual connection to obtain a first residual feature sequence; Performing a layer normalization operation on the first residual feature sequence to obtain a standardized first residual feature sequence; Inputting the standardized first residual feature sequence into a feedforward fully connected network to obtain a feedforward feature sequence; The feedforward feature sequence is discarded and then added to the first residual feature sequence through a residual connection to obtain a second residual feature sequence; Repeat the above layer normalization, multi-head attention mechanism, residual connection, layer normalization, feedforward fully connected network, and residual connection operations multiple times to gradually extract features and obtain a feature sequence.
5. A fundus disease type recognition system based on a visual transformer model, characterized in that: include: A data acquisition module, used for acquiring fundus disease images screened by the image quality screening model, and preprocessing the fundus disease images to form a fundus disease image data set; each fundus disease image in the fundus disease image data set has been labeled with a corresponding fundus disease; The fundus disease labels annotated for the fundus disease image include: age-related macular degeneration, diabetic retinopathy, myopic maculopathy, suspected glaucoma, cataract, epiretinal membrane, macular hole, retinal vein occlusion, retinal artery occlusion, retinitis pigmentosa, retinal detachment, vitreous degeneration, myelinated nerve fibers, any one or more combinations thereof; A model building module, used for inputting the fundus disease image data set into a visual transformer model for training to build a fundus disease type recognition model; A model deployment module, used to deploy the fundus disease type recognition model to perform real-time recognition of the fundus disease type according to the currently input fundus disease image; The training process of the image quality screening model includes: Acquire a non-mydriatic fundus image, and annotate the fundus image with a quality label to form a fundus image quality data set; the quality label includes any one or more combinations of a field of view centered on the macula and the optic nerve head, macula visibility, optic nerve head visibility, eyelid artifacts, camera artifacts, media opacity, blood vessel visibility, and illumination; Inputting the fundus image quality data set into an image classification model for training to obtain a trained image quality screening model; the image quality screening model is used to perform quality screening on the currently input fundus image; The image quality screening model is connected to the fundus disease type recognition model to input the screened fundus disease images into the fundus disease type recognition model.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.
7. A computer program product, characterized in that The computer program product includes computer program codes, and when the computer program codes are executed on a computer, the computer is enabled to implement the method according to any one of claims 1 to 4.
8. An electronic terminal comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Fundus illumination multi-disease detection system based on regional feature set neural network
CN111046835A
Multi-model sugar network lesion automatic screening method based on convolutional neural network
CN112200794A
Ophthalmology imaging method and system for automatic retina feature detection
CN115100729A