Establishment system and method of fundus anomaly recognition model and medium
By combining the architecture of UNet and discriminant model, using the AutoEncoder model for pre-training, and using Berlin noise to generate abnormal images, an efficient fundus anomaly recognition model was constructed, solving the problems of high false alarm rates and missed alarm rates in the existing technology, and achieving more accurate fundus color image abnormal recognition.
Patent Information
- Application Number
- CN202411939966.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art has high false alarm rates or missed alarm rates in fundus color abnormality recognition, and the AutoEncoder model lacks the integration of knowledge in specific fields, making it difficult to improve the model effect.
Using a hybrid UNet and discriminant model architecture, pre-training is performed through the AutoEncoder model, fused with the Unet model, and using the Berlin noise generation mask matrix to create anomaly images, and data training is performed to build a fundus anomaly recognition model.
It effectively reduces the work burden of doctors/image processing analysts, improves the accuracy and diagnostic efficiency of fundus color abnormalities, and can identify general fundus color abnormalities without additional abnormal labeling data.
Smart Images

Figure CN120047989A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of abnormal fundus color photograph image recognition, and in particular to a system and method for constructing an abnormal fundus recognition model and a medium. Background Art
[0002] With the maturity of deep learning technology and the increasingly complete data sets, various CV (computer image) models based on deep learning have become possible, which has further promoted the implementation of computer vision in various fields, from autonomous driving, intelligent monitoring to industrial automation, greatly improving production efficiency and convenience of life. In the medical field, the application of computer vision technology is particularly significant. High-precision models are used in medical image analysis, pathological detection and surgical assistance, which significantly improves the accuracy and speed of diagnosis.
[0003] Currently, in the task of abnormality recognition in fundus color photographs, accurately identifying and locating abnormal areas is crucial for early diagnosis and treatment of eye diseases. However, manual analysis of fundus images is not only time-consuming and laborious, but also easily affected by human factors. How to apply deep learning models to the task of abnormality recognition in fundus color photographs is of great research significance.
[0004] Patent document CN118542639A discloses a fundus image diagnosis and analysis system and method based on pattern recognition. By combining pattern recognition technology and image processing technology, comprehensive and multi-scale feature extraction is performed on the fundus color photos of the diagnosed patient object to learn and capture the abnormal state of the fundus, and further realize the discrimination of the type of fundus abnormality, so as to assist doctors in screening fundus lesions and reduce the workload of doctors / image processing analysts. However, this diagnostic method requires a lot of feature extraction and manual labeling work.
[0005] The existing technology also uses the Autoencoder model to reduce manual labeling, but traditional unsupervised anomaly detection usually uses the Autoencoder model to perform anomaly recall by measuring the difference between the reconstruction result of the model and the original image with a threshold. How to choose an appropriate reconstruction error threshold in this way is a challenge. Improper threshold selection may lead to high false alarm rate (normal data is misjudged as abnormal) or false alarm rate (failure to detect anomalies). In addition, as a general model, the Autoencoder model lacks the integration of specific domain knowledge, and it is very difficult to improve the model effect in the later stage of use.
[0006] In existing abnormality recognition, it is often necessary to segment fundus images. For example, the Unet image segmentation model is used for image segmentation. However, the segmentation task in the Unet image segmentation model has high requirements on the data volume and data quality, and it is impossible to model generic abnormality types.
[0007] Based on the above problems, the applicant proposed the technical solution of this application. Summary of the invention
[0008] In view of the above-mentioned defects of the prior art, the present invention can effectively segment and identify abnormal areas in fundus images by using a deep learning model, especially a hybrid UNet and discriminant model architecture, thereby greatly reducing the workload of doctors / image processing analysts and improving diagnostic efficiency and accuracy.
[0009] In order to achieve the above object, the present invention discloses a method for constructing a fundus abnormality recognition model, comprising the following steps:
[0010] Collecting fundus color images, wherein the fundus color images include a plurality of normal fundus color images;
[0011] Constructing a training set and a validation set according to the normal fundus color photograph images, and using the training set and the validation set to train the AutoEncoder model to obtain a basic pre-trained model capable of reconstructing normal fundus images;
[0012] Constructing a Unet model, freezing the encoder in the basic pre-trained model, and fusing the Unet model with the frozen basic pre-trained model to obtain a fusion model;
[0013] Using Perlin noise to generate a mask matrix, and performing pixel replacement on a plurality of normal fundus color photographs based on the mask matrix to obtain a plurality of abnormal color photographs with abnormal areas;
[0014] The abnormal color photograph image and the normal color photograph image of the fundus are used to perform data training on the fusion model to obtain a fundus abnormality recognition model.
[0015] Preferably, in the training of the AutoEncoder model, ResNet50 is used as the encoder layer and 5 upsample layers are used as the decoder layer. After the training of the AutoEncoder model is completed, the encoder layer of ResNet50 is taken out and frozen.
[0016] Preferably, the step of generating a mask matrix using Perlin noise comprises the following steps:
[0017] Generate a two-dimensional Perlin noise matrix, each element of the Perlin noise matrix corresponds to a pixel position in the image;
[0018] Normalizing the Perlin noise matrix so that the value range of each element thereof is between 0 and 1;
[0019] Each element in the Perlin noise matrix is binarized based on a set threshold to generate a mask matrix.
[0020] Preferably, the pixel replacement of a plurality of normal fundus color photographs based on the mask matrix is specifically performed as follows:
[0021] For pixel positions whose mask values in the mask matrix are the first set value, the pixel values of the normal image are retained, and for pixel positions whose mask values in the mask matrix are the second set value, the pixel values of the normal image are replaced with the pixel values of the preset abnormal area.
[0022] Preferably, ResNet50 is used as the backbone network for feature extraction in the fusion model, and the convolution layer of ResNet50 gradually downsamples the input image to extract multi-level feature representation.
[0023] Preferably, the fusion model uses a Unet encoder to gradually downsample the output feature map of ResNet50 to extract image features of different scales. In each downsampling step, the frozen ResNet50 feature map is spliced with the feature map of the current layer to fuse the pre-trained features and the current features.
[0024] Preferably, the fusion model uses a Unet decoder to gradually upsample, and in each upsampling step, the upsampled feature map is jump-connected with the feature map of the corresponding layer of the encoder to fuse multi-scale features.
[0025] Preferably, the loss function of the fusion model includes a segmentation task loss function, a classification task loss function and a total loss function. The segmentation task loss function uses cross entropy loss to measure the difference between the predicted result of the segmentation mask and the true label, and the classification task loss function uses cross entropy loss to measure the difference between the classification result and the true label. The total loss function is the weighted sum of the segmentation task loss and the classification task loss.
[0026] The present invention also discloses a system for constructing a fundus abnormality recognition model, comprising the following modules:
[0027] A normal image acquisition module, used for acquiring fundus color images, wherein the fundus color images include a plurality of normal fundus color images;
[0028] A pre-training module, used to construct a training set and a validation set according to the normal fundus color photograph image, and use the training set and the validation set to train the AutoEncoder model to obtain a basic pre-training model capable of reconstructing normal fundus images;
[0029] A model building module, used to build a Unet model, freeze the encoder in the basic pre-trained model, and fuse the Unet model with the frozen basic pre-trained model to obtain a fusion model;
[0030] An abnormal image generation module is used to generate a mask matrix using Perlin noise, and perform pixel replacement on a plurality of normal fundus color images based on the mask matrix to obtain a plurality of abnormal color images with abnormal areas;
[0031] The model training module is used to perform data training on the fusion model using the abnormal color photo image and the normal color photo image of the fundus to obtain a fundus abnormality recognition model.
[0032] The present invention also discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned knowledge base question-answering method based on intent recognition is implemented.
[0033] Compared with the prior art, the present invention has the following beneficial effects:
[0034] The method for constructing a fundus abnormality recognition model provided by the present invention is to obtain a basic pre-training model by collecting normal fundus color photos to train the AutoEncoder model, construct a Unet model and fuse it with the basic pre-training model to form a fusion model, and then construct an abnormal color photo with an abnormal area through Perlin noise, and train the fusion model with the normal color photo and the abnormal color photo to form a fundus abnormality recognition model. After the fundus abnormality recognition model obtained in this way is used for fundus abnormality recognition, it can more accurately perform abnormality recognition of fundus color photos. Since the noise-based abnormal data enhancement is used to realize the general type of fundus color photo abnormality recognition model, it is no longer necessary to use any abnormal annotation data in the model training, and only the generated abnormal data is used to realize the abnormality recognition of fundus color photos, so that the abnormal categories that were originally difficult to collect in large quantities can be recognized.
[0035] The concept, specific structure and technical effects of the present invention will be further described below in conjunction with the accompanying drawings to fully understand the purpose, characteristics and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a flow chart of the method for constructing the fundus abnormality recognition model of the present invention. DETAILED DESCRIPTION
[0037] In order to make the technical means, creative features, objectives and effects of the invention easier to understand, the invention is further described below with reference to specific diagrams. However, the invention is not limited to the following implementation cases.
[0038] It should be noted that the structures, proportions, sizes, etc. illustrated in the drawings in this specification are only used to match the contents disclosed in the specification so as to facilitate understanding and reading by persons familiar with this technology. They are not used to limit the conditions under which the present invention can be implemented, and therefore have no substantive technical significance. Any structural modification, change in proportion or adjustment of size, without affecting the effects and purposes that can be achieved by the present invention, should still fall within the scope of the technical contents disclosed by the present invention.
[0039] like Figure 1 As shown, the first embodiment of the present invention discloses a method for constructing a fundus abnormality recognition model, comprising the following steps:
[0040] Step S1, collecting fundus color photograph images, wherein the fundus color photograph images include a plurality of normal fundus color photograph images.
[0041] In this embodiment, fundus color photographs are used, including 1500 normal fundus color photographs in all directions and 79 abnormal images collected as abnormal images as a test set. The following Table 1 lists in detail the data used in this embodiment.
[0042] Data Dimensions Training set Validation set Test Set Total Normal color photo 1000 300 200 1500 Abnormal color photos - - 79 79
[0043] Table 1. Data volume and data segmentation logic of this embodiment
[0044] Fundus Photography is a medical imaging technology used to record and analyze fundus structures. The fundus refers to the posterior area inside the eyeball, including structures such as the retina, optic disc (optic nerve head), macula, and blood vessels. Fundus photography is taken with special photographic equipment and can provide high-resolution color images for the diagnosis and monitoring of various ophthalmic diseases, such as diabetic retinopathy, macular degeneration, glaucoma, etc.
[0045] Since this embodiment aims to train a classification and segmentation multi-task model by using normal color photos and generated noise, in order to ensure that data is not leaked, only 1000 normal color photos are used as the training set in the training stage. Then, an AutoEncoder model is pre-trained using the training set to reconstruct normal images, as well as a multi-task anomaly recognition model. The annotation data of this multi-task anomaly recognition model only comes from the corresponding image and mask generated by the anomaly type generator, so there is no need to perform any additional anomaly annotation data.
[0046] Step S2, constructing a training set and a validation set according to the normal fundus color photograph image, and using the training set and the validation set to train the AutoEncoder model to obtain a basic pre-trained model capable of reconstructing normal fundus images.
[0047] Specifically, AutoEncoder is a neural network model that is mainly used for unsupervised learning tasks. It encodes the input data into a low-dimensional latent space representation, and then reconstructs the original data from this representation. The core idea of AutoEncoder is to minimize the difference between the input data and the reconstructed data by training the network. In anomaly detection, the unsupervised nature of AutoEncoder is particularly important. It does not require labeled data, but only needs to be trained with normal data to learn the characteristics of normal data.
[0048] The AutoEncoder model usually consists of two main parts, the encoder and the decoder. The encoder compresses the input data into a low-dimensional latent representation, while the decoder reconstructs the original data from this latent representation. In this way, the AutoEncoder can effectively capture the main features of the input data. During the training process, the model will learn how to reconstruct normal data, but for abnormal data, the reconstruction error will increase significantly because its characteristics are different from normal data. Therefore, by detecting the reconstruction error, abnormal data can be effectively identified.
[0049] In image processing tasks, AutoEncoder can learn the features of normal images, which is very helpful for subsequent image segmentation (semantic segmentation) tasks. Specifically, the encoder layer can extract high-level features of the image, which can be used as input for subsequent segmentation models to enhance the segmentation effect. In this way, AutoEncoder can not only be used for anomaly detection, but also provide strong support for other computer vision tasks.
[0050] Step S3, constructing a Unet model, freezing the encoder in the basic pre-trained model, and fusing the Unet model with the frozen basic pre-trained model to obtain a fusion model.
[0051] Specifically, in this embodiment, when training the AutoEncoder model, ResNet50 is used as the encoder layer and 5 upsample layers are used as the decoder layer. After the AutoEncoder model training is completed, the encoder layer of ResNet50 is taken out for freezing. The fusion model obtained through such training has a feature extractor capable of modeling normal images.
[0052] Step S4, using Perlin noise to generate a mask matrix, and performing pixel replacement on a plurality of normal fundus color photographs based on the mask matrix to obtain a plurality of abnormal color photographs with abnormal areas.
[0053] In ophthalmology, abnormal images are usually divided into two categories: one is global abnormality and the other is local abnormality. Global abnormality may be caused by changes in the image color gamut, such as fundus images appearing red due to bleeding, or lesions appearing green. Or it may be blurring due to vitreous opacity or camera shake. Since the entire image is abnormal, data augmentation will generate the corresponding image and a full-area mask. Local abnormalities may be various sized dot-shaped or mass-shaped patches.
[0054] In this embodiment, Perlin noise is used for data enhancement. In the process of generating abnormal samples, a technology based on Perlin noise as a mask is adopted to add abnormal areas on normal images.
[0055] Specifically, the method of generating a mask matrix using Perlin noise includes the following steps:
[0056] Generate a two-dimensional Perlin noise matrix, each element of which corresponds to a pixel position in the image; normalize the Perlin noise matrix so that the value range of each element is between 0 and 1; and binarize each element of the Perlin noise matrix based on a set threshold to generate a mask matrix.
[0057] Generating the Perlin noise matrix specifically involves generating a two-dimensional Perlin noise matrix. Perlin noise is a smooth pseudo-random noise function that can generate natural and continuous textures, and each element of the noise matrix corresponds to a pixel position of the image.
[0058] The specific method of generating the mask matrix is to normalize the generated Perlin noise matrix so that its value range is between 0 and 1. By setting a threshold, the noise matrix is binarized to generate the mask matrix. Specifically, if the noise value is greater than or equal to the threshold, the mask value is 1; otherwise, the mask value is 0. The selection of this threshold can be adjusted according to the density and distribution of the required abnormal area, and this is not specifically limited in this embodiment.
[0059] Based on the mask matrix, the pixels of the normal fundus color photographs are replaced, specifically: for the pixel positions where the mask value in the mask matrix is the first set value, the pixel values of the normal image are retained, and for the pixel positions where the mask value in the mask matrix is the second set value, the pixel values of the normal image are replaced with the pixel values of the preset abnormal area. In this embodiment, the first set value is selected as 0, and the second set value is selected as 1, that is, for the pixel positions where the mask value is 0, the pixel values of the normal image are retained; for the pixel positions where the mask value is 1, the pixel values of the normal image are replaced with the pixel values of the predefined abnormal area, and the predefined abnormal area can be a specific color, texture or other features.
[0060] Through the above method, an abnormal area with a natural transition can be generated on a normal image, and the distribution of the abnormal area on the mask can be controlled by adjusting the Frequency, Scale, and Threshold, thereby effectively simulating the abnormal conditions that may appear in real color photos. This abnormal sample generation method based on Perlin noise is highly controllable and flexible, and is suitable for generating abnormal area half blocks of various sizes and shapes.
[0061] Step S5, using the abnormal color photo image and the normal fundus color photo image to perform data training on the fusion model to obtain a fundus abnormality recognition model.
[0062] The fusion model combines image segmentation and anomaly classification tasks, uses ResNet50 as the backbone network for feature extraction, and the convolution layer of ResNet50 gradually downsamples the input image to extract multi-level feature representations; uses the Unet encoder to gradually downsample the output feature map of ResNet50 to extract image features of different scales. In each downsampling step, the frozen ResNet50 feature map is spliced with the feature map of the current layer to fuse the pre-trained features and the current features; uses the Unet decoder to gradually upsample, and in each upsampling step, the upsampled feature map is jump-connected with the feature map of the corresponding layer of the encoder to fuse multi-scale features.
[0063] Specifically, the model architecture of the fusion model includes the input layer, feature extraction layer, UNet encoder, classification task branch, UNet decoder and output layer. ResNet50 is used as the encoder layer, 5 layers of upsampling are used as the decoder layer, and the autoencoder model is used as an unsupervised learning autoencoder model. All normal data are used as training data to perform image reconstruction tasks and train this autoencoder model. Through such training, the model can acquire the ability to model normal images. After the training, the encoder layer of Resnet50 is taken out and frozen as a feature supplement for a Unet.
[0064] The input layer input is a three-channel color image of size 3*512*512. The feature extraction layer is a ResNet50 backbone network, using a pre-trained ResNet50 model as a feature extractor. The convolution layer of ResNet50 gradually downsamples the input image to extract multi-level feature representations. The output feature map of ResNet50 will be used as the input of the UNet encoder. In the UNet encoder, the encoder feature map of the frozen autoencoder is used as part of the input of the UNet encoder, and the input image is gradually downsampled to extract features of different scales. In each downsampling step, the spatial resolution of the feature map decreases, but the number of channels increases. Moreover, in each downsampling step, the frozen ResNet50 feature map is concatenated with the feature map of the current layer to fuse the pre-trained features and the current features. In the classification task branch, a global average pooling layer is added to the output of the last convolutional layer of the encoder's ResNet50. The output of the global average pooling layer passes through a fully connected layer, and then passes through a Softmax activation function to output the abnormal classification result. In the UNet decoder, the decoder part gradually upsamples the feature map to restore the spatial resolution of the original input image. In each upsampling step, the decoder skips the upsampled feature map with the feature map of the corresponding layer of the encoder to fuse multi-scale features. The final decoder output passes through a convolutional layer to generate a segmentation mask of the same size as the input image. In the output layer, the output of the segmentation task is a segmentation mask of size 512*512, where each pixel corresponds to a category, and the output of the classification task is a category label, which represents the overall classification result of the input image.
[0065] For the loss function, the loss function of the fusion model includes a segmentation task loss function, a classification task loss function and a total loss function. The segmentation task loss function uses cross entropy loss to measure the difference between the prediction result of the segmentation mask and the true label. In an example, the segmentation task loss function uses cross entropy loss plus Dice coefficient (DiceLoss) to measure the difference between the prediction result of the segmentation mask and the true label. The classification task loss function uses cross entropy loss to measure the difference between the classification result and the true label. The total loss function is the weighted sum of the segmentation task loss and the classification task loss.
[0066] During the training process, the model optimizes the loss functions of the segmentation task and the classification task at the same time. Through multi-task learning, the model can share the parameters of the feature extraction part, thereby improving the generalization ability of the feature representation. Ultimately, the model can accurately classify images while maintaining high segmentation accuracy. This multi-task model uses the segmentation task to improve the classification ability and completes the image segmentation and anomaly classification tasks at the same time.
[0067] In the inference stage, the joint analysis of the anomaly classification results and the segmentation results can more accurately determine whether the image has anomalies.
[0068] Specifically, this embodiment adopts the following strategy to recall abnormal samples:
[0069] First, the model infers the input three-channel color image of size 3*512*512 and generates classification results and segmentation masks. The classification result represents the overall category of the image, while the segmentation mask identifies the category of each pixel in the image.
[0070] In the anomaly detection process, the classification results and segmentation results are combined for comprehensive analysis. If the classification result shows that the image belongs to the abnormal category, the image is directly regarded as an abnormal sample. In addition, the segmentation mask is analyzed to calculate the proportion of positive areas (i.e., areas segmented as abnormal categories) in the entire image. If the proportion of positive areas is greater than 10%, the image is regarded as an abnormal sample for recall.
[0071] Through this joint analysis strategy, abnormal samples can be detected more accurately and the possibility of misjudgment can be reduced. Samples marked as abnormal will be further processed and analyzed so that appropriate measures can be taken. This joint analysis method not only improves the accuracy of anomaly detection, but also makes better use of the complementary information of classification and segmentation tasks, providing a reliable basis for subsequent anomaly processing.
[0072] It should be noted that the normal fundus color photo data set collected is trained through the autoencoder model to obtain a basic pre-trained model, and then a segmentation model is merged. The pre-trained model is used as part of the parameters and the pre-trained model is frozen to train the basic unet model. At the same time, the abnormal color photo is generated by superimposing random noise through the normal fundus reference. If a normal picture is enhanced by the above noise, it is an abnormal picture, otherwise it is a normal picture. Then the segmentation+classification model is trained through the constructed segmentation model and the generated abnormal data to complete the model construction.
[0073] After the fundus abnormality recognition model obtained by the method in the first embodiment is used for fundus abnormality recognition, it can more accurately recognize the abnormalities of fundus color photos. Since the noise-based abnormal data enhancement is used to realize the general type of fundus color photo abnormality recognition model, no abnormal annotation data is needed in model training. Only the generated abnormal data is used to realize the abnormal recognition of fundus color photos, so that the abnormal categories that were originally difficult to collect in large quantities can be recognized, and more general type abnormality detection is supported. Since this method does not require a large amount of labeled data, it is convenient to produce the first version of the model. You can choose to use relatively less data to train the first version of the model, which can better solve the cold start of the abnormality recognition model.
[0074] The second embodiment of the present invention discloses a system for constructing a fundus abnormality recognition model, including a normal image acquisition module, a pre-training module, a model construction module, an abnormal image generation module and a model training module.
[0075] The normal image acquisition module is used to acquire fundus color photograph images, wherein the fundus color photograph images include a plurality of fundus normal color photograph images.
[0076] The pre-training module is used to construct a training set and a validation set according to the normal fundus color photograph image, and use the training set and the validation set to train the AutoEncoder model to obtain a basic pre-training model capable of reconstructing normal fundus images.
[0077] The model building module is used to build a Unet model, freeze the encoder in the basic pre-trained model, and fuse the Unet model with the frozen basic pre-trained model to obtain a fusion model.
[0078] The abnormal image generation module is used to generate a mask matrix using Perlin noise, and perform pixel replacement on a plurality of normal fundus color images based on the mask matrix to obtain a plurality of abnormal color images with abnormal areas.
[0079] The model training module is used to perform data training on the fusion model using the abnormal color photo image and the normal fundus color photo image to obtain a fundus abnormality recognition model.
[0080] Since the first embodiment corresponds to this embodiment, this embodiment can be implemented in conjunction with the first embodiment. The relevant technical details mentioned in the first embodiment are still valid in this embodiment, and the technical effects that can be achieved in the first embodiment can also be achieved in this embodiment. In order to reduce repetition, they are not repeated here. Accordingly, the relevant technical details mentioned in this embodiment can also be applied in the first embodiment.
[0081] The third embodiment of the present invention relates to a computer-readable storage medium having a computer program / instruction stored thereon, wherein the computer program / instruction implements the steps of the method in the first embodiment when executed by a processor.
[0082] The preferred specific embodiments of the present invention are described in detail above. It should be understood that ordinary technicians in the field can make many modifications and changes based on the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by technicians in the technical field based on the concept of the present invention through logical analysis, reasoning or limited experiments on the basis of the prior art should be within the scope of protection determined by the claims.
Claims
1. A method for constructing a fundus abnormality recognition model, characterized in that: The following steps are involved: Collecting fundus color images, wherein the fundus color images include a plurality of normal fundus color images; Constructing a training set and a validation set according to the normal fundus color photograph images, and using the training set and the validation set to train the AutoEncoder model to obtain a basic pre-trained model capable of reconstructing normal fundus images; Constructing a Unet model, freezing the encoder in the basic pre-trained model, and fusing the Unet model with the frozen basic pre-trained model to obtain a fusion model; Using Perlin noise to generate a mask matrix, and performing pixel replacement on a plurality of normal fundus color photographs based on the mask matrix to obtain a plurality of abnormal color photographs with abnormal areas; The abnormal color photograph image and the normal color photograph image of the fundus are used to perform data training on the fusion model to obtain a fundus abnormality recognition model.
2. The method for constructing a fundus abnormality recognition model according to claim 1, characterized in that: In the training of the AutoEncoder model, ResNet50 is used as the encoder layer and 5 upsample layers are used as the decoder layer. After the training of the AutoEncoder model is completed, the encoder layer of ResNet50 is taken out for freezing.
3. The method for constructing a fundus abnormality recognition model according to claim 1, characterized in that: The method of generating a mask matrix using Perlin noise comprises the following steps: Generate a two-dimensional Perlin noise matrix, each element of the Perlin noise matrix corresponds to a pixel position in the image; Normalizing the Perlin noise matrix so that the value range of each element thereof is between 0 and 1; Each element in the Perlin noise matrix is binarized based on a set threshold to generate a mask matrix.
4. The method for constructing a fundus abnormality recognition model according to claim 1, characterized in that: The pixel replacement of the plurality of normal color fundus photographs based on the mask matrix is specifically performed as follows: For pixel positions whose mask values in the mask matrix are the first set value, the pixel values of the normal image are retained, and for pixel positions whose mask values in the mask matrix are the second set value, the pixel values of the normal image are replaced with the pixel values of the preset abnormal area.
5. The method for constructing a fundus abnormality recognition model according to claim 1, characterized in that: In the fusion model, ResNet50 is used as the backbone network for feature extraction. The convolution layer of ResNet50 gradually downsamples the input image and extracts multi-level feature representations.
6. The method for constructing a fundus abnormality recognition model according to claim 5, characterized in that: In the fusion model, the output feature map of ResNet50 is gradually downsampled using the Unet encoder to extract image features of different scales. In each downsampling step, the frozen ResNet50 feature map is concatenated with the feature map of the current layer to fuse the pre-trained features with the current features.
7. The method for constructing a fundus abnormality recognition model according to claim 5, characterized in that: The fusion model uses a Unet decoder to gradually upsample, and in each upsampling step, the upsampled feature map is jump-connected with the feature map of the corresponding layer of the encoder to fuse multi-scale features.
8. The method for constructing a fundus abnormality recognition model according to claim 1, characterized in that: The loss function of the fusion model includes a segmentation task loss function, a classification task loss function and a total loss function. The segmentation task loss function uses cross entropy loss to measure the difference between the predicted result of the segmentation mask and the true label. The classification task loss function uses cross entropy loss to measure the difference between the classification result and the true label. The total loss function is a weighted sum of the segmentation task loss and the classification task loss.
9. A system for constructing a fundus abnormality recognition model, characterized in that: Includes the following modules: A normal image acquisition module, used for acquiring fundus color images, wherein the fundus color images include a plurality of normal fundus color images; A pre-training module, used to construct a training set and a validation set according to the normal fundus color photograph image, and use the training set and the validation set to train the AutoEncoder model to obtain a basic pre-training model capable of reconstructing normal fundus images; A model building module, used to build a Unet model, freeze the encoder in the basic pre-trained model, and fuse the Unet model with the frozen basic pre-trained model to obtain a fusion model; An abnormal image generation module is used to generate a mask matrix using Perlin noise, and perform pixel replacement on a plurality of normal fundus color images based on the mask matrix to obtain a plurality of abnormal color images with abnormal areas; The model training module is used to perform data training on the fusion model using the abnormal color photo image and the normal color photo image of the fundus to obtain a fundus abnormality recognition model.
10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, which, when executed by a processor, implements the knowledge base question-answering method based on intent recognition as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Eye fundus image diagnostic analysis system and method based on pattern recognition
CN118542639A
Cited By
Medical image abnormal region identification method based on large model self-supervised learning
CN120655643A
A medical image abnormal region identification method based on large model self-supervised learning
CN120655643B