Glaucoma different stage discrimination system based on multi-modal information network

By combining a multimodal information network with fundus images and medical records, a glaucoma discrimination system has been developed, which solves the problems of long diagnostic time and low accuracy of existing diagnostic methods and achieves efficient and accurate identification of different stages of glaucoma.

CN119541835BActive Publication Date: 2025-12-26HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411694302.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-25
Publication Date
2025-12-26
Estimated Expiration
2044-11-25

AI Technical Summary

Technical Problem

Existing glaucoma diagnostic methods are time-consuming, susceptible to human factors, and have low diagnostic efficiency. Automated diagnostic technologies require large computational resources and are difficult to effectively identify different stages of glaucoma.

Method used

A glaucoma discrimination system based on multimodal information networks is adopted, including data acquisition, preprocessing, super-resolution reconstruction and image-text multimodal convolutional network model. Combining fundus optical coherence tomography (OCT), OCTA and RNFL images and medical record information, different stages of glaucoma are predicted through super-resolution reconstruction and multimodal feature fusion.

Benefits of technology

It improves the accuracy and efficiency of glaucoma diagnosis, reduces the amount of image data required, and improves the accuracy of the glaucoma diagnosis process by combining image and text information. It also improves the probability of early, middle and late stages of glaucoma by combining image and text information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119541835B_ABST
    Figure CN119541835B_ABST
Patent Text Reader

Abstract

The present application relates to a glaucoma different stage discrimination system based on a multi-modal information network, belongs to computer vision and machine learning technology, and aims to solve the problem of low discrimination accuracy of the existing glaucoma different stage discrimination system. The system comprises a data acquisition module, a data preprocessing module, a super-resolution reconstruction network model, a picture-text multi-modal convolution network model and a to-be-tested module. The data acquisition module is used to acquire OCT images, OCTA images, RNFL images and medical record information. The data preprocessing module is used to obtain preprocessed OCT image training sets, OCTA image training sets, RNFL image training sets and medical record information. The super-resolution reconstruction network model is used to obtain a trained super-resolution reconstruction network model. The picture-text multi-modal convolution network model is used to obtain a trained picture-text multi-modal convolution network model. The to-be-tested module is used to output the probability of the different stages of glaucoma of the to-be-tested data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a glaucoma stage discrimination system, belonging to the technical fields of computer vision and machine learning. BACKGROUND

[0002] Glaucoma is characterized by its progression, and if not detected and treated in time, it can cause vision loss and irreversible blindness. Early identification and prevention of glaucoma are crucial to addressing this serious eye disease.

[0003] However, traditional glaucoma diagnosis methods not only take a long time, but are also easily affected by human factors during operation, resulting in errors in diagnosis results, and the overall diagnosis efficiency is not high. These problems have limited the application effect of traditional diagnosis methods in glaucoma screening and management.

[0004] In order to improve the accuracy of diagnosis and simplify the diagnosis process, the introduction of automatic diagnosis technology becomes particularly important and necessary.

[0005] There is currently a glaucoma automatic diagnosis system based on GoogleNet. The system uses a sliding window method combined with a network, combined with manually extracted OCTA structures and region of interest (ROI) sub-images for training. After training, even with images of poor quality databases, the algorithm shows good accuracy. The pre-processing step of the database image increases the computational overhead, which may slow down the entire workflow and increase resource requirements.

[0006] The accuracy rate (AR) of DenseNet169 reached 85.19%, indicating that it is effective in distinguishing between mild and severe glaucoma. But DenseNet169 is a deep neural network with a large number of parameters, making its training computationally intensive and requiring a large amount of computing resources.

[0007] The CAD framework that assists social professionals in glaucoma screening achieved a high accuracy of 96% in glaucoma classification, showing its potential utility in detecting various eye conditions in glaucoma detection. But the CAD framework with DLN requires a large number of different data sets for effective training, which can be challenging, especially for rare diseases or specific patient groups. SUMMARY

[0008] The purpose of the present application is to solve the problem of low accuracy in discriminating different stages of glaucoma, and to propose a glaucoma stage discrimination system based on a multi-modal information network.

[0009] The glaucoma stage discrimination system based on a multi-modal information network comprises a data acquisition module, a data preprocessing module, a super-resolution reconstruction network model, a graph-text multi-modal convolution network model and a to-be-tested module.

[0010] The data acquisition module is configured to collect an optical coherence tomography (OCT) image training set, an optical coherence tomography angiography (OCTA) image training set, an RNFL image training set, and medical record information, wherein the medical record information includes gender and age information of a patient.

[0011] The data preprocessing module is configured to preprocess the OCT image training set, the OCTA image training set, the RNFL image training set, and the medical record information collected by the data acquisition module, to obtain preprocessed OCT image training set, OCTA image training set, RNFL image training set, and medical record information, wherein the medical record information includes gender and age information of a patient.

[0012] The super-resolution reconstruction network model is configured to take the preprocessed OCT image training set, OCTA image training set, and RNFL image training set as inputs of the super-resolution reconstruction network model, and take high-resolution OCT image training set, OCTA image training set, and RNFL image training set as outputs of the super-resolution reconstruction network model, to obtain a trained super-resolution reconstruction network model.

[0013] The image-text multimodal convolutional network model is configured to take the high-resolution OCT image, OCTA image, RNFL image, and age, gender, and joint features output by the trained super-resolution reconstruction network model as inputs of the image-text multimodal convolutional network model, and take a probability of glaucoma at different stages as an output of the image-text multimodal convolutional network model, to obtain a trained image-text multimodal convolutional network model.

[0014] The to-be-tested module is configured to obtain a to-be-tested OCT image, OCTA image, RNFL image, age, and gender information in a case, to preprocess the to-be-tested OCT image, OCTA image, RNFL image, age, and gender information in the case, to obtain preprocessed OCT image, OCTA image, RNFL image, gender, age, and joint indicators, to input the preprocessed OCT image, OCTA image, and RNFL image into the trained super-resolution reconstruction network model, to output high-resolution OCT image, OCTA image, and RNFL image by a generator in the trained super-resolution reconstruction network model, and to input the high-resolution OCT image, OCTA image, and RNFL image output by the generator and the preprocessed gender, age, and joint indicators into the trained image-text multimodal convolutional network model, to output a probability of glaucoma at different stages by the trained image-text multimodal convolutional network model, including a probability of glaucoma at an early stage, a middle stage, and a late stage.

[0015] Preferably, the data collection module is used to collect an optical coherence tomography (OCT) image training set, an OCT angiography (OCTA) image training set, an RNFL image training set, and medical record information; the specific process is as follows:

[0016] The OCT image training set includes labeled normal OCT images, early glaucoma OCT images, medium-term glaucoma OCT images, and late glaucoma OCT images.

[0017] The OCTA image training set includes labeled normal OCTA images, early glaucoma OCTA images, medium-term glaucoma OCTA images, and late glaucoma OCTA images.

[0018] The RNFL image training set includes labeled normal RNFL images, early glaucoma RNFL images, medium-term glaucoma RNFL images, and late glaucoma RNFL images.

[0019] The OCT is an optical coherence tomography technology.

[0020] The OCTA is an optical coherence tomography angiography.

[0021] The RNFL is a retinal nerve fiber layer.

[0022] Preferably, the data preprocessing module is used to preprocess the OCT image training set, the OCTA image training set, the RNFL image training set, and the medical record information collected by the data collection module, to obtain preprocessed OCT image training set, OCTA image training set, RNFL image training set, and medical record information; the specific process is as follows:

[0023] 1) The OCT image training set collected by the data collection module is preprocessed to obtain a preprocessed OCT image training set; the specific process is as follows:

[0024] 11) The OCT image is denoised using a median filter technique to obtain a denoised OCT image.

[0025] 12) The boundaries of the denoised OCT image are located using a sobel edge detection technique to obtain an OCT image with clear boundaries.

[0026] 13) The OCT image with clear boundaries is segmented using an adaptive threshold technique into vitreous, retinal nerve fiber layer, ganglion cell layer, inner nuclear layer, outer plexiform layer, outer nuclear layer, and retinal pigment epithelial layer.

[0027] 14) Adjust the layered OCT images to 224×224 pixels, and normalize the OCT images adjusted to 224×224 pixels to obtain the preprocessed OCT image training set.

[0028] 2) Preprocess the OCTA (Optical Coherence Tomography) image training set acquired by the data acquisition module to obtain the preprocessed OCTA image training set; the specific process is as follows:

[0029] 21) Use the diffusion model to denoise the OCTA image to obtain the denoised OCTA image;

[0030] 22) Calculate the Hessian matrix of each pixel in the denoised OCTA image;

[0031] 23) Calculate the eigenvalues ​​of the Hessian matrix for each pixel and identify potential vascular structures in OCTA images based on the eigenvalues;

[0032] 24) Design a matched filter to enhance the identified potential vascular structures and suppress background noise to obtain OCTA images;

[0033] 25) The OCTA image obtained in 24) is processed using adaptive thresholding technology to obtain the filtered OCTA image;

[0034] 26) Adjust the filtered OCTA images obtained in 25) to 224×224 pixels, and normalize the adjusted OCTA images to obtain the preprocessed OCTA image training set.

[0035] 3) Preprocess the RNFL image training set acquired by the data acquisition module to obtain the preprocessed RNFL image training set; the specific process is as follows:

[0036] The RNFL image is adjusted to 224×224 pixels, and the adjusted RNFL image is normalized to obtain the preprocessed RNFL image training set.

[0037] 4) Preprocess the medical record information collected by the data acquisition module to obtain preprocessed medical record information; the preprocessed medical record information includes preprocessed age, preprocessed gender, and preprocessed joint features.

[0038] Preferably, in step 4), the medical record information collected by the data acquisition module is preprocessed to obtain preprocessed medical record information; the preprocessed medical record information includes preprocessed age, processed gender, and processed joint features;

[0039] The specific process is:

[0040] 41) Preprocessing the age data to obtain preprocessed age data; the specific process is:

[0041] The age is divided into four stages: 0-17 years old, 18-39 years old, 40-79 years old, and 80 years old and above;

[0042] The value corresponding to 0-17 years old is 1; the value corresponding to 18-39 years old is 2; the value corresponding to 40-79 years old is 3; and the value corresponding to 80 years old and above is 4;

[0043] 42) Preprocessing the gender data to obtain preprocessed gender data; the specific process is:

[0044] The gender variable is converted into a dummy variable using the get_dummies function in the pandas library;

[0045] 43) Based on the preprocessed age data and the preprocessed gender data, calculate the joint feature; represented as:

[0046] Joint feature = Beta1 x age + Beta2 x gender + Beta3

[0047] Where Beta1, Beta2, and Beta3 are coefficients;

[0048] The process of obtaining the coefficients Beta1, Beta2, and Beta3 is as follows:

[0049] Based on the joint feature, train a Logistic logistic regression model to obtain the prediction probability of having glaucoma, and call the roc_auc_score function in the scikit-learn library to calculate the area AUC under the ROC curve corresponding to the Logistic logistic regression model. Thus, one cycle is completed;

[0050] In each cycle, modify the values of the coefficients Beta1, Beta2, and Beta3, and take the set of Beta1, Beta2, and Beta3 coefficients that make the AUC value closest to 1 as the optimal Beta1, Beta2, and Beta3 coefficients;

[0051] Based on the optimal Beta1, Beta2, and Beta3 coefficients, calculate the joint feature.

[0052] Preferably, the super-resolution reconstruction network model is used to take the preprocessed fundus optical coherence tomography (OCT) image training set, the fundus optical coherence tomography angiography (OCTA) image training set and the RNFL image training set as inputs of the super-resolution reconstruction network model, and take the high-resolution OCT image training set, the OCTA image training set and the RNFL image training set as outputs of the super-resolution reconstruction network model, so as to obtain the trained super-resolution reconstruction network model; the specific process is as follows:

[0053] 1) The super-resolution reconstruction network model sequentially comprises a data conversion module, a generator network and a discriminator network.

[0054] The data conversion module comprises a high-resolution conversion module and a low-resolution conversion module.

[0055] The generator network sequentially comprises an input layer, an initial convolutional layer, a residual block, an up-sampling block, a final convolutional layer, and an output layer.

[0056] The discriminator network sequentially comprises an input layer, a convolutional layer, a batch normalization layer, a LeakyReLU activation function layer, an adaptive average pooling layer, a fully connected layer, and an output layer.

[0057] 2) The preprocessed OCT image is input into the data conversion module, and the data conversion module outputs a pair of low-resolution and high-resolution images corresponding to the OCT image.

[0058] The preprocessed OCTA image is input into the data conversion module, and the data conversion module outputs a pair of low-resolution and high-resolution images corresponding to the OCTA image.

[0059] The preprocessed RNFL image is input into the data conversion module, and the data conversion module outputs a pair of low-resolution and high-resolution images corresponding to the RNFL image.

[0060] 3) The low-resolution images corresponding to the OCT image, the OCTA image and the RNFL image output by the data conversion module are input into the generator network, and the generator network outputs high-resolution images.

[0061] 4) The high-resolution image corresponding to the OCT image output by the data conversion module or the high-resolution image corresponding to the OCT image output by the generator network is input into the discriminator network, and the discriminator network judges the probability P that the output image is the high-resolution image corresponding to the OCT image output by the data conversion module.

[0062] The high-resolution image corresponding to the OCTA image output by the data conversion module or the high-resolution image corresponding to the OCTA image output by the generator network is input into the discriminator network, and the discriminator network judges the probability P that the output image is the high-resolution image corresponding to the OCTA image output by the data conversion module.

[0063] inputting the high-resolution image corresponding to the RNFL image output by the data conversion module or the high-resolution image corresponding to the RNFL image output by the generator network into the discriminator network, and the discriminator network judging a probability P that the output image is the high-resolution image corresponding to the RNFL image output by the data conversion module;

[0064] 5), repeating 1) to 5) until the probability P that the output image is the high-resolution image corresponding to the OCT image output by the data conversion module, the probability P that the output image is the high-resolution image corresponding to the OCTA image output by the data conversion module, and the probability P that the output image is the high-resolution image corresponding to the RNFL image output by the data conversion module all reach the maximum, and a trained deep learning module is obtained.

[0065] Preferably, in the 2), the preprocessed OCT image is input into the data conversion module, and the data conversion module outputs a pair of low-resolution and high-resolution images corresponding to the OCT image.

[0066] The preprocessed OCTA image is input into the data conversion module, and the data conversion module outputs a pair of low-resolution and high-resolution images corresponding to the OCTA image.

[0067] The preprocessed RNFL image is input into the data conversion module, and the data conversion module outputs a pair of low-resolution and high-resolution images corresponding to the RNFL image.

[0068] The specific process is as follows:

[0069] The data conversion module comprises a high-resolution conversion module and a low-resolution conversion module.

[0070] The preprocessed OCT image, OCTA image and RNFL image are input into the high-resolution conversion module respectively, and the high-resolution conversion module outputs high-resolution images respectively.

[0071] The preprocessed OCT image, OCTA image and RNFL image are input into the low-resolution conversion module respectively, and the low-resolution conversion module outputs low-resolution images respectively.

[0072] The preprocessed OCT image, OCTA image and RNFL image are input into the high-resolution conversion module respectively, and the high-resolution conversion module outputs high-resolution images respectively. The process is as follows:

[0073] The preprocessed OCT image, OCTA image and RNFL image are randomly cropped respectively, and the randomly cropped OCT image, OCTA image and RNFL image are converted into tensors respectively.

[0074] The pre-processed OCT image, OCTA image and RNFL image are respectively input into a low-resolution conversion module, and the low-resolution conversion module outputs a low-resolution image; the process is as follows:

[0075] The pre-processed OCT image, OCTA image and RNFL image are respectively converted into a PIL image, and the size of the PIL image is adjusted to be equal to that of the A image by using bicubic interpolation, and the PIL image after size adjustment is converted into a tensor.

[0076] The size of the A image is the size of the image after random cropping in the high-resolution conversion module divided by the up-sampling factor.

[0077] Preferably, the low-resolution images corresponding to the OCT image, OCTA image and RNFL image output by the data conversion module in 3) are respectively input into a generator network, and the generator network outputs a high-resolution image.

[0078] The specific process is as follows:

[0079] The generator network sequentially includes an input layer, an initial convolution layer, a residual block, an up-sampling block, a final convolution layer, and an output layer.

[0080] The initial convolution layer sequentially includes a convolution layer and a pre-activation function layer.

[0081] The residual block sequentially includes a convolution layer, a batch normalization layer, a pre-activation function layer, a convolution layer, and a batch normalization layer.

[0082] The up-sampling block sequentially includes a convolution layer, a PixelShuffle up-sampling, and a pre-activation function layer.

[0083] The final convolution layer sequentially includes a convolution layer and a tanh activation function layer.

[0084] The low-resolution images corresponding to the OCT image, OCTA image and RNFL image output by the data conversion module are respectively input into the initial convolution layer, the residual block, the up-sampling block, the final convolution layer, and the output layer in sequence through the input layer, and the output layer outputs a high-resolution image.

[0085] Preferably, the high-resolution image corresponding to the OCT image output by the data conversion module or the high-resolution image corresponding to the OCT image output by the generator network in 4) is input into a discriminator network, and the discriminator network judges the probability P that the output image is the high-resolution image corresponding to the OCT image output by the data conversion module.

[0086] The high-resolution image corresponding to the OCTA image output by the data conversion module or the high-resolution image corresponding to the OCTA image output by the generator network is input into the discriminator network, and the discriminator network judges the probability P that the output image is the high-resolution image corresponding to the OCTA image output by the data conversion module.

[0087] The high-resolution image corresponding to the RNFL image output by the data conversion module or the high-resolution image corresponding to the RNFL image output by the generator network is input into the discriminator network, and the discriminator network judges the probability P that the output image is the high-resolution image corresponding to the RNFL image output by the data conversion module.

[0088] The specific process is as follows:

[0089] The discriminator network sequentially includes an input layer, a convolution layer, a batch normalization layer, a LeakyReLU activation function layer, a self-adaptive average pooling layer, a full connection layer, and an output layer.

[0090] The high-resolution image corresponding to the OCT image output by the data conversion module or the high-resolution image corresponding to the OCT image output by the generator network is sequentially input into the convolution layer, the batch normalization layer, the LeakyReLU activation function layer, the self-adaptive average pooling layer, the full connection layer, and the output layer through the input layer, and the output layer judges the probability P that the output image is the high-resolution image corresponding to the OCT image output by the data conversion module.

[0091] The high-resolution image corresponding to the OCTA image output by the data conversion module or the high-resolution image corresponding to the OCTA image output by the generator network is sequentially input into the convolution layer, the batch normalization layer, the LeakyReLU activation function layer, the self-adaptive average pooling layer, the full connection layer, and the output layer through the input layer, and the output layer judges the probability P that the output image is the high-resolution image corresponding to the OCTA image output by the data conversion module.

[0092] The high-resolution image corresponding to the RNFL image output by the data conversion module or the high-resolution image corresponding to the RNFL image output by the generator network is sequentially input into the convolution layer, the batch normalization layer, the LeakyReLU activation function layer, the self-adaptive average pooling layer, the full connection layer, and the output layer through the input layer, and the output layer judges the probability P that the output image is the high-resolution image corresponding to the RNFL image output by the data conversion module.

[0093] Preferably, 5) is repeatedly performed 1) to 5) until the probabilities P that the output image is the high-resolution image corresponding to the OCT image output by the data conversion module, the probabilities P that the output image is the high-resolution image corresponding to the OCTA image output by the data conversion module, and the probabilities P that the output image is the high-resolution image corresponding to the RNFL image output by the data conversion module are all maximized, and a trained deep learning module is obtained.

[0094] The specific process is as follows:

[0095] The total discriminator loss Loss_D is set as 1-Loss_real+Loss_fake

[0096] The loss Loss_real is -log(D(x))

[0097] D(x) is the average probability that the high-resolution image output by the discriminator network is the high-resolution image converted by the data conversion module;

[0098] The loss Loss_fake is -log(1-D(G(z)))

[0099] G(z) is the average probability that the high-resolution image output by the discriminator network is the high-resolution image generated by the generator network;

[0100] The discriminator parameter weight is updated using the Adam optimizer, so that the high-resolution image output by the discriminator network is closer to the high-resolution image converted by the data conversion module;

[0101] The total generator loss Loss_G is set as:

[0102] Loss_G=Adversarial_Loss+λ p Perception_Loss+λ i Image_Loss+λ tv TV_Loss

[0103] Adversarial_Loss is the adversarial loss;

[0104] Perception_Loss is the perception loss;

[0105] Image_Loss is the image loss;

[0106] TV_Loss is the total variation loss;

[0107] wherein λ p , λ i , and λ tv are weights.

[0108] updating the generator parameter weights using an Adam optimizer so that the high-resolution images generated by the generator are closer to the high-resolution images converted by the data conversion module;

[0109] repeating 1) to 5), alternately updating the generator and the discriminator, until the total discriminator loss and the total generator loss converge, to obtain the trained super-resolution reconstruction network model.

[0110] Preferably, the picture-text multi-modal convolutional network model is used to input the high-resolution OCT image, OCTA image, RNFL image and age, gender, combined features output by the trained super-resolution reconstruction network model as the input of the picture-text multi-modal convolutional network model, and to obtain the trained picture-text multi-modal convolutional network model by taking the disease probability of glaucoma at different stages as the output of the picture-text multi-modal convolutional network model; the specific process is as follows:

[0111] The picture-text multi-modal convolutional network model comprises a feature extraction network and a feature fusion network;

[0112] The feature extraction network comprises an image feature extraction network and a text information feature extraction network;

[0113] The image feature extraction network comprises, in sequence, an input layer, a convolutional layer, a batch normalization layer, a ReLU activation function layer, 50 residual blocks, a global average pooling layer, and an output layer;

[0114] Each residual block comprises a convolutional layer and an identity connection;

[0115] The text information feature extraction network comprises, in sequence, an input layer, an embedding layer, a convolutional layer, a max-pooling layer, and an output layer;

[0116] The feature fusion network comprises, in sequence, an input layer, a concatenation layer, a fully connected layer, a softmax layer, and an output layer;

[0117] The high-resolution OCT image, OCTA image, and RNFL image are respectively input into the image feature extraction network, and the image feature extraction network respectively outputs an OCT image feature vector, an OCTA image feature vector, and an RNFL image feature vector;

[0118] The age, gender, and combined features are respectively input into the text information feature extraction network, and the text information feature extraction network respectively outputs a feature vector of the age, gender, and combined features;

[0119] The OCT image feature vector, OCTA image feature vector, and RNFL image feature vector and the feature vector of the age, gender, and combined features are taken as the input of the feature fusion network, and the disease probability of glaucoma at different stages is output by the feature fusion network;

[0120] define a cross-entropy loss function;

[0121] update the parameters of the feature extraction network and the feature fusion network using the Adam optimization algorithm; until the cross-entropy loss function converges, a trained image-text multi-modal convolutional network model is obtained.

[0122] The beneficial effects of the present application are:

[0123] The present application first inputs the fundus optical coherence tomography OCT image, the labeled fundus optical coherence tomography OCTA image and the retinal nerve fiber layer RNFL image into the super-resolution reconstruction module to improve the resolution of the image, which is convenient for subsequent diagnosis and prediction. Then when the image data is insufficient, the text information in the medical record is also added to the diagnosis and prediction of glaucoma, which helps to reduce the amount of image data required. The text information in the medical record is jointly analyzed to find the joint index that has the greatest impact on glaucoma diagnosis. Finally, the image and text information, joint index are input into the image-text multi-modal convolutional neural network, and the probability of different stages of glaucoma is output, including the initial stage, the middle stage and the late stage of glaucoma. BRIEF DESCRIPTION OF DRAWINGS

[0124] Figure 1 The system flowchart of the present application. DETAILED DESCRIPTION

[0125] Specific implementation one: the glaucoma different stage discrimination system based on the multi-modal information network of the present embodiment comprises a data acquisition module, a data preprocessing module, a super-resolution reconstruction network model, an image-text multi-modal convolutional network model and a to-be-tested module.

[0126] The data acquisition module is used for acquiring the fundus optical coherence tomography OCT image training set, the fundus optical coherence tomography OCTA image training set, the RNFL image training set and the medical record information; the medical record information contains the gender and age information of the patient.

[0127] The data preprocessing module is used for preprocessing the fundus optical coherence tomography OCT image training set, the fundus optical coherence tomography OCTA image training set, the RNFL image training set and the medical record information collected by the data acquisition module, respectively, to obtain the preprocessed fundus optical coherence tomography OCT image training set, the fundus optical coherence tomography OCTA image training set, the RNFL image training set and the medical record information; the medical record information contains the gender and age information of the patient.

[0128] The super-resolution reconstruction network model is used for taking the preprocessed fundus optical coherence tomography (OCT) image training set, the fundus optical coherence tomography angiography (OCTA) image training set and the RNFL image training set as inputs of the super-resolution reconstruction network model, taking the high-resolution OCT image training set, the OCTA image training set and the RNFL image training set as outputs of the super-resolution reconstruction network model, and obtaining the trained super-resolution reconstruction network model;

[0129] The image-text multimodal convolution network model is used for taking the high-resolution OCT image, the OCTA image, the RNFL image and the age, gender and joint features output by the trained super-resolution reconstruction network model as inputs of the image-text multimodal convolution network model, taking the disease probability of glaucoma at different stages as an output of the image-text multimodal convolution network model, and obtaining the trained image-text multimodal convolution network model;

[0130] The to-be-tested module is used for obtaining the to-be-tested OCT image, the OCTA image, the RNFL image, the age and the gender information in the case, preprocessing the to-be-tested OCT image, the OCTA image, the RNFL image, the age and the gender information in the case to obtain the preprocessed OCT image, the OCTA image, the RNFL image, the gender, the age and the joint indicators, inputting the preprocessed OCT image, the OCTA image and the RNFL image into the trained super-resolution reconstruction network model, outputting the high-resolution OCT image, the OCTA image and the RNFL image by the generator in the trained super-resolution reconstruction network model, inputting the high-resolution OCT image, the OCTA image and the RNFL image output by the generator and the preprocessed gender, age and joint indicators into the trained image-text multimodal convolution network model, and outputting the probability of glaucoma at different stages by the trained image-text multimodal convolution network model, including the initial stage, the middle stage and the late stage of glaucoma.

[0131] The multi-modal information refers to the OCT image, the OCTA image, the RNFL image, the gender, the age and the joint indicators;

[0132] The multi-modal information network refers to the super-resolution reconstruction network model and the image-text multimodal convolution network model.

[0133] Specific implementation manner two: the difference between the embodiment and the specific implementation manner one is that the data collection module is used for collecting the fundus optical coherence tomography (OCT) image training set, the fundus optical coherence tomography angiography (OCTA) image training set, the RNFL image training set and the medical record information, and the specific process is as follows:

[0134] The fundus optical coherence tomography (OCT) image training set includes labeled normal OCT images, early glaucoma OCT images, middle glaucoma OCT images and late glaucoma OCT images.

[0135] The fundus optical coherence tomography OCTA image training set comprises labeled normal OCTA images, early glaucoma OCTA images, medium-term glaucoma OCTA images and late glaucoma OCTA images.

[0136] The RNFL image training set comprises labeled normal RNFL images, early glaucoma RNFL images, medium-term glaucoma RNFL images and late glaucoma RNFL images.

[0137] The OCT is an optical coherence tomography technology.

[0138] The OCTA is an optical coherence tomography angiography.

[0139] The RNFL is a retinal nerve fiber layer.

[0140] The other steps and parameters are the same as those in the first embodiment.

[0141] The third embodiment is different from the first or second embodiment in that the data preprocessing module is used to preprocess the fundus optical coherence tomography OCT image training set, the fundus optical coherence tomography OCTA image training set, the RNFL image training set and the medical record information collected by the data collection module, to obtain the preprocessed fundus optical coherence tomography OCT image training set, the fundus optical coherence tomography OCTA image training set, the RNFL image training set and the medical record information; the specific process is as follows:

[0142] 1) The fundus optical coherence tomography OCT image training set collected by the data collection module is preprocessed to obtain the preprocessed fundus optical coherence tomography OCT image training set; the specific process is as follows:

[0143] 11) The OCT image is denoised using a median filtering technology to obtain a denoised OCT image.

[0144] 12) The boundaries of the denoised OCT image are located using a sobel edge detection technology to obtain an OCT image with clear boundaries.

[0145] Since the boundaries between the retinal layers in the OCT image often exhibit obvious gradient changes, the sobel edge detection technology is used to locate the boundaries of the denoised OCT image.

[0146] 13) The OCT image with clear boundaries is segmented using an adaptive threshold technology into a vitreous body, a retinal nerve fiber layer, a ganglion cell layer, an inner nuclear layer, an outer plexiform layer, an outer nuclear layer and a retinal pigment epithelial layer.

[0147] 14), adjust the layered OCT image to 224x224 pixels (there are 7 layers inside the OCT image), and perform minimum-maximum normalization on the OCT image adjusted to 224x224 pixels to obtain the preprocessed OCT image training set;

[0148] By using the different reflection intensities of different retinal layers, an adaptive threshold technique is used to segment the boundary. The retina is divided into seven main layers;

[0149] 2), pre-process the fundus optical coherence tomography OCTA image training set collected by the data acquisition module to obtain the pre-processed fundus optical coherence tomography OCTA image training set; the specific process is:

[0150] 21), use a diffusion model to denoise the OCTA image to obtain a denoised OCTA image;

[0151] 22), calculate the Hessian matrix of each pixel point in the denoised OCTA image;

[0152] 23), calculate the eigenvalues of the Hessian matrix of each pixel point, and identify the potential blood vessel structure in the OCTA image based on the eigenvalues;

[0153] 24), design a matched filter to enhance the identified potential blood vessel structure and suppress background noise to obtain an OCTA image;

[0154] 25), use an adaptive threshold technique to process the OCTA image obtained in 24) to obtain a screened OCTA image;

[0155] 26), adjust the screened OCTA image obtained in 25) to 224x224 pixels, and perform minimum-maximum normalization on the OCTA image adjusted to 224x224 pixels to obtain the pre-processed OCTA image training set;

[0156] 3), pre-process the RNFL image training set collected by the data acquisition module to obtain the pre-processed RNFL image training set; the specific process is:

[0157] Adjust the RNFL image to 224x224 pixels, and perform minimum-maximum normalization on the RNFL image adjusted to 224x224 pixels to obtain the pre-processed RNFL image training set;

[0158] 4), pre-process the medical record information collected by the data acquisition module to obtain the pre-processed medical record information; the pre-processed medical record information includes pre-processed age, pre-processed gender, and pre-processed joint features.

[0159] Other steps and parameters are the same as embodiment one or two.

[0160] Embodiment four: different from one of embodiments one to three is that the medical record information collected by the data collection module is preprocessed in 4) to obtain preprocessed medical record information; the preprocessed medical record information includes preprocessed age, preprocessed gender, and preprocessed joint features.

[0161] The specific process is:

[0162] 41), pre-processing the age data to obtain preprocessed age data; the specific process is:

[0163] The age is divided into four stages: 0-17 years old, 18-39 years old, 40-79 years old, and 80 years old and above.

[0164] The value corresponding to 0-17 years old is 1; the value corresponding to 18-39 years old is 2; the value corresponding to 40-79 years old is 3; and the value corresponding to 80 years old and above is 4.

[0165] Since age is continuous data, converting it to discrete data can simplify the analysis process, so the age data is binned. According to research, there is one glaucoma patient for every 200 people in the population over 40 years old, and one glaucoma patient for every 8 people in the population over 80 years old. Therefore, the age is divided into 0-17 years old, 18-39 years old, 40-79 years old, and 80 years old and above.

[0166] 42), pre-processing the gender data to obtain preprocessed gender data; the specific process is:

[0167] The get_dummies function in the pandas library is used to convert the gender variable into a dummy variable.

[0168] Since gender is a non-numeric data type, the get_dummies function in the pandas library is used to convert the gender variable into a dummy variable, that is, one-hot encoding, that is, a new binary column (0 or 1) is created for each gender category, 1 represents belonging to the category, and 0 represents not belonging to the category.

[0169] 43), based on the preprocessed age data and the preprocessed gender data, calculate the joint features; represented as:

[0170] Joint features = Beta1 x age + Beta2 x gender + Beta3

[0171] Wherein, Beta1, Beta2, Beta3 are coefficients.

[0172] The coefficient Beta1, Beta2, Beta3 is obtained as follows:

[0173] Based on the joint feature training Logistic logistic regression model, the prediction probability of glaucoma is obtained, and the roc_auc_score function in the scikit-learn library is called to calculate the area AUC under the ROC curve corresponding to the Logistic logistic regression model, and thus a cycle is completed.

[0174] The values of the coefficients Beta1, Beta2, Beta3 are modified in each cycle, and the set of Beta1, Beta2, Beta3 coefficients that make the AUC value closest to 1 is taken as the optimal Beta1, Beta2, Beta3 coefficients.

[0175] Based on the optimal Beta1, Beta2, Beta3 coefficients, the joint feature is calculated.

[0176] The medical record contains the gender and age information of the patient. Glaucoma can occur in people of any age, but the probability of disease in people over 40 years old will greatly increase. Therefore, age is positively correlated with the probability of glaucoma, and the higher the age, the higher the probability of glaucoma. For women, before and after menopause, due to changes in hormone levels, the risk of glaucoma will also increase accordingly. Therefore, the gender factor also has a certain influence on the probability of glaucoma.

[0177] For continuous diagnostic indicators, a linear combination of multiple diagnostic indicators is usually used as a new indicator, and this new indicator is used as the basis for diagnosis. Since the outcome variable of diagnosis is a binary variable, i.e. whether or not to have glaucoma, the Logistic regression model is used to predict the probability of glaucoma based on the joint indicator, and the Beta coefficient (OR = exp(Beta), Beta = log(OR)) in the Logistic regression model is used to explain the influence of the independent variables (features) in the model on the dependent variable (outcome variable).

[0178] A for loop structure is used, and each time the Beta coefficients Beta1, Beta2 and Beta3 of the age, gender and intercept term are input, the joint feature of this cycle is obtained, a Logistic logistic regression model is trained using the joint feature, the prediction probability of glaucoma is obtained, and the roc_auc_score function in the scikit-learn library is called to calculate the area AUC under the ROC curve. Thus, a cycle is completed. In each cycle, the Beta coefficient is modified so that the calculated AUC value is close to 1. The set of Beta coefficients that make the AUC value closest to 1 is the Beta coefficient sought.

[0179] Other steps and parameters are the same as one of the first to third embodiments.

[0180] The fifth embodiment is different from one of the first to fourth embodiments in that: the super-resolution reconstruction network model is used to take the preprocessed fundus optical coherence tomography (OCT) image training set, the fundus optical coherence tomography angiography (OCTA) image training set and the RNFL image training set as inputs of the super-resolution reconstruction network model, and the high-resolution OCT image training set, the OCTA image training set and the RNFL image training set as outputs of the super-resolution reconstruction network model, to obtain the trained super-resolution reconstruction network model; and the specific process is:

[0181] 1) The super-resolution reconstruction network model sequentially includes a data conversion module, a generator network and a discriminator network.

[0182] The data conversion module includes a high-resolution conversion module and a low-resolution conversion module.

[0183] The generator network sequentially includes an input layer, an initial convolutional layer, a residual block, an up-sampling block, a final convolutional layer, and an output layer.

[0184] The discriminator network sequentially includes an input layer, a convolutional layer (Conv), a batch normalization layer (BatchNorm2d), a LeakyReLU activation function layer, an adaptive average pooling layer (AdaptiveAvgPool2d), a fully connected layer, and an output layer.

[0185] 2) The preprocessed OCT image is input into the data conversion module, and the data conversion module outputs a pair of low-resolution and high-resolution images corresponding to the OCT image.

[0186] The preprocessed OCTA image is input into the data conversion module, and the data conversion module outputs a pair of low-resolution and high-resolution images corresponding to the OCTA image.

[0187] The preprocessed RNFL image is input into the data conversion module, and the data conversion module outputs a pair of low-resolution and high-resolution images corresponding to the RNFL image.

[0188] 3) The low-resolution images corresponding to the OCT image, the OCTA image and the RNFL image output by the data conversion module are input into the generator network, and the generator network outputs high-resolution images.

[0189] 4) The high-resolution image corresponding to the OCT image output by the data conversion module or the high-resolution image corresponding to the OCT image output by the generator network is input into the discriminator network, and the discriminator network judges the probability P that the output image is the high-resolution image corresponding to the OCT image output by the data conversion module.

[0190] input the high-resolution image corresponding to the OCTA image output by the data conversion module or the high-resolution image corresponding to the OCTA image output by the generator network into the discriminator network, and the discriminator network judges the probability P that the output image is the high-resolution image corresponding to the OCTA image output by the data conversion module;

[0191] input the high-resolution image corresponding to the RNFL image output by the data conversion module or the high-resolution image corresponding to the RNFL image output by the generator network into the discriminator network, and the discriminator network judges the probability P that the output image is the high-resolution image corresponding to the RNFL image output by the data conversion module;

[0192] 6), repeat 1) to 5) until the probability P that the output image is the high-resolution image corresponding to the OCT image output by the data conversion module, the probability P that the output image is the high-resolution image corresponding to the OCTA image output by the data conversion module, and the probability P that the output image is the high-resolution image corresponding to the RNFL image output by the data conversion module all reach the maximum, and a trained deep learning module is obtained.

[0193] The other steps and parameters are the same as one of the first to fourth embodiments.

[0194] Embodiment six: the difference between this embodiment and one of the first to fifth embodiments is that in the 2), the preprocessed OCT image is input into the data conversion module, and the data conversion module outputs a pair of low-resolution and high-resolution images corresponding to the OCT image;

[0195] the preprocessed OCTA image is input into the data conversion module, and the data conversion module outputs a pair of low-resolution and high-resolution images corresponding to the OCTA image;

[0196] the preprocessed RNFL image is input into the data conversion module, and the data conversion module outputs a pair of low-resolution and high-resolution images corresponding to the RNFL image;

[0197] The specific process is as follows:

[0198] The data conversion module comprises a high-resolution conversion module and a low-resolution conversion module.

[0199] The preprocessed OCT image, OCTA image and RNFL image are respectively input into the high-resolution conversion module, and the high-resolution conversion module outputs high-resolution images;

[0200] The preprocessed OCT image, OCTA image and RNFL image are respectively input into the low-resolution conversion module, and the low-resolution conversion module outputs low-resolution images;

[0201] The pair of low-resolution and high-resolution images refers to images obtained by inputting the same preprocessed fundus optical coherence tomography (OCT) image into a high-resolution conversion module and a low-resolution conversion module, respectively.

[0202] The preprocessed OCT image, OCTA image and RNFL image are respectively input into the high-resolution conversion module, and the high-resolution conversion module outputs high-resolution images.

[0203] The preprocessed OCT image, OCTA image and RNFL image are respectively randomly cropped, and the randomly cropped OCT image, OCTA image and RNFL image are respectively converted into tensors.

[0204] The preprocessed OCT image, OCTA image and RNFL image are respectively input into the low-resolution conversion module, and the low-resolution conversion module outputs low-resolution images.

[0205] The preprocessed OCT image, OCTA image and RNFL image are respectively converted into PIL images, the sizes of the PIL images are adjusted to be equal to the size of the A image by bicubic interpolation, and the PIL images after size adjustment are converted into tensors.

[0206] The size of the A image is the size of the image after random cropping in the high-resolution conversion module divided by the up-sampling factor.

[0207] The up-sampling factor is set.

[0208] The other steps and parameters are the same as one of the first to fifth embodiments.

[0209] The seventh embodiment is different from one of the first to sixth embodiments in that the low-resolution images corresponding to the OCT image, OCTA image and RNFL image output by the data conversion module in the step 3) are respectively input into a generator network, and the generator network outputs high-resolution images.

[0210] The specific process is as follows:

[0211] The generator network sequentially includes an input layer, an initial convolution layer, a residual block, an up-sampling block, a final convolution layer, and an output layer.

[0212] The initial convolution layer sequentially includes a convolution layer (Conv) and a pre-activation function layer (PReLU).

[0213] The residual block sequentially includes a convolution layer (Conv), a batch normalization layer (BatchNorm2d), a pre-activation function layer (PReLU), a convolution layer (Conv), and a batch normalization layer (BatchNorm2d).

[0214] The upsampling block sequentially comprises a convolution layer (Conv), a PixelShuffle upsampling, and a pre-activation function layer (PReLU);

[0215] The final convolution layer sequentially comprises a convolution layer (Conv) and a tanh activation function layer;

[0216] The low-resolution images corresponding to the OCT image, the OCTA image, and the RNFL image output by the data conversion module are respectively input into the initial convolution layer, the residual block, the upsampling block, the final convolution layer, and the output layer in sequence through the input layer, and the output layer outputs the high-resolution images.

[0217] The other steps and parameters are the same as one of the first to sixth embodiments.

[0218] The eighth embodiment is different from one of the first to seventh embodiments in that the high-resolution image corresponding to the OCT image output by the data conversion module or the high-resolution image corresponding to the OCT image output by the generator network is input into the discriminator network in the 4), and the discriminator network judges the probability P that the output image is the high-resolution image corresponding to the OCT image output by the data conversion module;

[0219] The high-resolution image corresponding to the OCTA image output by the data conversion module or the high-resolution image corresponding to the OCTA image output by the generator network is input into the discriminator network, and the discriminator network judges the probability P that the output image is the high-resolution image corresponding to the OCTA image output by the data conversion module;

[0220] The high-resolution image corresponding to the RNFL image output by the data conversion module or the high-resolution image corresponding to the RNFL image output by the generator network is input into the discriminator network, and the discriminator network judges the probability P that the output image is the high-resolution image corresponding to the RNFL image output by the data conversion module;

[0221] The specific process is as follows:

[0222] The discriminator network sequentially comprises an input layer, a convolution layer (Conv), a batch normalization layer (BatchNorm2d), a LeakyReLU activation function layer, an adaptive average pooling layer (AdaptiveAvgPool2d), a fully connected layer, and an output layer;

[0223] The high-resolution image corresponding to the OCT image output by the data conversion module or the high-resolution image corresponding to the OCT image output by the generator network is sequentially input into a convolution layer (Conv), a batch normalization layer (BatchNorm2d), a LeakyReLU activation function layer, an adaptive average pooling layer (AdaptiveAvgPool2d), a full connection layer, and an output layer through an input layer, and the output layer determines a probability P that the output image is the high-resolution image corresponding to the OCT image output by the data conversion module.

[0224] The high-resolution image corresponding to the OCTA image output by the data conversion module or the high-resolution image corresponding to the OCTA image output by the generator network is sequentially input into a convolution layer (Conv), a batch normalization layer (BatchNorm2d), a LeakyReLU activation function layer, an adaptive average pooling layer (AdaptiveAvgPool2d), a full connection layer, and an output layer through an input layer, and the output layer determines a probability P that the output image is the high-resolution image corresponding to the OCTA image output by the data conversion module.

[0225] The high-resolution image corresponding to the RNFL image output by the data conversion module or the high-resolution image corresponding to the RNFL image output by the generator network is sequentially input into a convolution layer (Conv), a batch normalization layer (BatchNorm2d), a LeakyReLU activation function layer, an adaptive average pooling layer (AdaptiveAvgPool2d), a full connection layer, and an output layer through an input layer, and the output layer determines a probability P that the output image is the high-resolution image corresponding to the RNFL image output by the data conversion module.

[0226] The other steps and parameters are the same as one of the first to seventh embodiments.

[0227] The ninth embodiment is different from one of the first to eighth embodiments in that the steps 1) to 5) are repeatedly performed in the step 5) until the probabilities P that the output image is the high-resolution image corresponding to the OCT image output by the data conversion module, the high-resolution image corresponding to the OCTA image output by the data conversion module, and the high-resolution image corresponding to the RNFL image output by the data conversion module determined by the discriminator network all reach the maximum, and a trained deep learning module is obtained.

[0228] The specific process is as follows:

[0229] Define the loss function: the loss function includes GAN loss, perception loss PerceptionLoss (usually using the difference between pre-trained CNN feature maps to measure), image loss ImageLoss (such as L1 or L2 loss), and total variation loss TVLoss (for smoothing images); the specific process of each training cycle is as follows:

[0230] Set the total discriminator loss Loss_D = 1 - Loss_real + Loss_fake

[0231] Loss_real = -log(D(x))

[0232] D(x) is the average probability of the discriminator network outputting that the high-resolution image is the high-resolution image x converted by the data conversion module;

[0233] Loss_fake = -log(1-D(G(z)))

[0234] G(z) is the average probability of the discriminator network outputting that the high-resolution image is the high-resolution image generated by the generator network;

[0235] Use the Adam optimizer to update the discriminator parameter weight, so that the high-resolution image output by the discriminator network is closer to the high-resolution image converted by the data conversion module (the probability is close to 1);

[0236] Set the total generator loss Loss_G:

[0237] Loss_G = Adversarial_Loss + λ p Perception_Loss + λ i Image_Loss + λ tv TV_Loss

[0238] Adversarial_Loss is the adversarial loss, which is the average probability of the discriminator network judging that the generated high-resolution image is not the high-resolution image converted by the data conversion module;

[0239] Perception_Loss is the perception loss, which is the feature maps of the high-resolution image generated by the generator and the high-resolution image converted by the data conversion module input into the VGG16 network, respectively, and the VGG16 network outputs the feature maps of the high-resolution image generated by the generator and the high-resolution image converted by the data conversion module at the 31st layer, respectively. Then use the mean square error loss function MSELoss() to calculate the difference between the two feature maps, and the difference between the two feature maps is taken as the perception loss;

[0240] Image_Loss is an image loss, which is a difference between a pixel square sum of a high-resolution image generated by the generator and a pixel square sum of a high-resolution image converted by the data conversion module, calculated using a mean square error loss function MSELoss();

[0241] TV_Loss is a total variation loss, which is a total variation obtained by adding a square sum of differences between adjacent pixels in each row of the high-resolution image generated by the generator and a square sum of differences between adjacent pixels in each column of the high-resolution image generated by the generator, then dividing the square sum of the differences by a total number of pixels of the high-resolution image generated by the generator to obtain a loss value irrelevant to the image size, and finally multiplying the loss value by a set weight coefficient to obtain the final total variation loss;

[0242] wherein λ p , λ i , and λ tv are weights for balancing different loss terms;

[0243] The generator parameter weights are updated using the Adam optimizer so that the high-resolution image generated by the generator is closer to the high-resolution image converted by the data conversion module;

[0244] The steps 1) to 5) are repeatedly executed to alternately update the generator and the discriminator until the total discriminator loss and the total generator loss converge, and a trained super-resolution reconstruction network model is obtained.

[0245] “Alternately” means that the discriminator is fixed first and the generator is updated, and then the generator is fixed and the discriminator is updated. “Converge” means that the total discriminator loss and the total generator loss gradually stabilize and do not change significantly.

[0246] The other steps and parameters are the same as one of the first to eighth embodiments.

[0247] The tenth embodiment is different from one of the first to ninth embodiments in that the image-text multi-modal convolutional network model is used to take the high-resolution OCT image, the OCTA image, the RNFL image, and the age, gender, and joint features output by the trained super-resolution reconstruction network model as inputs of the image-text multi-modal convolutional network model, and take the glaucoma probability in different periods (including the early, middle, and late periods of glaucoma) as outputs of the image-text multi-modal convolutional network model to obtain a trained image-text multi-modal convolutional network model; and the specific process is as follows:

[0248] The image-text multi-modal convolutional network model comprises a feature extraction network and a feature fusion network;

[0249] The feature extraction network comprises an image feature extraction network and a text information feature extraction network;

[0250] The image feature extraction network sequentially comprises: an input layer, a convolution layer (Conv), a batch normalization layer (BatchNorm2d), a ReLU activation function layer, 50 residual blocks, a global average pooling layer (GlobalAveragePooling Layer), and an output layer.

[0251] Each residual block contains a convolution layer and an identity connection (the working process of each residual block is: the input of the convolution layer and the output of the convolution layer are added, and the added features are used as the input of the next residual block); the convolution group is composed of residual blocks, and between the convolution groups, a max pooling layer is used for down-sampling to reduce the size of the feature map and increase the receptive field.

[0252] The text information feature extraction network sequentially comprises: an input layer, an embedding layer, a convolution layer, a max pooling layer (MaxPooling Layer), and an output layer.

[0253] The feature fusion network sequentially comprises: an input layer, a concatenation layer, a fully connected layer, a softmax layer, and an output layer.

[0254] The high-resolution OCT image, OCTA image, and RNFL image are respectively input into the image feature extraction network, and the image feature extraction network respectively outputs the OCT image feature vector, OCTA image feature vector, and RNFL image feature vector.

[0255] The age, gender, and joint features are respectively input into the text information feature extraction network, and the text information feature extraction network respectively outputs the feature vectors of the age, gender, and joint features.

[0256] The OCT image feature vector, OCTA image feature vector, and RNFL image feature vector, and the feature vectors of the age, gender, and joint features are used as the input of the feature fusion network, and the feature fusion network outputs the disease probability of different stages of glaucoma (including the initial, medium, and late stages of glaucoma).

[0257] The cross-entropy loss function is defined.

[0258] The parameters of the feature extraction network and the feature fusion network are updated using the Adam optimization algorithm; until the cross-entropy loss function converges, the trained image-text multi-modal convolutional network model is obtained.

[0259] The other steps and parameters are the same as one of the first to ninth specific embodiments.

[0260] The present application can have other various embodiments, and those skilled in the art can make various corresponding changes and modifications according to the present application without departing from the spirit and essence of the present application, and these corresponding changes and modifications shall all belong to the protection scope of the claims of the present application.

Claims

1. A glaucoma different stage discrimination system based on a multi-modal information network, characterized in that: The system comprises a data acquisition module, a data preprocessing module, a super-resolution reconstruction network model, a graph-text multi-modal convolution network model and a to-be-tested module. The data acquisition module is used for acquiring an optical coherence tomography (OCT) image training set, an OCT angiography (OCTA) image training set, an RNFL image training set and medical record information, wherein the medical record information comprises gender and age information of a patient. The data preprocessing module is used for preprocessing the OCT image training set, the OCTA image training set, the RNFL image training set and the medical record information acquired by the data acquisition module, to obtain preprocessed OCT image training set, OCTA image training set, RNFL image training set and medical record information, wherein the medical record information comprises gender and age information of a patient. The super-resolution reconstruction network model is used for taking the preprocessed OCT image training set, OCTA image training set and RNFL image training set as inputs of the super-resolution reconstruction network model, taking high-resolution OCT image training set, OCTA image training set and RNFL image training set as outputs of the super-resolution reconstruction network model, and obtaining a trained super-resolution reconstruction network model. The graph-text multi-modal convolution network model is used for taking the high-resolution OCT image, OCTA image, RNFL image and age, gender and joint features output by the trained super-resolution reconstruction network model as inputs of the graph-text multi-modal convolution network model, taking a probability of glaucoma at different stages as an output of the graph-text multi-modal convolution network model, and obtaining a trained graph-text multi-modal convolution network model. The graph-text multi-modal convolution network model comprises a feature extraction network and a feature fusion network. The feature extraction network comprises an image feature extraction network and a text information feature extraction network. The image feature extraction network comprises, in sequence, an input layer, a convolution layer, a batch normalization layer, a ReLU activation function layer, 50 residual blocks, a global average pooling layer and an output layer. Each residual block comprises a convolution layer and an identity connection. The text information feature extraction network comprises, in sequence, an input layer, an embedding layer, a convolution layer, a max pooling layer and an output layer. The feature fusion network comprises, in sequence, an input layer, a concatenation layer, a full connection layer, a softmax layer and an output layer. The high-resolution OCT image, OCTA image and RNFL image are input into the image feature extraction network, and the image feature extraction network outputs an OCT image feature vector, an OCTA image feature vector and an RNFL image feature vector, respectively. The age, gender and joint features are input into the text information feature extraction network, and the text information feature extraction network outputs a feature vector of the age, gender and joint features, respectively. The OCT image feature vector, the OCTA image feature vector, the RNFL image feature vector and the feature vector of the age, gender and joint features are taken as inputs of the feature fusion network, and a probability of glaucoma at different stages is output by the feature fusion network. define a cross-entropy loss function; update the parameters of the feature extraction network and the feature fusion network using the Adam optimization algorithm; until the cross-entropy loss function converges, a trained image-text multi-modal convolutional network model is obtained; The acquisition process of the joint feature is: Based on the preprocessed age data and the preprocessed gender data, the joint feature is calculated; represented as: Joint feature = Beta1 x age + Beta2 x gender + Beta3 Wherein, Beta1, Beta2, Beta3 are coefficients; The acquisition process of the coefficients Beta1, Beta2, Beta3 is: Based on the joint feature, a Logistic logistic regression model is trained to obtain the prediction probability of glaucoma, and the area AUC under the ROC curve corresponding to the Logistic logistic regression model is calculated by calling the roc_auc_score function in the scikit-learn library, and thus a cycle is completed; In each cycle, the values of the coefficients Beta1, Beta2, Beta3 are modified, and the set of Beta1, Beta2, Beta3 coefficients that make the AUC value closest to 1 is taken as the optimal Beta1, Beta2, Beta3 coefficients; Based on the optimal Beta1, Beta2, Beta3 coefficients, the joint feature is calculated; The test module is used to obtain the OCT image, OCTA image, RNFL image, age and gender information in the case to be tested, and the OCT image, OCTA image, RNFL image, age and gender information in the case to be tested are preprocessed to obtain the preprocessed OCT image, OCTA image, RNFL image, gender, age and joint feature; the preprocessed OCT image, OCTA image and RNFL image are input into the trained super-resolution reconstruction network model, and the generator in the trained super-resolution reconstruction network model outputs high-resolution OCT image, OCTA image and RNFL image; the high-resolution OCT image, OCTA image and RNFL image output by the generator and the preprocessed gender, age and joint feature are input into the trained image-text multi-modal convolutional network model, and the trained image-text multi-modal convolutional network model outputs the probability of different stages of glaucoma, including early, middle and late stages of glaucoma. 2.The glaucoma different stage discrimination system based on the multi-modal information network of claim 1, wherein: The data acquisition module is used to collect fundus optical coherence tomography OCT image training set, fundus optical coherence tomography OCTA image training set, RNFL image training set, and medical record information; the specific process is: The fundus optical coherence tomography OCT image training set includes labeled normal OCT image, early glaucoma OCT image, middle glaucoma OCT image and late glaucoma OCT image; The fundus optical coherence tomography OCTA image training set includes labeled normal OCTA image, early glaucoma OCTA image, middle glaucoma OCTA image and late glaucoma OCTA image; The fundus optical coherence tomography OCTA image training set includes labeled normal OCTA image, early glaucoma OCTA image, middle glaucoma OCTA image and late glaucoma OCTA image; The RNFL image training set comprises labeled normal RNFL images, early glaucoma RNFL images, medium-term glaucoma RNFL images and late glaucoma RNFL images; The OCT is an optical coherence tomography technology; The OCTA is an optical coherence tomography angiography; The RNFL is a retinal nerve fiber layer. 3.The glaucoma different stage discrimination system based on the multi-modal information network of claim 2, wherein: The data preprocessing module is configured to preprocess the fundus optical coherence tomography (OCT) image training set, the fundus optical coherence tomography angiography (OCTA) image training set, the RNFL image training set and the medical record information collected by the data collection module, to obtain the preprocessed fundus optical coherence tomography (OCT) image training set, the preprocessed fundus optical coherence tomography angiography (OCTA) image training set, the preprocessed RNFL image training set and the preprocessed medical record information; and the specific process is as follows: 1) The fundus optical coherence tomography (OCT) image training set collected by the data collection module is preprocessed to obtain the preprocessed fundus optical coherence tomography (OCT) image training set; and the specific process is as follows: 11) The OCT image is denoised using a median filter technology to obtain a denoised OCT image; 12) The boundary of the denoised OCT image is located using a sobel edge detection technology to obtain an OCT image with clear boundaries; 13) The OCT image with clear boundaries is segmented using an adaptive threshold technology into a vitreous body, a retinal nerve fiber layer, a ganglion cell layer, an inner nuclear layer, an outer plexiform layer, an outer nuclear layer and a retinal pigment epithelial layer; 14) The segmented OCT image is adjusted to 224x224 pixels, and the OCT image adjusted to 224x224 pixels is normalized to obtain the preprocessed OCT image training set; 2) The fundus optical coherence tomography angiography (OCTA) image training set collected by the data collection module is preprocessed to obtain the preprocessed fundus optical coherence tomography angiography (OCTA) image training set; and the specific process is as follows: 21) The OCTA image is denoised using a diffusion model to obtain a denoised OCTA image; 22) The Hessian matrix of each pixel point in the denoised OCTA image is calculated; 23) The eigenvalues of the Hessian matrix of each pixel point are calculated, and the potential blood vessel structure in the OCTA image is identified based on the eigenvalues; 24) A matched filter is designed to enhance the identified potential blood vessel structure and suppress background noise to obtain an OCTA image; 25) The OCTA image obtained in 24) is processed using an adaptive threshold technology to obtain a screened OCTA image; 26) The screened OCTA image obtained in 25) is adjusted to 224x224 pixels, and the OCTA image adjusted to 224x224 pixels is normalized to obtain the preprocessed OCTA image training set; 3) The RNFL image training set collected by the data collection module is preprocessed to obtain the preprocessed RNFL image training set; and the specific process is as follows: The RNFL image is adjusted to 224*224 pixels, the RNFL image adjusted to 224*224 pixels is normalized, and a pretreated RNFL image training set is obtained; 4) The medical record information collected by the data collection module is preprocessed to obtain pretreated medical record information; The pretreated medical record information includes pretreated age, pretreated gender and pretreated joint features. 4.The glaucoma different stage discrimination system based on the multi-modal information network of claim 3, wherein: The pretreated medical record information includes pretreated age, pretreated gender and pretreated joint features. The specific process is as follows: 41) The age data is preprocessed to obtain pretreated age data; the specific process is as follows: The age is divided into four stages: 0-17 years old, 18-39 years old, 40-79 years old and 80 years old and above; The value corresponding to 0-17 years old is 1; the value corresponding to 18-39 years old is 2; the value corresponding to 40-79 years old is 3; and the value corresponding to 80 years old and above is 4; 42) The gender data is preprocessed to obtain pretreated gender data; the specific process is as follows: The gender variable is converted into a dummy variable using the get_dummies function in the pandas library.

5. The glaucoma different stage discrimination system based on a multi-modal information network according to claim 4, characterized in that: The super-resolution reconstruction network model is used to input the pretreated fundus optical coherence tomography (OCT) image training set, the pretreated fundus optical coherence tomography angiography (OCTA) image training set and the pretreated RNFL image training set as the input of the super-resolution reconstruction network model, and the high-resolution OCT image training set, the OCTA image training set and the RNFL image training set as the output of the super-resolution reconstruction network model, to obtain a trained super-resolution reconstruction network model; the specific process is as follows: 1) The super-resolution reconstruction network model sequentially includes a data conversion module, a generator network and a discriminator network; The data conversion module includes a high-resolution conversion module and a low-resolution conversion module; The generator network sequentially includes an input layer, an initial convolution layer, a residual block, an up-sampling block, a final convolution layer and an output layer; The discriminator network sequentially includes an input layer, a convolution layer, a batch normalization layer, a LeakyReLU activation function layer, an adaptive average pooling layer, a fully connected layer and an output layer; 2) The pretreated OCT image input data is converted into the data conversion module, and the data conversion module outputs a pair of low-resolution and high-resolution images corresponding to the OCT image; The pretreated OCTA image input data is converted into the data conversion module, and the data conversion module outputs a pair of low-resolution and high-resolution images corresponding to the OCTA image; The pretreated RNFL image is input into the data conversion module, and the data conversion module outputs a pair of low-resolution and high-resolution images corresponding to the RNFL image; 3) The low-resolution images corresponding to the OCT image, the OCTA image and the RNFL image output by the data conversion module are input into the generator network, and the generator network outputs high-resolution images respectively; 4) input the high-resolution image corresponding to the OCT image output by the data conversion module or the high-resolution image corresponding to the OCT image output by the generator network into the discriminator network, and the discriminator network determines the probability P that the output image is the high-resolution image corresponding to the OCT image output by the data conversion module; input the high-resolution image corresponding to the OCTA image output by the data conversion module or the high-resolution image corresponding to the OCTA image output by the generator network into the discriminator network, and the discriminator network determines the probability P that the output image is the high-resolution image corresponding to the OCTA image output by the data conversion module; input the high-resolution image corresponding to the RNFL image output by the data conversion module or the high-resolution image corresponding to the RNFL image output by the generator network into the discriminator network, and the discriminator network determines the probability P that the output image is the high-resolution image corresponding to the RNFL image output by the data conversion module; 5) repeat 1) to 5) until the probability P that the output image is the high-resolution image corresponding to the OCT image output by the data conversion module determined by the discriminator network, the probability P that the output image is the high-resolution image corresponding to the OCTA image output by the data conversion module determined by the discriminator network, and the probability P that the output image is the high-resolution image corresponding to the RNFL image output by the data conversion module determined by the discriminator network all reach the maximum, and a trained deep learning module is obtained.

6. The glaucoma different stage discrimination system based on a multi-modal information network according to claim 5, characterized in that: in 2), input the preprocessed OCT image into the data conversion module, and the data conversion module outputs a pair of low-resolution and high-resolution images corresponding to the OCT image; input the preprocessed OCTA image into the data conversion module, and the data conversion module outputs a pair of low-resolution and high-resolution images corresponding to the OCTA image; input the preprocessed RNFL image into the data conversion module, and the data conversion module outputs a pair of low-resolution and high-resolution images corresponding to the RNFL image; The specific process is as follows: The data conversion module comprises a high-resolution conversion module and a low-resolution conversion module; input the preprocessed OCT image, OCTA image and RNFL image into the high-resolution conversion module, respectively, and the high-resolution conversion module outputs a high-resolution image, respectively; input the preprocessed OCT image, OCTA image and RNFL image into the low-resolution conversion module, respectively, and the low-resolution conversion module outputs a low-resolution image, respectively; The process of inputting the preprocessed OCT image, OCTA image and RNFL image into the high-resolution conversion module, respectively, and the high-resolution conversion module outputting a high-resolution image, respectively, is as follows: randomly crop the preprocessed OCT image, OCTA image and RNFL image, respectively, and convert the randomly cropped OCT image, OCTA image and RNFL image into tensors, respectively; The process of inputting the preprocessed OCT image, OCTA image and RNFL image into the low-resolution conversion module, respectively, and the low-resolution conversion module outputting a low-resolution image, respectively, is as follows: The pre-processed OCT image, OCTA image and RNFL image are respectively converted into a PIL image, and the size of the PIL image is adjusted to be equal to that of the A image by bicubic interpolation, and the PIL image with the adjusted size is converted into a tensor; The size of the A image is the size of the image randomly cropped in the high-resolution conversion module divided by the up-sampling factor.

7. The glaucoma different stage discrimination system based on a multi-modal information network according to claim 6, characterized in that: The low-resolution images corresponding to the OCT image, OCTA image and RNFL image output by the data conversion module in 3) are respectively input into the generator network, and the generator network outputs high-resolution images; The specific process is as follows: The generator network sequentially comprises an input layer, an initial convolution layer, a residual block, an up-sampling block, a final convolution layer and an output layer. The initial convolution layer sequentially comprises a convolution layer and a pre-activation function layer. The residual block sequentially comprises a convolution layer, a batch normalization layer, a pre-activation function layer, a convolution layer and a batch normalization layer. The up-sampling block sequentially comprises a convolution layer, a PixelShuffle up-sampling and a pre-activation function layer. The final convolution layer sequentially comprises a convolution layer and a tanh activation function layer. The low-resolution images corresponding to the OCT image, OCTA image and RNFL image output by the data conversion module are respectively input into the initial convolution layer, residual block, up-sampling block, final convolution layer and output layer in sequence through the input layer, and the output layer outputs high-resolution images. 8.The glaucoma different stage discrimination system based on the multi-modal information network of claim 7, wherein: The high-resolution image corresponding to the OCT image output by the data conversion module or the high-resolution image corresponding to the OCT image output by the generator network in 4) is input into the discriminator network, and the discriminator network judges the probability P that the output image is the high-resolution image corresponding to the OCT image output by the data conversion module. The high-resolution image corresponding to the OCTA image output by the data conversion module or the high-resolution image corresponding to the OCTA image output by the generator network is input into the discriminator network, and the discriminator network judges the probability P that the output image is the high-resolution image corresponding to the OCTA image output by the data conversion module. The high-resolution image corresponding to the RNFL image output by the data conversion module or the high-resolution image corresponding to the RNFL image output by the generator network is input into the discriminator network, and the discriminator network judges the probability P that the output image is the high-resolution image corresponding to the RNFL image output by the data conversion module. The specific process is as follows: The discriminator network sequentially comprises an input layer, a convolution layer, a batch normalization layer, a LeakyReLU activation function layer, an adaptive average pooling layer, a fully connected layer and an output layer. The high-resolution image corresponding to the OCT image output by the data conversion module or the high-resolution image corresponding to the OCT image output by the generator network is input into the convolution layer, batch normalization layer, LeakyReLU activation function layer, adaptive average pooling layer, fully connected layer and output layer in sequence through the input layer, and the output layer judges the probability P that the output image is the high-resolution image corresponding to the OCT image output by the data conversion module. The high-resolution image corresponding to the OCTA image output by the data conversion module or the high-resolution image corresponding to the OCTA image output by the generator network is sequentially input to the convolution layer, the batch normalization layer, the LeakyReLU activation function layer, the adaptive average pooling layer, the full connection layer and the output layer through the input layer, and the output layer judges the probability P that the output image is the high-resolution image corresponding to the OCTA image output by the data conversion module; The high-resolution image corresponding to the RNFL image output by the data conversion module or the high-resolution image corresponding to the RNFL image output by the generator network is sequentially input to the convolution layer, the batch normalization layer, the LeakyReLU activation function layer, the adaptive average pooling layer, the full connection layer and the output layer through the input layer, and the output layer judges the probability P that the output image is the high-resolution image corresponding to the RNFL image output by the data conversion module.

9. The glaucoma different stage discrimination system based on a multi-modal information network according to claim 8, characterized in that: The 5) is repeatedly executed 1) to 5) until the probability P that the output image is the high-resolution image corresponding to the OCT image output by the data conversion module is judged by the discriminator network, the probability P that the output image is the high-resolution image corresponding to the OCTA image output by the data conversion module is judged by the discriminator network, and the probability P that the output image is the high-resolution image corresponding to the RNFL image output by the data conversion module is judged by the discriminator network all reach the maximum, and a trained deep learning module is obtained; The specific process is: Set the total discriminator loss Loss_D = 1-Loss_real+Loss_fake Loss_real = -log(D(x)) D(x) is the average probability that the discriminator network outputs the high-resolution image as the high-resolution image x converted by the data conversion module; Loss_fake = -log(1-D(G(z))) G(z) is the average probability that the discriminator network outputs the high-resolution image as the high-resolution image generated by the generator network; The Adam optimizer is used to update the discriminator parameter weight, so that the high-resolution image output by the discriminator network is closer to the high-resolution image converted by the data conversion module; Set the total generator loss Loss_G: Loss_G = Adversarial_Loss + λ p Perception_Loss + λ i Image_Loss + λ tv TV_Loss Adversarial_Loss is the adversarial loss; Perception_Loss is the perception loss; Image_Loss is the image loss, which is the difference between the pixel square sum of the high-resolution image generated by the generator and the pixel square sum of the high-resolution image converted by the data conversion module, calculated using the mean square error loss function MSELoss(); TV_Loss is the total variation loss; where λ p , λ i , λ tv are weights; The Adam optimizer is used to update the generator parameter weight, so that the high-resolution image generated by the generator is closer to the high-resolution image converted by the data conversion module; Repeat 1) to 5), alternately update the generator and the discriminator, until the total discriminator loss and the total generator loss converge, and a trained super-resolution reconstruction network model is obtained.

Citation Information

Patent Citations

  • Anterior chamber angle image grading method based on visual text fusion

    CN115423790A

  • Multi-mode glaucoma image recognition method and system

    CN118196584A