Newborn fundus image classification method and imaging method based on multimodal data
Through the neonatal fundus image classification method with multimodal data, the image and text feature generator are used and combined with multiple modules for training, the reliability and accuracy of single modal classification is solved, and more efficient neonatal fundus image classification is achieved.
Patent Information
- Application Number
- CN202510614711.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-05-14
AI Technical Summary
In the prior art, the classification of fundus images of newborns relies on a single mode, resulting in poor classification reliability and accuracy, and it is difficult to introduce additional modal data for massive amounts of manual labeling information.
The neonatal fundus image classification method adopts multimodal data, obtains the neonatal fundus image data, performs image processing and text annotation, extracts text and image features, and uses pre-trained image encoder and text encoder to generate pseudo-text features. The classification model is trained in combination with the image prediction module, pseudo-text prediction module and fusion module, and finally generates accurate results for neonatal fundus image classification.
The classification of fundus images of neonatal babies based on multimodal data is realized, which improves the reliability and accuracy of classification, which is significantly better than the single-modal method.
Smart Images

Figure CN120126204B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing, and particularly relates to a method for classifying neonatal fundus images based on multi-modal data and an imaging method. Background Art
[0002] The classification of neonatal fundus images is of great significance in both clinical and basic medical research. At present, for the classification of neonatal fundus images, it is still clinical medical imaging personnel who classify the fundus images according to their own experience and technical level. However, this manual classification scheme for neonatal fundus images is not only time-consuming and laborious, but also has poor reliability.
[0003] At present, although some researchers have proposed a classification scheme for neonatal fundus images based on deep learning, this type of scheme mainly relies on a single image modality for image classification. Neonatal fundus images have characteristics such as light color, high transparency, uneven sparse blood vessels, and unclear features. The existing classification schemes based on a single fundus image modality also have relatively poor accuracy. In addition, although introducing additional modality data can improve the accuracy of neonatal fundus images, this requires a large amount of manually labeled information, which is difficult to achieve at present. Summary of the Invention
[0004] One of the purposes of the present invention is to provide a method for classifying neonatal fundus images based on multi-modal data with high reliability and good accuracy.
[0005] Another purpose of the present invention is to provide an imaging method including the method for classifying neonatal fundus images based on multi-modal data.
[0006] The method for classifying neonatal fundus images based on multi-modal data provided by the present invention includes the following steps:
[0007] S1. Obtain existing neonatal fundus image data;
[0008] S2. Perform image processing on the neonatal fundus images obtained in step S1 to construct a training dataset; at the same time, select several neonatal fundus images for text annotation;
[0009] S3. Extract the text features of the neonatal fundus images with text annotation in an offline state, and at the same time extract the image features of the corresponding neonatal fundus images;
[0010] S4. Use the text features and the corresponding image features obtained in step S3 as a first training set, and train a text feature generator based on a pre-trained image encoder and text encoder; the text feature generator is used to generate instance-level pseudo-text features of the input image;
[0011] S5. Construct an initial model for classifying neonatal fundus images that includes an image prediction module, a pseudo-text prediction module, and a fusion module;
[0012] Among them, construct an image prediction module based on a pre-trained image encoder, an attention mechanism, and a linear layer, which is used to generate an image classification prediction result for the input image; construct a pseudo-text prediction module based on a pre-trained image encoder, the trained text feature generator obtained in step S4, an attention mechanism, and a linear layer, which is used to first generate instance-level pseudo-text features for the input image, and then generate a pseudo-text classification prediction result based on the instance-level pseudo-text features; construct a fusion module based on an information entropy scheme, which is used to fuse the generated image classification prediction result and pseudo-text classification prediction result to obtain the final neonatal fundus image classification result;
[0013] S6. Use the training dataset constructed in step S2 to train the initial model for classifying neonatal fundus images constructed in step S5 to obtain a trained neonatal fundus image classification model;
[0014] S7. Use the neonatal fundus image classification model obtained in step S6 to classify actual neonatal fundus images.
[0015] The described step S2 specifically includes the following steps:
[0016] Obtain neonatal fundus images: Among them, for each neonate, obtain several fundus images, label the category of the neonate, but do not label the category of each fundus image;
[0017] Adjust the resolution of the obtained neonatal fundus images and perform a normalization operation to construct a training dataset;
[0018] Among the obtained neonatal fundus images, select several neonatal fundus images for text description annotation.
[0019] The described step S3 specifically includes the following steps:
[0020] In an offline state, use a text-based model to extract text features from the neonatal fundus images with text annotations;
[0021] At the same time, use a vision-based model to extract image features from the neonatal fundus images with text annotations.
[0022] The described text-based model includes the Bert-base-Chinese model.
[0023] The described visual base model includes the ResNet-50 model or the RETFound model; when the ResNet-50 model is adopted, the extracted image features are the input of the global average pooling operation in the ResNet-50 model; when the RETFound model is adopted, the extracted image features are the input of the last layer transformer in the RETFound model.
[0024] The described step S4 specifically includes the following steps:
[0025] Use the text features and corresponding image features obtained in step S3 as the first training set;
[0026] Process the image features sequentially through a pre-trained image encoder and a text feature generator to be trained to generate pseudo-text features;
[0027] Process the text features through a pre-trained text encoder to generate real text features;
[0028] During training, calculate the error loss between the pseudo-text features and the real text features to train the text feature generator.
[0029] The described pre-trained image encoder includes a pre-trained RETFound model or a pre-trained ResNet-50 model; the described pre-trained text encoder includes a pre-trained Bert-base-Chinese model; the text feature generator is specifically a linear mapping layer.
[0030] The described calculation of the error loss between the pseudo-text features and the real text features is specifically the calculation of the mean square error loss between the pseudo-text features and the real text features.
[0031] The described step S5 includes the following steps:
[0032] Image prediction module: Based on a pre-trained image encoder, a multi-head self-attention mechanism, a linear mapping layer, and an attention weight calculation scheme, construct an image prediction module; the image prediction module is used to generate an image classification prediction result of the input image;
[0033] Pseudo-text prediction module: Based on a pre-trained image encoder, the trained text feature generator obtained in step S4, an attention weight calculation scheme, and a linear layer, construct a pseudo-text prediction module; the input image first passes through a pre-trained image encoder and the trained text feature generator obtained in step S4 to generate instance-level pseudo-text features, and then the instance-level pseudo-text features are processed based on the attention weight calculation scheme and the linear layer to generate a pseudo-text classification prediction result;
[0034] Fusion module: Calculate the weight values of the image classification prediction result and the pseudo-text classification prediction result based on the information entropy calculation scheme, and combine the weighted summation scheme to fuse the prediction results, and finally obtain the classification result of the neonatal fundus image.
[0035] The processing process of the image prediction module specifically includes the following steps:
[0036] The input image is processed by a pre-trained image encoder to obtain image dense features;
[0037] Concatenate a learnable classification token to the obtained image dense features, denoted as: where is the token group after adding position encoding; is the classification token; is the image patch token; is the and concatenation operation; is the position encoding;
[0038] Then, the obtained is processed through a linear mapping layer, information interaction is performed through a multi-head self-attention mechanism, and processing is performed through a linear layer. Then, the result is summed with to obtain instance-level image features , denoted as: In the formula is the linear layer processing function; is the linear mapping layer processing function; is the multi-head self-attention mechanism processing function;
[0039] The obtained instance-level image features are processed through an attention weight calculation scheme and then processed through a linear layer to obtain an image classification prediction result, denoted as: In the formula is the image classification prediction result; is the processing process of the attention weight calculation scheme of the image prediction module, and , is the encoding of the ii-th image instance-level feature, is the weight value of and , w is the first parameter to be learned, and V is the second parameter to be learned.
[0040] The processing process of the pseudo-text prediction module specifically includes the following steps:
[0041] The input image generates instance-level pseudo-text features through a pre-trained image encoder and the trained text feature generator obtained in step S4;
[0042] The instance-level pseudo-text features are processed through an attention weight calculation scheme and then through a linear layer to obtain the pseudo-text classification prediction result, expressed as: In the formula is the pseudo-text classification prediction result; is the instance-level pseudo-text feature; is the processing process of the attention weight calculation scheme of the pseudo-text prediction module, and , where is the k-th text instance-level feature encoding, is the weight value of and , is the third parameter to be learned, is the fourth parameter to be learned.
[0043] The processing process of the fusion module specifically includes the following steps:
[0044] Calculate the image classification prediction result of the image information entropy is , calculate the text information entropy of the pseudo-text classification prediction result is for ; where is the information entropy calculation function;
[0045] Take the maximum value of the image information entropy and the text information entropy to obtain the extreme value information entropy is ;
[0046] Use the following formula to calculate the image weight and text weight: In the formula is the image weight; is the text weight;
[0047] Finally, calculate the prediction result is .
[0048] The present invention also provides an imaging method, which includes the above-mentioned neonatal fundus image classification method based on multi-modal data, and further includes the following steps:
[0049] S8. Label and re-image the classification result of the neonatal fundus image obtained in step S7 on the neonatal fundus image to obtain a neonatal fundus image with the classification result.
[0050] The neonatal fundus image classification method and imaging method based on multimodal data provided by the present invention generate pseudo-text features through a pre-trained text feature generator, and construct a neonatal fundus image classification model based on image modal data and text modal data based on the text feature generator, attention mechanism, linear layer and information entropy scheme. Therefore, the present invention can not only realize the classification of neonatal fundus images based on multimodal data, but also has higher reliability and better accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a schematic flowchart of the method of the classification method of the present invention.
[0052] Figure 2 It is a schematic flowchart of the method of the imaging method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0053] As Figure 1 shown is a schematic flowchart of the method of the method of the present invention: The neonatal fundus image classification method based on multimodal data provided by the present invention includes the following steps:
[0054] S1. Obtain existing neonatal fundus image data.
[0055] S2. Perform image processing on the neonatal fundus images obtained in step S1 to construct a training dataset; at the same time, select several neonatal fundus images for text annotation; specifically include the following steps:
[0056] Obtain neonatal fundus images: Among them, for each neonate, obtain several fundus images, label the category of the neonate, but do not label the category of each fundus image;
[0057] Adjust the resolution of the obtained neonatal fundus images (preferably adjusted to ), and perform a normalization operation to construct a training dataset;
[0058] Among the obtained neonatal fundus images, select several neonatal fundus images for text description annotation.
[0059] S3. In the offline state, extract the text features of the neonatal fundus images with text annotation, and at the same time extract the image features of the corresponding neonatal fundus images; specifically include the following steps:
[0060] In the offline state, use a text-based model to extract text features from the neonatal fundus images with text annotation;
[0061] At the same time, use a vision-based model to extract image features from the neonatal fundus images with text annotation;
[0062] Among them, the text base model includes the Bert-base-Chinese model; the visual base model includes the ResNet-50 model or the RETFound model; when the ResNet-50 model is adopted, the extracted image features are the input of the global average pooling operation in the ResNet-50 model; when the RETFound model is adopted, the extracted image features are the input of the last layer transformer in the RETFound model.
[0063] S4. Use the text features and corresponding image features obtained in step S3 as the first training set, and based on the pre-trained image encoder and text encoder, train the text feature generator; the text feature generator is used to generate instance-level pseudo-text features of the input image.
[0064] Specifically, it includes the following steps:
[0065] Use the text features and corresponding image features obtained in step S3 as the first training set;
[0066] Process the image features sequentially through the pre-trained image encoder and the text feature generator to be trained to generate pseudo-text features;
[0067] Process the text features through the pre-trained text encoder to generate real text features;
[0068] During training, calculate the error loss between the pseudo-text features and the real text features to achieve the training of the text feature generator;
[0069] Among them, the pre-trained image encoder includes the pre-trained RETFound model or the pre-trained ResNet-50 model; the pre-trained text encoder includes the pre-trained Bert-base-Chinese model; the text feature generator is specifically a linear mapping layer;
[0070] During training, calculating the error loss between the pseudo-text features and the real text features is specifically calculating the mean square error loss between the pseudo-text features and the real text features;
[0071] At the same time, during training, only the parameters of the text feature generator are trained, and the parameters of the image encoder and the text encoder remain unchanged.
[0072] S5. Build an initial neonatal fundus image classification model including an image prediction module, a pseudo-text prediction module, and a fusion module;
[0073] Among them, an image prediction module is constructed based on a pre-trained image encoder, an attention mechanism, and a linear layer, which is used to generate an image classification prediction result of the input image; a pseudo-text prediction module is constructed based on the pre-trained image encoder, the trained text feature generator obtained in step S4, the attention mechanism, and the linear layer, which is used to first generate instance-level pseudo-text features of the input image, and then generate a pseudo-text classification prediction result based on the instance-level pseudo-text features; a fusion module is constructed based on an information entropy scheme, which is used to fuse the generated image classification prediction result and the pseudo-text classification prediction result to obtain the final neonatal fundus image classification result.
[0074] Specifically, the implementation includes the following steps:
[0075] Image prediction module: An image prediction module is constructed based on a pre-trained image encoder, a multi-head self-attention mechanism, a linear mapping layer, and an attention weight calculation scheme; the image prediction module is used to generate an image classification prediction result of the input image;
[0076] Pseudo-text prediction module: A pseudo-text prediction module is constructed based on a pre-trained image encoder, the trained text feature generator obtained in step S4, an attention weight calculation scheme, and a linear layer; the input image first generates instance-level pseudo-text features through the pre-trained image encoder and the trained text feature generator obtained in step S4, and then the instance-level pseudo-text features are processed based on the attention weight calculation scheme and the linear layer to generate a pseudo-text classification prediction result;
[0077] Fusion module: Based on an information entropy calculation scheme, the weight values of the image classification prediction result and the pseudo-text classification prediction result are calculated, and the prediction results are fused by combining a weighted summation scheme, and finally the classification result of the neonatal fundus image is obtained.
[0078] Among them, the processing process of the image prediction module specifically includes the following steps:
[0079] The input image is processed by a pre-trained image encoder to obtain image dense features;
[0080] The obtained image dense features are concatenated with a learnable classification token, expressed as: Among them is the token group after adding position encoding; is the classification token; is the image patch token; is to and The operation of concatenating; is the position encoding;
[0081] Then the obtained Processed sequentially through a linear mapping layer, information interaction is carried out through a multi-head self-attention mechanism and processed through a linear layer, and then the result is combined with for summation to obtain instance-level image features , expressed as: In the formula is the processing function of the linear layer; is the processing function of the linear mapping layer; is the processing function of the multi-head self-attention mechanism;
[0082] The obtained instance-level image features are processed through an attention weight calculation scheme and then processed through a linear layer to obtain an image classification prediction result, expressed as: In the formula is the image classification prediction result; is the processing process of the attention weight calculation scheme of the image prediction module, and , is the ii-th image instance-level feature encoding, is the weight value of and , w is the first parameter to be learned, and V is the second parameter to be learned.
[0083] The processing process of the pseudo-text prediction module specifically includes the following steps:
[0084] The input image generates instance-level pseudo-text features through a pre-trained image encoder and the trained text feature generator obtained in step S4;
[0085] The instance-level pseudo-text features are processed through an attention weight calculation scheme and then processed through a linear layer to obtain a pseudo-text classification prediction result, expressed as: In the formula is the pseudo-text classification prediction result; is the instance-level pseudo-text feature; is the processing process of the attention weight calculation scheme of the pseudo-text prediction module, and , where is the kk-th text instance-level feature encoding, is the weight value of and , is the third parameter to be learned, is the fourth parameter to be learned.
[0086] The processing process of the fusion module specifically includes the following steps:
[0087] Considering that there are decision differences between the pseudo-text classification prediction results and the image classification prediction results, but since they are different modalities and have undergone different processing processes, each has a certain decision-making ability. Therefore, information entropy is used to weight the two decision scores, aiming to accurately quantify the data purity and uncertainty under each result; the image classification prediction result is calculated of the image information entropy is , and the pseudo-text classification prediction result is calculated of the text information entropy is ; where is the information entropy calculation function;
[0088] Take the maximum value of the image information entropy and the text information entropy to obtain the extreme value information entropy which is ;
[0089] The image weight and the text weight are calculated using the following formula: In the formula is the image weight; is the text weight; a higher information entropy means more instability and a lower weight will be obtained, while a lower information entropy will obtain a higher weight to affect the final decision;
[0090] Finally, the calculated prediction result is .
[0091] S6. Use the training dataset constructed in step S2 to train the initial neonatal fundus image classification model constructed in step S5 to obtain the trained neonatal fundus image classification model;
[0092] During the training process, the parameters of the pre-trained image encoder and the parameters of the trained text feature generator obtained in step S4 remain unchanged.
[0093] S7. Use the neonatal fundus image classification model obtained in step S6 to classify actual neonatal fundus images.
[0094] The following combines an embodiment to illustrate the effect of the method of the present invention:
[0095] Dataset introduction: The data in this embodiment comes from a multi-center research project on neonatal intraocular disease screening led by the Second Xiangya Hospital of Central South University and approved by the ethics committee. The dataset includes 115,621 fundus images obtained from 8,886 neonates using RetCam3 during the period from 2015 to 2019; the resolution of these images is Pixels, and multiple retinal images of each newborn were taken at different angles; four professionals assigned classification labels based on the corresponding multiple retinal images; to ensure accuracy and consistency, all four professionals jointly reviewed the uncertain or ambiguous samples and obtained clear annotations; 8,886 newborns were randomly divided into three subsets: a training set, a validation set, and a test set according to the ratio; the dataset used in the first stage was a subset of the training set, which contained 466 sample numbers and a total of 4,188 instances; this means that a total of 4,188 image-text pairs participated in the training of the pseudo-text generator in the first stage.
[0096] The experimental table shows the experimental results of different methods on the above dataset, including the comparison results of four performance indicators: sensitivity, specificity, F1-score, and AUC; the experiment is mainly divided into two modules, which are experiments using the features extracted by different feature extractors; the models are all trained on the training set and the validation set, and then tested on the test set.
[0097] The method of the present invention is compared with existing solutions; in the comparative experiments, the visual foundation model uses the ResNet-50 model or the RETFound model, and comparative experiments are carried out separately; in the existing solutions, the MaxMIL solution is the one proposed by Hu J, Chen Y, Zhong J, et al. in the 2019 paper "Automated analysis for retinopathy of prematurity by deep neural network."; ABMIL is the one proposed by Ilse M, Tomczak J, Welling M. in the 2018 paper "Attention-based deep multiple instance learning."; Gated-ABMIL is the one proposed by Ilse M, Tomczak J, Welling M. in the 2018 paper "Attention-based deep multiple instance learning."; DSMIL is the one proposed by Li B, Li Y, Eliceiri K W. in the 2021 paper "Dual-stream multiple instance learning network for whole slide image classification with self-supervised contrastive learning."; TransMIL is the one proposed by Shao Z, Bian H, Chen Y, et al in the 2021 paper "TransMIL: Transformer based correlated multiple instance learning for whole slide image classification."; R2T-MIL is the one proposed by Tang W, Zhou F, Huang S, et al. in the 2024 paper "Feature Re-Embedding: Towards foundation model-level performance in computational pathology."
[0098] The experimental data are shown in Table 1:
[0099] Table 1 Schematic table of experimental data comparison
[0100]
[0101] As can be seen from Table 1, when using ResNet-50 as the visual base model, the proposed solution of the present invention performs excellently in all evaluation metrics. Specifically, the sensitivity of the proposed solution of the present invention reaches 89.91%, the specificity is 94.70%, the F1 score is 88.68%, and the AUC value is as high as 96.67%; these values are significantly better than other solutions such as MaxMIL, ABMIL, Gated-ABMIL, DSMIL, TransMIL, and R2T-MIL, etc.; for example, the sensitivity of MaxMIL is 75.09%, the specificity is 87.91%, the F1 score is 79.01%, and the AUC value is 90.32%; while the corresponding values of ABMIL are 74.39%, 86.77%, 80.36%, and 93.62% respectively; this indicates that the proposed solution of the present invention has obvious advantages in terms of recognition accuracy and generalization ability. When using RETFound as the visual base model, the proposed solution of the present invention also performs excellently; its sensitivity is 89.55%, the specificity is 95.00%, the F1 score is 88.20%, and the AUC value is 96.48%; in comparison, the sensitivity of MaxMIL is 73.12%, the specificity is 86.22%, the F1 score is 74.44%, and the AUC value is 92.33%; the corresponding values of ABMIL are 74.93%, 87.63%, 77.87%, and 93.05% respectively; this further verifies the effectiveness and robustness of the proposed solution of the present invention.
[0102] In summary, whether using ResNet-50 or RETFound as the visual base model, the proposed solution of the present invention has achieved the best results in multiple key metrics such as sensitivity, specificity, F1 score, and AUC value, demonstrating its excellent performance and application potential in the relevant field.
[0103] As Figure 2 shown in the schematic flowchart of the method of the imaging method of the present invention: The imaging method disclosed in the present invention includes the above-mentioned method for classifying neonatal fundus images based on multimodal data; this imaging method includes the following steps:
[0104] S1. Obtain existing neonatal fundus image data;
[0105] S2. Perform image processing on the neonatal fundus images obtained in step S1 to construct a training dataset; at the same time, select several neonatal fundus images for text annotation;
[0106] S3. In the offline state, extract the text features of the neonatal fundus images with text annotation, and at the same time extract the image features of the corresponding neonatal fundus images;
[0107] S4. Use the text features and corresponding image features obtained in step S3 as the first training set, and based on the pre-trained image encoder and text encoder, train the text feature generator; the text feature generator is used to generate instance-level pseudo-text features of the input image;
[0108] S5. Construct an initial neonatal fundus image classification model including an image prediction module, a pseudo-text prediction module, and a fusion module;
[0109] Among them, the image prediction module is constructed based on the pre-trained image encoder, the attention mechanism, and the linear layer, and is used to generate the image classification prediction result of the input image; the pseudo-text prediction module is constructed based on the pre-trained image encoder, the trained text feature generator obtained in step S4, the attention mechanism, and the linear layer, and is used to first generate the instance-level pseudo-text features of the input image, and then generate the pseudo-text classification prediction result based on the instance-level pseudo-text features; the fusion module is constructed based on the information entropy scheme, and is used to fuse the generated image classification prediction result and the pseudo-text classification prediction result to obtain the final neonatal fundus image classification result;
[0110] S6. Use the training data set constructed in step S2 to train the initial neonatal fundus image classification model constructed in step S5 to obtain the trained neonatal fundus image classification model;
[0111] S7. Use the neonatal fundus image classification model obtained in step S6 to perform the classification of actual neonatal fundus images;
[0112] S8. Label and re-image the classification result of the neonatal fundus image obtained in step S7 on the neonatal fundus image to obtain a neonatal fundus image with the classification result.
[0113] The imaging method provided by the present invention can be directly applied to existing neonatal fundus image devices (such as neonatal fundus image imaging systems), or directly applied to a terminal (such as a computer); specifically in application, use the existing scheme to obtain actual neonatal fundus images, and then input the obtained data into the corresponding machine device or terminal. At this time, the machine device or terminal can obtain the actual neonatal fundus image classification result according to the imaging method disclosed in the present invention, and display the classification result on the original image through different types of representations (such as colors), and then perform re-imaging and output; at this time, the output image is the image with the fundus image classification result, and this image can reflect the actual neonatal fundus image and the classification result, thus greatly facilitating the subsequent work of clinical medical staff and laboratory experimenters.
Claims
1. A method for classifying neonatal fundus images based on multimodal data, characterized in that It includes the following steps: S1. Obtain the existing neonatal fundus image data; S2. Perform image processing on the neonatal fundus images obtained in step S1 to construct a training dataset; meanwhile, select several neonatal fundus images for text annotation; S3. In the offline state, extract the text features of the neonatal fundus images with text annotation, and at the same time extract the image features of the corresponding neonatal fundus images; S4. Use the text features and the corresponding image features obtained in step S3 as the first training set, and based on the pre-trained image encoder and text encoder, train the text feature generator; the text feature generator is used to generate instance-level pseudo-text features of the input image; S5. Construct an initial neonatal fundus image classification model including an image prediction module, a pseudo-text prediction module, and a fusion module; Among them, the image prediction module is constructed based on the pre-trained image encoder, attention mechanism, and linear layer, and is used to generate the image classification prediction result of the input image; the pseudo-text prediction module is constructed based on the pre-trained image encoder, the trained text feature generator obtained in step S4, attention mechanism, and linear layer, and is used to first generate the instance-level pseudo-text features of the input image, and then generate the pseudo-text classification prediction result based on the instance-level pseudo-text features; the fusion module is constructed based on the information entropy scheme and is used to fuse the generated image classification prediction result and pseudo-text classification prediction result to obtain the final neonatal fundus image classification result; S6. Use the training dataset constructed in step S2 to train the initial neonatal fundus image classification model constructed in step S5 to obtain a trained neonatal fundus image classification model; S7. Use the neonatal fundus image classification model obtained in step S6 to perform the classification of actual neonatal fundus images.
2. The method for classifying neonatal fundus images based on multi-modal data according to claim 1, wherein The specific steps of step S2 are as follows: Obtain neonatal fundus images: Among them, for each neonate, obtain several fundus images, label the category of the neonate, but do not label the category of each fundus image; Adjust the resolution of the obtained neonatal fundus images and perform normalization operations to construct a training dataset; Among the obtained neonatal fundus images, select several neonatal fundus images for text description annotation; The specific steps of step S3 are as follows: In the offline state, use the text base model to extract the text features of the neonatal fundus images with text annotation; At the same time, use the vision base model to extract the image features of the neonatal fundus images with text annotation.
3. The method for classifying neonatal fundus images based on multimodal data according to claim 2, wherein The text base model includes the Bert-base-Chinese model; the vision base model includes the ResNet-50 model or the RETFound model; when using the ResNet-50 model, the extracted image features are the input of the global average pooling operation in the ResNet-50 model; when using the RETFound model, the extracted image features are the input of the last layer of the transformer in the RETFound model.
4. The method for classifying neonatal fundus images based on multimodal data according to claim 3, wherein The said step S4 specifically includes the following steps: Use the text features and corresponding image features obtained in step S3 as the first training set; Process the image features sequentially through a pre-trained image encoder and a text feature generator to be trained to generate pseudo-text features; Process the text features through a pre-trained text encoder to generate real text features; During training, calculate the error loss between the pseudo-text features and the real text features to train the text feature generator.
5. The method for classifying neonatal fundus images based on multimodal data according to claim 4, wherein The said pre-trained image encoder includes a pre-trained RETFound model or a pre-trained ResNet-50 model; the said pre-trained text encoder includes a pre-trained Bert-base-Chinese model; the text feature generator is specifically a linear mapping layer; calculating the error loss between the pseudo-text features and the real text features is specifically calculating the mean square error loss between the pseudo-text features and the real text features.
6. The method for classifying neonatal fundus images based on multimodal data according to claim 5, wherein The said step S5 includes the following steps: Image prediction module: Based on a pre-trained image encoder, a multi-head self-attention mechanism, a linear mapping layer, and an attention weight calculation scheme, construct an image prediction module; the image prediction module is used to generate an image classification prediction result of the input image; Pseudo-text prediction module: Based on a pre-trained image encoder, the trained text feature generator obtained in step S4, an attention weight calculation scheme, and a linear layer, construct a pseudo-text prediction module; the input image first generates instance-level pseudo-text features through a pre-trained image encoder and the trained text feature generator obtained in step S4, and then processes the instance-level pseudo-text features based on the attention weight calculation scheme and the linear layer to generate a pseudo-text classification prediction result; Fusion module: Calculate the weight values of the image classification prediction result and the pseudo-text classification prediction result based on the information entropy calculation scheme, and combine the weighted summation scheme to fuse the prediction results, and finally obtain the classification result of the neonatal fundus image.
7. The method for classifying neonatal fundus images based on multi-modal data according to claim 6, wherein The processing process of the image prediction module specifically includes the following steps: The input image is processed through a pre-trained image encoder to obtain image dense features; Concatenate the obtained image dense features with a classification token to be learned, denoted as: where is the token group after adding position encoding; is the classification token; is the image patch token; is to and the concatenation operation; is the position encoding; Then, the obtained is processed through a linear mapping layer, information interaction is carried out through a multi-head self-attention mechanism, and then processed through a linear layer. Then, the result is summed with to obtain instance-level image features , which is expressed as: In the formula is the processing function of the linear layer; is the processing function of the linear mapping layer; is the processing function of the multi-head self-attention mechanism; The obtained instance-level image features , are processed through an attention weight calculation scheme and then through a linear layer to obtain the image classification prediction result, expressed as: In the formula is the image classification prediction result; is the processing process of the attention weight calculation scheme of the image prediction module, and , is the ii-th image instance-level feature encoding, is the weight value of and , w is the first parameter to be learned, and V is the second parameter to be learned.
8. The method for classifying neonatal fundus images based on multi-modal data according to claim 7, wherein The processing process of the pseudo-text prediction module specifically includes the following steps: The input image generates instance-level pseudo-text features through a pre-trained image encoder and the trained text feature generator obtained in step S4; The instance-level pseudo-text features are processed through an attention weight calculation scheme and then through a linear layer to obtain the pseudo-text classification prediction result, expressed as: In the formula is the pseudo-text classification prediction result; are the instance-level pseudo-text features; is the processing process of the attention weight calculation scheme of the pseudo-text prediction module, and , where is the k-th text instance-level feature encoding, is the weight value of and , is the third parameter to be learned, is the fourth parameter to be learned.
9. The method for classifying neonatal fundus images based on multimodal data according to claim 8, wherein The processing process of the fusion module specifically includes the following steps: The predicted result of image classification is calculated of the image information entropy is ; the predicted result of pseudo-text classification is calculated of the text information entropy is ; where is the information entropy calculation function; Obtain the image information entropy and the text information entropy to get the extreme value information entropy which is ; The image weight and the text weight are calculated using the following formula: In the formula is the image weight; is the text weight; Finally, the predicted result is calculated. For .
10. An imaging method, comprising the method for classifying neonatal fundus images based on multimodal data according to any one of claims 1 to 9, characterized in that It also includes the following steps: S8. Mark and re-image the classification result of the neonatal fundus image obtained in step S7 on the neonatal fundus image to obtain a neonatal fundus image with the classification result.
Citation Information
Patent Citations
Bimodal-fused forged information detection method based on prompt learning
CN118364421A
Newborn fundus image classification method based on multi-instance classification, imaging method and storage medium
CN119229513A