Vision-language continuous comparison pre-training method for multi-modal eye fundus image
By designing a continuous comparison pre-training framework, integrating multimodal fundus images and text features, alleviating catastrophic forgetting, and improving the learning ability and diagnostic efficiency of fundus image analysis.
Patent Information
- Application Number
- CN202510447982.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-22
AI Technical Summary
The existing pre-training methods lack the ability to integrate multimodal information in fundus image analysis, making them difficult to adapt to the multimodal data arrived in increments, and alleviate catastrophic misfortune.
A continuous contrast pre-training framework is designed to integrate multimodal fundus images and text features through comparative learning strategies, combine representative joint embedding playback strategies and non-diagonal information distillation losses to build a unified multimodal representation space to alleviate catastrophic forgetting.
It significantly improves the generalization ability and diagnostic efficiency of multimodal fundus image analysis, effectively alleviates catastrophic forgetting problems, and improves the stability and generalization performance of the model.
Smart Images

Figure CN120356035A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a vision - language contrast pre - training method for multi - modal fundus images, belonging to the field of continuous contrast pre - training. Background Art
[0002] Fundus imaging technology plays an important role in the early diagnosis and intervention of ophthalmic and systemic diseases. By analyzing fundus images, doctors can identify various disease signs and thus take timely treatment measures. However, traditional fundus image analysis methods usually design independent models for a single modality (such as color fundus photography, fluorescein fundus angiography, or optical coherence tomography). This mode limits the sharing of cross - modal knowledge and the generalization ability of the model. With the development of deep learning technology, base models pre - trained on large - scale datasets have shown great potential in fundus image analysis, being able to provide powerful initial weights for downstream tasks and significantly improving the efficiency of transfer learning. Nevertheless, most current base models are still limited to a single modality and lack the integration of multi - modal information, which restricts their performance in complex clinical tasks.
[0003] Fundus imaging technology covers multiple modalities, and each modality provides different perspectives on the retinal structure and diseases. For example, color fundus photography (CFP) can capture the overall morphology of the retinal surface, fluorescein fundus angiography (FFA) focuses on the dynamic changes of the vascular system, and optical coherence tomography (OCT) provides high - resolution cross - sectional images of the retina. Multi - modal base pre - training can integrate the information of these different modalities, thereby improving the diagnostic ability and generalization performance of the model. However, in a dynamic environment, it is often unrealistic to collect data of all modalities because data of different modalities usually arrive gradually. In addition, for new modalities, training a model from scratch is inefficient, and the model needs to be continuously updated and incrementally learned during the pre - training process.
[0004] During the continuous pre - training process, catastrophic forgetting (i.e., training in a new stage impairs previous knowledge) is a major challenge. Different from class - incremental learning, continuous pre - training lacks clear class concepts and requires the model to learn the general representations of each fundus modality and integrate them into a unified space. However, features in the later stage often disrupt the early representation space, weaken the generalization ability for previous modalities, and exacerbate the forgetting problem. Some existing methods (such as MedCoSS and CMDK) alleviate this problem through replay - based self - supervised learning, but these methods only rely on image data and ignore paired text descriptions. In fact, as a bridge connecting different imaging modalities, text can achieve semantic understanding aligned with images and further improve the generalization performance of the model. Especially in the medical field, text descriptions rich in expert knowledge can significantly improve the quality of pre - training.
[0005] Although continuous vision-language pre-training has been explored in the field of natural images (such as CTP and IncCLIP), these methods mainly target class-incremental tasks and are not directly applicable to the problem of medical modality increment. In fundus image analysis, how to design a continuous pre-training framework that can effectively integrate multi-modal information, alleviate catastrophic forgetting, and improve generalization performance has become an urgent problem that needs to be solved by those skilled in the art. Summary of the Invention
[0006] Technical Problem:
[0007] Aiming at the problems that existing pre-training methods lack the ability to integrate multi-modal information in fundus image analysis, are difficult to adapt to incrementally arriving multi-modal data, and are not good at alleviating catastrophic forgetting, the present invention proposes the first continuous contrast pre-training framework in the fundus field to achieve efficient representation learning and continuous knowledge update of multi-modal fundus images.
[0008] Technical Solution:
[0009] The object of the present invention can be achieved by the following technical solutions:
[0010] A vision-language continuous contrast pre-training method for multi-modal fundus images, which is used to incrementally integrate image and text features of different imaging modalities to construct a unified multi-modal representation space, including the following steps:
[0011] Step S1: Establish a continuous contrast pre-training framework (Continual Contrastive Pre-training) in the field of fundus images. Adopting a contrastive learning strategy, during the incremental input process of multi-modal data, make the representations of matching images and texts more consistent, construct a unified multi-modal representation space, and avoid catastrophic forgetting in subsequent stages. Specifically, the entire continuous contrast pre-training stage is divided into three stages, respectively labeled as stage1, stage2, and stage3. In the three-stage continuous vision-language pre-training, fundus image-text pairs from three fundus image modalities of color fundus photography (CFP), fluorescein fundus angiography (FFA), and optical coherence tomography (OCT) are used and to input into the model of the current stage and These modalities are introduced gradually to conform to the decreasing trend of their data distribution in real clinical practice.
[0012] Step S2: Design the Representative Joint Embedding Rehearsal strategy, construct a replay mechanism based on representative joint embedding. First, calculate the image-text joint embedding J of each sample pair, i.e., image-text. i Then, use the K-means sampling strategy to retain some representative image-text pairs from the previous stage and mix them with the modal data of the current stage. This design allows the model in the current stage to review and revise past knowledge during training.
[0013] Step S3: Incorporate the Off-Diagonal Information Distillation loss, save the model of the previous stage and calculate the similarity matrix and the distillation loss based on it. S t-1 and S t are the similarity matrices of stage t-1 and stage t respectively. The method perceives and maintains the alignment relationship between images and texts by distilling the off-diagonal information in the similarity matrices S t-1 and S t of the models in the two previous stages. This design introduces the distillation loss L OID to maintain the alignment of image and text features between different stages.
[0014] Beneficial effects:
[0015] Compared with the prior art, the present invention has the following advantages:
[0016] 1. The technical solution of the present invention proposes the first continuous vision-language pre-training framework in the field of fundus. By contrastive pre-training, it aligns the multi-modal fundus images and texts that will gradually arrive and integrates the multi-modal knowledge into a unified basic model, significantly improving the generalization ability and diagnostic efficiency of multi-modal fundus image analysis.
[0017] 2. The technical solution of the invention designs a replay strategy based on representative joint embedding. It calculates the image-text joint embedding of each stage through similarity-weighted representation and uses K-means sampling to select the most representative image-text pairs, effectively alleviating the catastrophic forgetting problem in continuous pre-training.
[0018] 3. The technical solution of the present invention introduces a non-diagonal information distillation strategy. By retaining the similarity matrix of the previous stage, it maintains the alignment relationship between images and texts in continuous contrastive learning, further improving the stability and generalization performance of the model.
[0019] 4. Experimental results show that the continuous contrast pre-training framework of the present invention performs optimally among all comparison methods, demonstrates excellent generalization ability and the lowest forgetting rate in downstream tasks, can effectively meet the analysis requirements of various fundus modalities at different training stages, and provides an efficient and robust solution for multi-modal representation learning of fundus images. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0021] Figure 1 It is a schematic diagram of jointly pre-training three types of fundus images, namely CFP, FFA, and OCT, in one stage.
[0022] Figure 2 It is a schematic diagram of the three fundus image modalities of CFP, FFA, and OCT incrementally arriving and continuously pre-training in three stages.
[0023] Figure 3 It is a schematic diagram of the overall framework of the present invention, including a continuous pre-training framework, as well as a designed representative joint embedding replay mechanism and non-diagonal information distillation strategy.
[0024] Figure 4 It is a schematic diagram of the continuous contrast pre-training framework proposed by the present invention, which is divided into four pre-training steps in total.
[0025] Figure 5 It is a flowchart of the representative joint embedding replay mechanism of the present invention.
[0026] Figure 6 It is a schematic diagram of the non-diagonal information distillation strategy of the present invention.
[0027] Figure 7 It is a flowchart of a vision-language continuous contrast pre-training method for multi-modal fundus images proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] The present invention will be described in detail below with reference to the drawings and specific embodiments. Note that the following description of the embodiments is only illustrative in nature, and the present invention is not intended to limit the objects or uses to which it is applied, nor is the present invention limited to the following embodiments.
[0029] For easy understanding, the following will explain the terms involved in this embodiment.
[0030] Fundus Image: A fundus image refers to an image of the posterior part of the eyeball (i.e., the fundus) taken by a specialized imaging device (such as a fundus camera), mainly used to observe the morphology and lesions of structures such as the retina, optic disc, macula, and blood vessels. Fundus images are important for the diagnosis of ophthalmic diseases (such as diabetic retinopathy, glaucoma, macular degeneration, etc.), and can also reflect the progression of systemic diseases (such as hypertension, diabetes, etc.). Common fundus imaging techniques include color fundus photography (CFP), fluorescein fundus angiography (FFA), and optical coherence tomography (OCT). Among them, color fundus photography (CFP) takes color images of the fundus through visible light, and can clearly show the overall morphology of the retina, blood vessel distribution, and lesion areas; fluorescein fundus angiography (FFA) injects sodium fluorescein dye and takes fundus images to observe the hemodynamic changes of retinal blood vessels; optical coherence tomography (OCT) generates cross-sectional images of each layer of the retina through the interference principle, and can accurately measure the thickness of the retina and observe the interlayer structure. Multi-modal fundus images refer to a collection of fundus images obtained by different imaging techniques, which can provide complementary diagnostic information and improve the comprehensiveness and accuracy of disease analysis.
[0031] Continual Pre-training: Continual pre-training is a training strategy in machine learning, aiming to continuously update and optimize the model by gradually introducing data of new modalities or new tasks to meet the multi-modal learning needs in a dynamic environment. Different from traditional pre-training, continual pre-training needs to solve the problem of catastrophic forgetting, that is, the training in the new stage should not damage the model's knowledge of previous modalities or tasks. Continual pre-training has been widely applied in fields such as medical image analysis and natural language processing.
[0032] Contrastive Pre-training: Contrastive pre-training is a self-supervised learning method that learns the general representation of data by maximizing the similarity of positive sample pairs (such as image-text pairs) and minimizing the similarity of negative sample pairs. In medical image analysis, contrastive pre-training can effectively integrate multi-modal information (such as images and texts) and improve the generalization ability and diagnostic performance of the model.
[0033] Catastrophic Forgetting: Catastrophic forgetting refers to the phenomenon that in the process of continuous learning or continuous pre-training, the model forgets previously learned knowledge when learning new knowledge. This phenomenon is particularly prominent in multi-modal learning and dynamic environments, leading to a significant decline in the performance of the model on previous modalities or tasks. Alleviating catastrophic forgetting is one of the core challenges in continuous pre-training.
[0034] Multi-modal Representation Learning: Multi-modal representation learning refers to learning general features that can represent multi-modal information simultaneously by integrating data of multiple modalities (such as images, text, audio, etc.). In medical image analysis, multi-modal representation learning can combine complementary information of different modalities to improve the diagnostic ability and robustness of the model.
[0035] Embodiment
[0036] As Figure 1 shown, in the case of joint learning, the model needs to simultaneously acquire and process data of three modalities (such as color fundus photography, fluorescein fundus angiography, and optical coherence tomography). This mode can achieve the optimal representation learning effect when multi-modal data is complete. However, in a dynamic medical clinical environment, multi-modal medical image data usually arrives incrementally, that is, data of different modalities is not acquired simultaneously but arrives gradually over time. In this case, if the model is retrained from scratch every time new modal data arrives, it will not only consume a large amount of time and computing resources but also may cause the model to be unable to effectively utilize the knowledge of previous modalities.
[0037] As Figure 2 shown, in the case of continual learning, the model incrementally pre-trains data of three different modalities in three stages. This mode can adapt to the scenario of dynamic data arrival but also faces a relatively serious problem of catastrophic forgetting, that is, the model trained on new modal data is prone to forgetting the modal knowledge learned in the previous stage, resulting in a significant decline in the performance of the model on previous modalities. Therefore, how to effectively integrate multi-modal information and alleviate catastrophic forgetting in the incremental learning process has become a key challenge in multi-modal representation learning in a dynamic medical clinical environment.
[0038] Figure 3 shows a schematic diagram of the overall framework of a vision-language continuous contrast pre-training method for multi-modal fundus images. The method includes the following steps:
[0039] Step S1: Establish the first continuous contrastive pre-training framework in the field of fundus images. Using a contrastive learning strategy, in the process of incremental input of multimodal data, the representation of matched images and texts is made more consistent, a unified multimodal representation space is constructed, and catastrophic forgetting in subsequent stages is avoided. Specifically, the entire continuous contrastive pre-training stage is divided into three stages, marked as stage1, stage2, and stage3. In the three-stage continuous visual-language pre-training, fundus image-text pairs from CFP, FFA, and OCT modalities are used respectively. and To input the model of the current stage and These modalities are introduced gradually to match their decreasing distribution in real clinical practice. Figure 4 shown.
[0040] In this embodiment, a stage0 is designed before the three stages of the continuous contrastive pre-training framework. In stage0, the text encoder is pre-trained using ophthalmology and general medical report text data through masked language modeling loss (MLM) to provide a shared semantic basis for different image modalities to better understand textual knowledge and provide rich semantic representation for image-text alignment in subsequent contrastive pre-training.
[0041] In this embodiment, the three stages of the continuous comparison pre-training framework pre-train the data of three common fundus image modalities, CFP, FFA and OCT, respectively, wherein the input model of each stage Each sample is a number of image-text pairs, where each sample is an image and its paired text description. The model includes a visual encoder based on ResNet50 and a text encoder based on BioClinicalBERT, which process image and text data respectively. Each stage independently learns a new modality representation, and in subsequent stages, only the weights of the previous stage model need to be loaded to continue incremental training, without reusing the modality data of the previous stage for training from scratch.
[0042] In this embodiment, the goal of contrastive pre-training is to bring the positive sample image-text pairs closer in the representation space, while pushing the negative sample pairs further away. Assume that there are N pairs of image-text pairs in each batch. After encoding and projection, the image feature representation and text feature representation are respectively recorded as: I = [I1, I2, ..., I N ] and T=[T1,T2,…,T N ]. The similarity matrix S is an N×N matrix, where each element S ijThe feature I of the i-th image i and the feature T of the j-th text j The cosine similarity between them is:
[0043]
[0044] where i, j ∈ 1, 2, …, N, and i and j represent the serial numbers of the image and the text respectively.
[0045] In this embodiment, an InfoNCE loss function similar to the CLIP method is adopted to train the model. S ii is the similarity of the positive sample pair, S ij and S ji are the similarities of the negative sample pairs, and τ is the temperature parameter. The formula is as follows:
[0046]
[0047] where is the contrastive loss.
[0048] Step S2: Design a Representative Joint Embedding Rehearsal strategy, construct a rehearsal mechanism based on representative joint embedding. First, calculate the image-text joint embedding representation J i of each sample pair, that is, the image-text pair. Then, use the K-means sampling strategy to retain some representative image-text pairs from the previous stage and mix them with the modal data of the current stage. This design allows the model of the current stage to review and revise past knowledge during the training process. The implementation process of the described representative joint embedding rehearsal mechanism is as Figure 5 shown.
[0049] In this embodiment, the representative judgment in the rehearsal mechanism of representative joint embedding is based on the image-text joint embedding J i of each sample pair. Specifically, the data of the previous stage is input into the model of the previous stage to obtain the image and text embedding representations. For the feature I i of the i-th image and the feature T i of the i-th text, the joint embedding J i is expressed as:
[0050] J i = S ii · I i + (1 - S ii ) · T i
[0051] where Sii is I i and T i is the similarity score between them. If the image and text are highly aligned, the joint embedding relies more on the features of the image I i ; if weakly aligned, the features of the text T i are enhanced in the joint embedding to ensure the integrity of the representation.
[0052] In this embodiment, after calculating the joint embedding representation J of the previous modality image-text pair, the replay mechanism of the representative joint embedding i adopts a k-means sampling strategy to select the most representative samples to construct a replay buffer, and the formula is as follows:
[0053]
[0054] where c k is the centroid of each cluster, and Q k represents the sample set in the k-th cluster. For each cluster, a subset of samples closest to the centroid is selected and added to the replay. K is the total number of clusters, and R is the final replay constructed for this stage.
[0055] In this embodiment, the replay mechanism of the representative joint embedding maintains a lean and fixed-size replay buffer throughout the continuous pre-training framework, in which all previous modality data is evenly distributed in the replay buffer. As new modalities are added, the buffer is dynamically updated during pre-training to prevent it from becoming too large, thus avoiding storage and computational burdens.
[0056] Step S3: Add off-diagonal information distillation loss, save the model of the previous stage and calculate the similarity matrix and distillation loss accordingly. S t-1 and S t are the similarity matrices of stage t-1 and stage t respectively. The method perceives and maintains the alignment relationship between the image and the text by distilling the off-diagonal information in the similarity matrices S t-1 and S t of the models in the two previous stages. This design introduces a distillation loss L ODID to maintain the alignment of the image and text features between different stages. The schematic diagram of the off-diagonal information distillation loss is as Figure 6 shown.
[0057] In this embodiment, before calculating the off-diagonal information distillation loss, it is necessary to correct the similarity matrix of the model in the previous stage. Specifically, S t-1and S t are the similarity matrices of the stage t-1 and the stage t respectively. For the frozen model of the previous stage where the similarity of the diagonal in the similarity matrix is not the largest, the model will misjudge the image-text alignment in the current stage. The method replaces the rows in S t with the corresponding values in S t-1 to avoid misleading updates:
[0058] If
[0059] In this embodiment, the non-diagonal information distillation loss distills the similarity distribution from the similarity matrix of the previous stage using KL divergence. This design explicitly constrains the update of the representation. The specific distillation formula is as follows:
[0060]
[0061] In this embodiment, the non-diagonal information distillation loss needs to be incorporated into the contrastive loss to obtain the total training loss function of the present invention as follows:
[0062]
[0063] Generally speaking, Figure 7 shows the overall flowchart of the training method of the present invention, including the above three steps of S1, S2 and S3.
[0064] In summary, the present invention provides an implementation method for the first multi-modal fundus image incremental learning framework based on continuous pre-training in the field of fundus. The core of this method is the representative joint embedding replay strategy and the non-diagonal information distillation module. The method of the present invention effectively alleviates the catastrophic forgetting problem by designing a memory replay strategy, and improves the performance of the model in the incremental learning scenario through the non-diagonal information distillation module. Experiments show that this method can significantly improve the continuous learning ability and generalization performance of the model, providing a new solution for incremental pre-training in a dynamic clinical environment. Therefore, this technology has high application value and promotion potential.
[0065] Specific application example:
[0066] For the construction of pre-training data, in a specific embodiment of the present invention, the pre-training data is constructed in multiple stages. In stage 0, 34 publicly available medical text datasets are used, covering the fields of ophthalmology and general medicine; in stages 1 to 3, image-text pairs are used for continuous vision-language pre-training. For the CFP modality, the MM-Retinal and Flair datasets are used, integrating 37 publicly available fundus datasets, covering 96 categories, with a total of 193,678 pairs of data; for the FFA modality, the MM-Retinal and FFA-IR subsets are combined, covering 46 retinal diseases, with a total of 701,098 pairs of data; for the OCT modality, 11 publicly available OCT datasets with category annotations and MM-Retinal are used, providing 183,817 pairs of data. All category labels are mapped to text prompts through prior knowledge.
[0067] For the evaluation data and metrics, in a specific embodiment of the present invention, the disease classification task is carried out on five multi-class fundus downstream datasets, covering three modalities of CFP, FFA, and OCT, and is evaluated under zero-shot, linear probing, and CLIP adapter settings. These datasets cover common retinal diseases such as diabetic retinopathy, glaucoma, and age-related macular degeneration. The results of the current stage are used to evaluate the plasticity of the model, and the results of the previous stage are used to reflect the degree of forgetting.
[0068] For the implementation details and comparison methods, in a specific embodiment of the present invention, ResNet-50 is used as the visual encoder, and the initialization parameters are from ImageNet1K, and BioClinicalBERT is used as the text encoder. The image resolution is adjusted to 512×512, and the length of the English text token is set to 256. The evaluation uses five-fold cross-validation, and the results are averaged. The training uses the AdamW optimizer, combined with a warm-up strategy, and is carried out on four NVIDIA A6000 GPUs with a batch size of 24. The comparison methods include some popular continual learning methods, covering SeqFT, LWF, EWC, ER, and ICARL, representing sequence fine-tuning, regularization methods, and replay methods respectively.
[0069] As shown in Table 1. Performance comparison on the CFP modality (FIVES and ODIR datasets). Training FFA and OCT in stages 2 and 3 results in forgetting of CFP. The table shows the downstream test results of CFP after all three stages, and the forgetting rates of stages 2 and 3 relative to stage 1 are shown in parentheses. Ours in the table is the method proposed in the present invention.
[0070] Table 1
[0071]
[0072]
[0073]
[0074] As shown in Table 2, the performance comparison of the FFA modality on the MPOS dataset. Training OCT in Stage 3 leads to forgetting of FFA. The table shows the downstream test results of Stages 2 and 3 on FFA, and the values in parentheses represent the forgetting rate relative to Stage 2. Ours in the table is the method proposed in the present invention.
[0075] Table 2
[0076]
[0077]
[0078] As shown in Table 3, the performance comparison (ACC and AUC, %) on the OCT modality. The data in the table are the averages under the zero-shot, linear probe, and ClipAdapter settings. Ours in the table is the method proposed in the present invention.
[0079] Table 3
[0080]
[0081] Performance analysis of the CFP modality: As shown in Table 1, the method proposed in the present invention shows the lowest forgetting rate and the best performance in all stages and metrics of the CFP modality. For example, in Stage 2 of the FIVES dataset, the zero-shot accuracy of the method proposed in the present invention only drops by 6.9%, significantly better than 31.4% of SeqFT; in Stage 3, the zero-shot accuracy of the method proposed in the present invention drops by 13.0%, while SeqFT drops by 37.0%. In addition, after training in Stage 2 of FFA and Stage 3 of OCT, the method proposed in the present invention still maintains a zero-shot accuracy of 45.9% in Stage 3, much higher than the performance of LWF and EWC. In the ODIR dataset, the method proposed in the present invention not only effectively alleviates forgetting (the zero-shot AUC in Stage 3 only drops by 12.3%), but also achieves performance improvement in some metrics (such as the ClipAdapter AUC in Stage 2 increases by 1.1%), fully verifying its superiority.
[0082] FFA Modal Performance Analysis: As shown in Table 2, in the experiments on the MPOS dataset (FFA modal), the method proposed in the present invention shows the lowest forgetting rate in Stage 3. The linear probing accuracy only drops by 3.2%, and the AUC only drops by 1.1%, significantly outperforming SeqFT and EWC. The method proposed in the present invention effectively balances cross-modal knowledge through the memory replay strategy and cross-modal information distillation, maintains stable performance on previous tasks while adapting to new modalities, fully demonstrating its effectiveness in the FFA modal.
[0083] OCT Modal Performance Analysis: As shown in Table 3, the method proposed in the present invention performs optimally on both the OCTID and OCTDL datasets in the OCT modal. In OCTID, the zero-shot accuracy and AUC of the method proposed in the present invention are 71.6% and 94.6% respectively; in OCTDL, they are 83.1% and 88.1% respectively, both superior to other methods. The experiments show that the method proposed in the present invention has better plasticity.
[0084] Those of ordinary skill in the art will realize that the embodiments described herein are for helping readers understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc., made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. A vision-language continuous contrast pre-training method for multimodal fundus images, characterized in that, To incrementally integrate image and text features of different imaging modalities to construct a unified multi-modal representation space, including the following steps: Step S1: Establish a continuous contrast pre-training framework. Adopt a contrastive learning strategy to make the representations of matching images and texts more consistent during the incremental input of multi-modal data, construct a unified multi-modal representation space, and avoid catastrophic forgetting in subsequent stages. The entire continuous contrast pre-training framework is divided into three stages, labeled as stage1, stage2, and stage3 respectively. In the three stages, fundus image-text pairs from three fundus image modalities, namely color fundus photography (CFP), fluorescein fundus angiography (FFA), and optical coherence tomography (OCT), are used and to input the model of the current stage and and these modalities are introduced step by step to conform to the decreasing trend of their data distribution in real clinical practice; Step S2: Construct a replay mechanism based on representative joint embedding. First, calculate the image-text joint embedding J of each sample pair, i.e., the image-text pair. i , and then use the K-means sampling strategy to retain some representative image-text pairs from the previous stage and mix them with the modality data of the current stage. Allow the model of the current stage to review and revise past knowledge during training. Step S3: Add non-diagonal information distillation loss and save the model of the previous stage and calculate the similarity matrix and distillation loss accordingly; where S t-1 and S t are the similarity matrices of stage t-1 and stage t respectively; the method perceives and maintains the alignment relationship between images and texts by distilling the non-diagonal information in the similarity matrices S t-1 and S t ; and introduce distillation loss L ODID to maintain the alignment of image and text features between different stages.
2. The vision-language continuous contrast pre-training method for multi-modal fundus images according to claim 1, wherein Before the three stages of the continuous contrast pre-training framework in step S1, a stage0 is designed. In stage0, the text encoder is pre-trained using ophthalmic and general medical report text data through masked language modeling loss to provide a shared semantic basis for different image modalities, so as to better understand text knowledge and provide rich semantic representations for image-text alignment in subsequent contrast pre-training.
3. The visual-language continuous contrast pre-training method for multimodal fundus images according to claim 1, wherein In step S1, the data of three fundus image modalities, namely CFP, FFA, and OCT, are continuously pre-trained in sequence in the three stages of the pre-training framework. In each stage, the input to the model is a number of image-text pairs, where each sample is an image and its paired text description; the model includes a vision encoder based on ResNet50 and a text encoder based on BioClinicalBERT, which process image and text data respectively; each stage independently learns a new modality representation, and in subsequent stages, only the weights of the previous stage model need to be loaded for incremental training, without having to retrain from scratch using the data of the previous stage modalities.
4. The visual-language continuous contrast pre-training method for multi-modal fundus images according to claim 1, wherein In step S1, the objective of the contrastive pre-training is to bring the image-text pairs of the positive sample pairs closer in the representation space while pushing the negative sample pairs farther away; assuming there are N pairs of image-text pairs in each batch; after encoding and projection, the image feature representation and the text feature representation are respectively denoted as: I = [I1, I2, …, I N and T = [T1, T2, …, T N ; the similarity matrix S is an N×N matrix, where each element S ij represents the cosine similarity between the feature I i of the i-th image and the feature T j of the j-th text: Where i, j ∈ 1, 2, …, N, and i and j represent the serial numbers of images and texts respectively.
5. The vision-language continuous contrast pre-training method for multi-modal fundus images according to claim 1, characterized in that In step S1, the model is trained using an InfoNCE loss function similar to the CLIP method, where S ii is the similarity of the positive sample pair, S ij and S ji are the similarities of the negative sample pairs, and τ is the temperature parameter. The formula is as follows: Among them, is the contrastive loss.
6. The vision-language continuous contrast pre-training method for multimodal fundus images according to claim 1, wherein, In step S2, the representative judgment in the replay mechanism of the representative joint embedding is based on the image-text joint embedding J of each sample pair i . Specifically, the data of the previous stage is input into the model of the previous stage to obtain the image and text embedding representations; for the feature I i of the i-th image and the feature T i of the i-th text, the joint embedding J i is expressed as: J i = S ii · I i + (1 - S ii )· T i Among them, S ii is the similarity score i between I i and T; if the image and text are highly aligned, the joint embedding relies more on the features I i of the image; if weakly aligned, the features T i of the text are enhanced in the joint embedding to ensure the integrity of the representation.
7. The vision-language continuous contrast pre-training method for multi-modal fundus images according to claim 1, wherein In step S2, after the replay mechanism of the representative joint embedding calculates the joint embedding J of the previous modality image-text pair, a k-means sampling strategy is adopted to select the most representative samples to construct the replay buffer, and the formula is as follows: i After that, a k-means sampling strategy is adopted to select the most representative samples to construct the replay buffer, and the formula is as follows: where, c k is the centroid of each cluster, and Q k represents the set of samples in the k-th cluster; for each cluster, select the subset of samples closest to the centroid and add them to the replay; K is the total number of clusters, and R is the final replay constructed for this stage.
8. The vision-language continuous contrast pre-training method for multimodal fundus images according to claim 1, wherein In step S3, the replay mechanism of representative joint embedding maintains a streamlined and fixed-size replay buffer throughout the continuous pre-training framework, in which all previous modality data is evenly distributed in the replay buffer; as new modalities are added, the buffer is dynamically updated during pre-training to prevent it from becoming too large, thus avoiding storage and computational burdens.
9. The visual-language continuous contrast pre-training method for multimodal fundus images according to claim 1, characterized in that Before calculating the non - diagonal information distillation loss in step S3, it is necessary to correct the similarity matrix of the previous - stage model; specifically, for the frozen model of the previous stage when the similarity of the diagonal in the similarity matrix is not the largest, the model will misjudge the image - text alignment in the current stage. Replace the rows in S t with the corresponding values in S t-1 to avoid misleading updates: If 10. The visual-language continuous contrast pre-training method for multimodal fundus images according to claim 1, characterized in that In step S3, the non-diagonal information distillation loss distills the similarity distribution from the similarity matrix of the previous stage using KL divergence to explicitly constrain the update of the representation. The specific distillation formula is as follows: Among them, is the non-diagonal information distillation loss; Incorporate the non - diagonal information distillation loss into the contrastive loss to obtain the total training loss function as follows: Among them, λ is the weight of the non - diagonal information distillation loss .
Citation Information
Cited By
Congenital heart disease perioperative period risk early warning method and system based on fundus color photo
CN120895240A
Continuous learning method of visual language model for image classification, equipment and medium
CN121746822A