A construction method of a macular degeneration classification system and a macular degeneration classification system

Through loose pairing technology and low-rank adaptive technology, multimodal data sets are expanded, combined with deep typical correlation analysis, the problems of scarcity of data and high computational complexity in AMD automated screening are solved, and efficient fusion and accurate diagnosis of multimodal information are achieved.

CN119904695BActive Publication Date: 2025-07-22BEIJING INSTITUTE OF PETROCHEMICAL TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411993405.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-07-22
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

The existing AMD automated screening method relies on high-quality labeled data sets, and the data is scarce and difficult to obtain. The single-modal model has incomplete feature capture when diagnosing complex lesions, the multimodal data fusion calculation is high, and modal heterogeneity increases the difficulty of information fusion.

Method used

The multimodal data set is expanded by loose pairing technology, combined with low-rank adaptive technology and deep typical correlation analysis, features are extracted through the ViT-large encoder, and efficient fusion of multimodal information is achieved through the feature fusion module and the classifier to reduce the computational burden.

Benefits of technology

It improves the accuracy and efficiency of AMD diagnosis, reduces the computational complexity, ensures the comprehensive capture and fusion of multimodal features, and improves the recognition ability of complex lesions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119904695B_ABST
    Figure CN119904695B_ABST
Patent Text Reader

Abstract

The present invention provides a method for constructing a macular degeneration classification system and a macular degeneration classification system. The method includes: using a loose pairing technique to pair the acquired multimodal images with the same AMD degeneration labels to expand the multimodal AMD image dataset; using a feature extraction module with a fused low-rank adaptive technique to extract features from the OCT-CFP training sample pairs to obtain multimodal feature vectors; and implementing the maximization of the correlation of multimodal features in a shared space based on a deep canonical correlation analysis module of a neural network; fusing the multimodal features based on a feature fusion module to output target fusion features for classifier training, and training to obtain a target classifier with a minimized loss function; integrating the loose pairing technique, the feature extraction module, the deep canonical correlation analysis module, the feature fusion module, and the target classifier to obtain a macular degeneration classification system. Through the present invention, computing resources are saved, and the degeneration classification efficiency and accuracy are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image classification, and particularly relates to a method for constructing a macular degeneration classification system and a macular degeneration classification system. Background Art

[0002] Age-related macular degeneration (AMD) is a retinal degenerative disease with a high incidence in the elderly population and is also one of the main causes of irreversible vision loss globally. According to pathological characteristics and disease progression, AMD can be further divided into dry AMD, wet AMD, and polypoidal choroidal vasculopathy (PCV). Given the differences in treatment methods among these subtypes, accurately distinguishing normal retina from different AMD subtypes has important clinical significance for formulating personalized treatment strategies and improving patient prognosis. In the diagnosis of AMD, color fundus photography (CFP) and optical coherence tomography (OCT) are two commonly used non-invasive fundus imaging techniques. CFP can provide high-resolution color images of the retinal surface, which helps detect surface abnormalities, but has limited ability to identify changes in the deep retinal structure; while OCT can detail the layered structure of the retina and reveal deep lesion characteristics, but is relatively insufficient in judging surface color and texture. However, when multi-modal information (CFP / OCT images and clinical records) is provided, experts make fewer diagnostic errors when recommending referrals. Due to the lack of experienced ophthalmologists, more and more research has been dedicated to automated screening methods for AMD based on CFP images, OCT images, or a combination of both.

[0003] Problems existing in the prior art are as follows:

[0004] ① The accuracy and robustness of most AMD automated screening methods rely on the support of a large number of high-quality labeled datasets. In the field of ophthalmology, due to high labeling costs and privacy issues, such datasets are scarce and difficult to obtain, which poses a huge challenge to the development and optimization of automated screening systems.

[0005] ②In existing research, using the weights pre-trained on ImageNet as the initial point of the model for transfer learning has gradually become a common strategy. Recently, a model called RETFound has received extensive attention. As a general retinal-based model based on VisionTransformer (ViT), this model learns rich visual feature representations through pre-training on a large-scale unlabeled retinal image dataset, and only requires limited labeled data for efficient fine-tuning in subsequent tasks. It is worth noting that this model performs excellently in identifying complex patterns and features related to eye health in the single-modal setting. However, as a single-modal model, the structure of RETFound can only process single-modal data, which to some extent limits its comprehensive capture of retinal lesion features. Especially when diagnosing complex lesions such as AMD, relying solely on single-modal data may lead to the loss of some key pathological information. In contrast, multi-modal data fusion can combine the advantages of different modalities and provide a more comprehensive and detailed description of the lesion. Therefore, integrating multi-modal information into the RETFound model during the downstream task fine-tuning process to improve its recognition ability and diagnostic accuracy for complex AMD lesions is an important direction worthy of exploration.

[0006] ③In practical applications, the training and fine-tuning processes of multi-modal models face the dual challenges of computational complexity and modal fusion. As the number of modalities increases, the number of model parameters will increase exponentially, resulting in a significant increase in computational overhead. This complexity is particularly evident when integrating CFP and OCT images for AMD diagnosis.

[0007] ④The heterogeneity between different modal features further increases the difficulty of effective information fusion. How to fully utilize the complementary advantages of each modality while ensuring computational efficiency remains a key issue. Summary of the Invention

[0008] To this end, the present invention provides a method for constructing a macular degeneration classification system and a macular degeneration classification system, aiming to solve the technical problems in the existing AMD automated screening methods, such as difficult acquisition of training data, insufficient recognition ability, low accuracy, and inability to ensure computational efficiency.

[0009] To achieve the above objectives, the present invention adopts the following technical solutions:

[0010] According to the first aspect of the present invention, the present invention provides a method for constructing a macular degeneration classification system, the method comprising:

[0011] Obtain multiple training samples carrying different AMD degeneration labels to form a multi-modal AMD image dataset; wherein, each of the training samples includes an OCT image and a CFP image; the AMD degeneration labels include normal, dry, wet, and polypoidal choroidal vasculopathy;

[0012] Use the loose pairing technique to pair the OCT images and CFP images in the multi-modal AMD image dataset with the same AMD degeneration label to obtain multiple groups of OCT-CFP training sample pairs;

[0013] Construct a feature extraction module integrating the low-rank and adaptive technique, and use the feature extraction module to extract features from the OCT-CFP training sample pairs to obtain an OCT high-dimensional feature vector and a CFP high-dimensional feature vector;

[0014] Construct a deep canonical correlation analysis module based on a neural network, and perform non-linear feature transformation on the OCT high-dimensional feature vector and the CFP high-dimensional feature vector respectively to obtain an OCT target feature vector and a CFP target feature vector with maximized correlation in the shared space;

[0015] Construct a multi-modal based feature fusion module to fuse the OCT target feature vector and the CFP target feature vector, and output a target fusion feature;

[0016] Construct a classifier based on a preset classification algorithm, use the target fusion feature as the input of the classifier, take the output sub-type category as the prediction result, and use the AMD degeneration label as the actual result to train and verify the classifier to obtain a target classifier with minimized loss function;

[0017] Integrate the loose pairing technique, the feature extraction module, the deep canonical correlation analysis module, the feature fusion module, and the target classifier to obtain a macular degeneration classification system for AMD degeneration classification.

[0018] Further, before using the loose pairing technique to pair the OCT images and CFP images in the multi-modal AMD image dataset with the same AMD sub-type label, the method further includes:

[0019] Perform image preprocessing on the multi-modal image dataset to obtain a preprocessed target image dataset; the target image dataset includes multiple training samples carrying different AMD degeneration labels; each training sample includes a preprocessed OCT image and a CFP image; wherein, the image preprocessing includes the following steps:

[0020] S1: Delete the abnormal image data in the multi-modal image dataset to obtain a first image dataset;

[0021] S2: Perform image normalization processing and image weighted enhancement processing on the first image data in the first image dataset to obtain a second image dataset;

[0022] S3: Perform intermediate image enhancement processing on the second image data in the second image dataset.

[0023] Further, the image normalization processing includes at least one of normalizing the channel value range of image pixels, equalizing the mean value of image pixels, standardizing the image pixels, and uniformly adjusting the image size;

[0024] and / or,

[0025] The processing formula for image weighted enhancement is expressed as:

[0026] I weight = I org *α + I blur *β + γ

[0027] I blur = I org *kernal h×w

[0028] where, I weight is the second image data; I org is the first image data; I blur is the image of the first image data after Gaussian blur, kernal h×w represents a Gaussian kernel with a size of h×w; α and β are weighting coefficients; γ is a constant;

[0029] and / or,

[0030] The image enhancement includes at least one of image rotation, image translation, random flipping, applying contrast-limited adaptive histogram equalization to CFP images, and applying median filtering method to OCT images.

[0031] Further, using the loose pairing technology to pair the OCT images and CFP images in the multimodal AMD image dataset with the same AMD degeneration label to obtain multiple groups of OCT-CFP training sample pairs, including:

[0032] For each training sample i in the target image dataset I aug , take the CFP image (I aug _f i ) and the OCT image (I aug _o i ) in it as the original pair;

[0033] Initialize an empty list of loose pairs;

[0034] For other training samples j (j ≠ i) in the target image dataset, if the AMD degeneration labels of the training sample i and the training sample j are the same, then add (I aug _f i , I aug _o j ) and (I aug _f j , I aug _o i ) to the loose pair list;

[0035] Merge the original pairs and the loose pair list to obtain a multi-modal AMD image set that expands multiple groups of OCT-CFP training sample pairs.

[0036] Furthermore, the feature extraction module that constructs the fusion low-rank adaptive technology is used to extract features from the OCT-CFP training sample pairs to obtain OCT high-dimensional feature vectors and CFP high-dimensional feature vectors, including:

[0037] Construct a feature extraction module that includes two independent ViT-large encoders; the ViT-large encoder includes a patch embedding layer, 24 Transformer blocks, and a 1024-dimensional embedding vector;

[0038] Input the CFP image (X cfp ∈R H×W×C ) and the OCT image (X oct ∈R H ×W×C ) in the OCT-CFP training sample pair into two independent ViT-large encoders respectively;

[0039] Divide them into patches of a fixed size through the patch embedding layer and map them to 1024-dimensional embedding vectors, denoted as P cfp ∈R 1×1024 and P oct ∈R 1×1024 ;

[0040] Process the embedding vectors using positional encoding to generate corresponding input feature vectors, denoted as Z 0 cfp ∈R 1×1024 and Z 0 oct ∈R 1×1024 ;

[0041] Input the input feature vector into the Transformer block pair for processing, and the output formula is expressed as:

[0042]

[0043] After being processed by 24 Transformer blocks, the high-dimensional feature vector Z to be transformed is obtained cfp ∈R 1×1024 and Z oct ∈R 1 ×1024 .

[0044] The method further includes:

[0045] Utilize the low-rank adaptation technology to perform low-rank decomposition on the query matrix WQ and the value projection matrix WV in the self-attention layer of the Transformer block to optimize the weight update, and the formula is expressed as follows:

[0046] ΔW = BA

[0047] where B ∈ R d×r and A ∈ R r×d are matrices generated by low-rank decomposition, and the rank r is much smaller than the dimension d;

[0048] Utilize the Transformer block with optimized weight update to process the input feature vector to obtain the OCT high-dimensional feature vector h oct ∈R 1×1024 and the CFP high-dimensional feature vector h cfp ∈R 1×1024 , and the formula is expressed as follows:

[0049] h = W0Z + ΔWZ = W0Z + BAZ

[0050] where W0 is the initial pre-trained weight; ΔW is the weight update term after low-rank decomposition; Z is the high-dimensional feature vector to be transformed.

[0051] Furthermore, build a deep canonical correlation analysis module based on a neural network, and perform non-linear feature transformation on the OCT high-dimensional feature vector and the CFP high-dimensional feature vector respectively to obtain the OCT target feature vector and the CFP target feature vector with maximized correlation in the shared space;

[0052] Build a deep canonical correlation analysis module including two independent deep neural networks, and input the OCT high-dimensional feature vector and the CFP high-dimensional feature vector into the two independent deep neural networks respectively for multiple non-linear transformations to generate the output feature O cfp ∈R N×d and O oct ∈R N×d, the formula is as follows:

[0053] O cfp = f cfp (h cfp , W1), O oct = f oct (h oct , W2)

[0054] Among them, are the CFP instance feature matrix and the OCT instance feature matrix respectively; N is the number of instances; d1 is the feature dimension of the CFP modality; d2 is the feature dimension of the OCT modality; W1 and W2 are the parameters for non-linear transformation by two deep neural networks respectively; d is the output dimension;

[0055] Center the output features to construct a cross-modal covariance matrix:

[0056]

[0057] Among them, Σ 11 and Σ 22 represent the covariance matrix of the CFP modality and the covariance matrix of the OCT modality respectively; Σ 12 is the cross-modal covariance matrix;

[0058] Maximize the correlation between the output features O cfp and O oct in the shared space to obtain the OCT target feature vector and the CFP target feature vector. The formula is as follows:

[0059]

[0060] Furthermore, construct a multi-modal based feature fusion module to fuse the OCT target feature vector and the CFP target feature vector, and output the target fusion feature, including:

[0061] Use the OCT target feature vector and the CFP target feature vector as input features, and initialize the moving average variable y avg as a zero matrix, with the same size as the input features;

[0062] Update the moving average variable y avg , the formula is as follows:

[0063] y avg = β·y avg + (1 - β)·y t

[0064] Among them, β is the moving average decay coefficient; y t is the input feature at the current moment; yt is the result of element-wise multiplication of the OCT target feature vector and the CFP target feature vector at time t;

[0065] Output the fused feature y output = y avg ;

[0066] Take y output as the input for the next time step t+1;

[0067] Select a specific time step t * as the termination and output the final target fused feature.

[0068] Furthermore, the preset classification algorithm includes at least one of a support vector machine, a random forest, a K-nearest neighbor algorithm, and a multi-layer perceptron;

[0069] The loss function formula is expressed as:

[0070]

[0071] where is the classification loss function; is the deep canonical correlation loss function; N is the number of training samples; C is the number of classes; y i,j is the actual result of training sample i; is the probability that the prediction result of training sample i is class j; λ is the loss weight used to balance the influence of the classification loss and the deep canonical correlation loss on the classifier training.

[0072] According to the second aspect of the present invention, the present invention provides a macular degeneration classification system constructed by the construction method of the macular degeneration classification system according to any one of the first aspects of the present invention. The system includes: a loose pairing module, a feature extraction module, a deep canonical correlation analysis module, a feature fusion module, and a target classifier;

[0073] The loose pairing module is used to obtain the OCT image and the CFP image to be classified, and use the loose pairing technology to pair the OCT image and the CFP image with the same AMD degeneration label to obtain multiple groups of OCT-CFP training sample pairs;

[0074] The feature extraction module is used to extract features from the OCT-CFP training sample pairs to obtain the OCT high-dimensional feature vector and the CFP high-dimensional feature vector after feature transformation based on the low-rank adaptation technology;

[0075] The deep canonical correlation analysis module is used to perform non-linear feature transformation on the OCT high-dimensional feature vector and the CFP high-dimensional feature vector to obtain the OCT target feature vector and the CFP target feature vector with maximized correlation in the shared space;

[0076] The feature fusion module is used to perform feature fusion on the OCT target feature vector and the CFP target feature vector to obtain a target fusion feature;

[0077] The target classifier is used to take the target fusion feature as an input and output the AMD degeneration category corresponding to the OCT image and the CFP image to be classified;

[0078] Wherein, the AMD degeneration category includes normal, dry, wet, and polypoidal choroidal vasculopathy.

[0079] The present invention adopts the above technical solutions and at least has the following beneficial effects:

[0080] Through the solution of the present invention, a large pre-trained single-modal retinal basic model is finely tuned to realize multi-modal AMD recognition. The present invention not only focuses on how to effectively expand the model to integrate multi-modal data during the fine-tuning process, but also focuses on reducing the computational cost and optimizing the information fusion between modalities. Considering the significant increase in computational complexity during the integration of multi-modal data, the present invention adopts the Low-Rank Adaptation (LoRA) method. By introducing the adaptation mechanism of low-rank matrices, the full-parameter matrix is decomposed into the product of two low-rank matrices, so that only part of the model parameters are optimized and adjusted during fine-tuning. This method significantly reduces the computational burden and memory requirements, while ensuring that the performance and accuracy of the model are not affected when processing multi-modal data. In order to bridge the heterogeneity of different modal features, the present invention also introduces the Deep Canonical Correlation Analysis (DCCA) method. Through two independent neural networks, the fine-grained cross-modal non-linear feature transformation between the CFP and OCT modal features is performed, and the classical Canonical Correlation Analysis (CCA) is combined for regularization to maximize the linear correlation between the two modal features. During the training process, the features of the two modalities are mapped into a coordinated high-dimensional hyperspace, enabling the model to learn highly correlated multi-modal features. Additionally, it should be noted that the present invention completely retains the original network structure of the large pre-trained single-modal retinal basic model, thereby ensuring that the feature processing ability in the single-modal case is not affected.

[0081] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. Brief Description of the Drawings

[0082] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0083] Figure 1 The flowchart showing the construction method of the macular degeneration classification system provided by an embodiment of the present invention is shown;

[0084] Figure 2 The module processing flowchart showing the construction method of the macular degeneration classification system provided by an embodiment of the present invention is shown;

[0085] Figure 3 The flowchart showing the image preprocessing provided by an embodiment of the present invention is shown;

[0086] Figure 4 The network structure diagram showing the low-rank adaptive technology provided by an embodiment of the present invention is shown;

[0087] Figure 5 The network structure diagram showing the deep canonical correlation analysis module provided by an embodiment of the present invention is shown;

[0088] Figure 6 The structure diagram showing the macular degeneration classification system provided by an embodiment of the present invention is shown. Detailed implementation manners

[0089] The following will describe the exemplary embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.

[0090] It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.

[0091] An embodiment of the present invention provides a method for constructing a macular degeneration classification system, as Figure 1 shown, which may at least include the following steps S101 to S107:

[0092] Step S101, obtaining a plurality of training samples carrying different AMD degeneration labels to form a multi-modal AMD image data set.

[0093] The multi-modal AMD image data set obtained in the embodiment of the present invention includes a subset corresponding to normal and different age-related macular degeneration (AMD) subclass disease labels, and each training sample in each subset includes a color fundus photograph (CFP) and an optical coherence tomography (OCT). That is to say, each training sample includes an OCT image and a CFP image. The AMD degeneration labels in the embodiment of the present invention include normal, dry, wet, and polypoidal choroidal vasculopathy.

[0094] In the research process of the present invention, in order to construct a multi-modal AMD image dataset, first, 1094 color fundus photography (CFP) images and 1289 optical coherence tomography (OCT) images were collected from the publicly available medical imaging dataset MMC-AMD. These images cover normal retinas and different subtypes of AMD from 1093 eyes, including dry AMD, wet AMD, and polypoidal choroidal vasculopathy (PCV). Each image was carefully annotated by professional ophthalmologists (i.e., AMD degeneration labels), ensuring accurate classification of lesion types. During the construction of the dataset, since CFP images are three-channel color images and OCT images are single-channel grayscale images, in order to ensure correct pairing and alignment of the two-modal images when input into the model, the embodiments of the present invention can use the method of channel replication to expand the single-channel OCT images into three channels, ensuring consistency in the channel dimension between CFP and OCT images. In addition, in order to reduce the impact of class imbalance on the classification results, resampling can also be performed on the categories with fewer samples.

[0095] As Figure 2 shown, it is a module processing flow chart of the construction method of the macular degeneration classification system. After obtaining a multi-modal AMD image dataset containing two-modal image data (OCT images / CFP images), image preprocessing can also be performed on the multi-modal image dataset to obtain a preprocessed target image dataset. As Figure 3 shown, it is a flow chart of image preprocessing in the embodiments of the present invention.

[0096] Specifically, image preprocessing may include the following steps:

[0097] Step S1: Delete abnormal image data in the multi-modal image dataset to obtain a first image dataset;

[0098] Step S2: Perform image normalization processing and image weighted enhancement processing on the first image data in the first image dataset to obtain a second image dataset;

[0099] Step S3: Perform image enhancement processing on the second image data in the second image dataset.

[0100] After the embodiments of the present invention perform operations such as abnormal data cleaning on the two-modal image data, data preprocessing steps including image normalization, image weighted enhancement, and image enhancement are executed to improve the image quality and meet the requirements of model input. During this process, the original image input is I org , and after image normalization and image weighted enhancement, image I weight is obtained, and finally, image enhancement operation is performed to obtain the image output as I aug。

[0101] Among them, the image normalization process includes normalizing the channel value range of image pixels, equalizing the image pixels, standardizing the image pixels, and uniformly adjusting the image size. Specifically, normalizing the channel value range of image pixels can be understood as performing range scaling on each channel (R, G, B) to ensure that the value ranges of all channels are consistent; the equalization process can make the image more consistent by calculating the average value of image pixels and subtracting this average value from each pixel in the image; the standardization process can reduce the variation range of image data and make the data more stable by calculating the standard deviation of image pixels and dividing each pixel in the image by this standard deviation; the size adjustment means adjusting the sizes of the two-modal images to be the same to ensure data consistency and meet the input requirements of the model. For example, the sizes are both adjusted to 224*224.

[0102] Image weighted enhancement refers to performing Gaussian Blurring on the original image I org to reduce the high-frequency detail information of the image, and weighting the original image I org and the image I blur after the Gaussian Blurring process to generate the enhanced image I weight . The processing formula can be expressed as:

[0103] I weight = I org *α + I blur *β + γ

[0104] The Gaussian Blurring process refers to performing a convolution operation on the original image I org using a Gaussian kernel to generate a blurred image I blur . The formula is expressed as follows:

[0105] I blur = I org *kernal h×w

[0106] Among them, I weight is the second image data; I org is the first image data; I blur is the image of the first image data after Gaussian Blurring, and kernal h×w represents a Gaussian kernel with a size of h×w; α and β are weighting coefficients; γ is a constant.

[0107] Image enhancement may include image rotation, image translation, random flipping, applying contrast-limited adaptive histogram equalization (CLAHE) to CFP images, and applying median filtering method to OCT images, etc. Specifically, it may include the following steps: rotating the image by 45° or 90°; randomly translating the image along the horizontal (x-axis) or vertical (y-axis) direction, and the translation distance range is between 0 and 10% of the image width or height; random flipping; applying contrast-limited adaptive histogram equalization (CLAHE) to CFP images to enhance the low-contrast regions in the images; using the median filtering method with a 3×3 kernel for OCT images, removing the salt-and-pepper noise in the images by replacing the pixel values with the median of the neighboring pixel values while preserving the edge information of the images.

[0108] Step S102: Using the loose pairing technique, pair the OCT images and CFP images in the multimodal AMD image dataset with the same AMD degeneration label to obtain multiple groups of OCT-CFP training sample pairs.

[0109] To obtain more training instances, the natural strategy is to strictly select CFP and OCT images from the same eye according to the disease label. Assume there are two eyes, eye a and eye b in the training set, that is, set and set where f and o represent CFP images and OCT images respectively. As can be seen from the above, there are only five groups with strict pairing. To increase the number of multimodal instances for training, the embodiment of the present invention constructs input pairs based on the disease category label rather than the same eye. That is, as long as the disease category labels are the same, the CFP image can be paired with the OCT image to construct an input pair. Based on the loose pairing method, the original five groups of available training instances for the two eyes, eye a and eye b are expanded to ten groups of available training instances: That is, as long as the disease category labels of the two-modal images are the same, they can be paired together to form loose image pairs. Through the loose pairing technique of the embodiment of the present invention, more training instances are constructed for the category label of the patient, effectively expanding the original training data. The specific implementation process of the loose pairing technique includes the following steps:

[0110] Step S1: Take I aug as the dataset input, and each sample in I aug contains a CFP image and an OCT image; there is a true AMD subtype label corresponding to each sample i;

[0111] Step S2: Denote the AMD subtype label of sample i as "Pi ", where P i is one of the four AMD subtype labels {"normal", "dryAMD", "wetAMD", "pcv"}; establish a disease category dictionary to map the AMD subtype label of each sample to the corresponding multimodal image data;

[0112] Step S3: For each training sample i in the target image dataset I aug , take the CFP image (I aug _f i ) and the OCT image (I aug _o i ) as the original pair;

[0113] Step S4: Initialize an empty loose pair list (LoosePairs = []);

[0114] Step S5: Traverse the target image dataset I aug , for other training samples j (j ≠ i) in it, if the AMD degeneration labels of training sample i and training sample j are the same, then construct a new multimodal image pair and add (I aug _f i , I aug _o j ) and (I aug _f j , I aug _o i ) to the loose pair list;

[0115] Step S6: Combine the original pair and the loose pair list to obtain an augmented multimodal AMD image set of multiple OCT-CFP training sample pairs.

[0116] Step S103, construct a feature extraction module using the fused low-rank adaptive technology, and use the feature extraction module to extract features from the OCT-CFP training sample pairs to obtain the OCT high-dimensional feature vector and the CFP high-dimensional feature vector.

[0117] The feature extraction module in the embodiments of the present invention is based on the ViT-large encoder in the RETFound model. This model is trained on a large-scale unlabeled retinal image through self-supervised learning and is applicable to the detection tasks of ocular and systemic diseases with explicit labels. The model adopts a self-encoder architecture with a specific configuration, including an encoder and a decoder. Among them, the encoder uses the ViT-large structure, including 24 Transformer blocks and a 1024-dimensional embedding vector. This encoder processes unmasked 16×16 pixel image patches and projects these image patches into 1024-dimensional feature vectors. The Transformer blocks generate high-level features through the multi-head self-attention mechanism and the multi-layer perceptron. The decoder part is the ViT-small structure, which reconstructs the masked image patches through 8 Transformer blocks and a 512-dimensional embedding vector. In the embodiments of the present invention, the decoder part is discarded, and only the ViT-large encoder is retained for feature extraction. The CFP image and the OCT image are processed through two independent ViT-large encoders respectively, so as to ensure the independent extraction of features of each modality. Specifically, let the input of the CFP image be X cfp ∈R H×W×C , and the input of the OCT image be X oct ∈R H×W×C . Each image first passes through the patch embedding layer, which divides it into patches of a fixed size, and then undergoes position encoding processing to generate the corresponding input feature vectors, denoted as P cfp ∈R 1×1024 and P oct ∈R 1×1024 . Each ViT-large encoder is processed through a series of Transformer blocks, and each block contains a multi-head self-attention mechanism and a feed-forward neural network, and finally two high-dimensional feature vectors Z cfp ∈R 1×1024 and Z oct ∈R 1×1024 are obtained. These feature vectors represent the representations of the images after deep encoding and capture the important information in different modalities. It should be noted that although the CFP and OCT images are processed in independent encoders, both follow the same feature extraction process, so as to ensure the consistency of cross-modal features for multi-modal interaction in the subsequent feature fusion stage. In summary, the feature extraction module in the embodiments of the present invention includes two independent ViT-large encoders, and each ViT-large encoder includes a patch embedding layer, 24 Transformer blocks, and a 1024-dimensional embedding vector. The specific implementation process of the feature extraction module includes the following steps:

[0118] Step S1: Input the CFP images (X cfp ∈R 224×224×3 ) and OCT images (X oct ∈R 224×224×3 ) in the OCT-CFP training sample pairs into two independent ViT-large encoders respectively;

[0119] Step S2: Divide them into patches of a fixed size (16×16) through the patch embedding layer, and then map each patch into an embedding vector of 1024 dimensions, denoted as P cfp ∈R 1×1024 and P oct ∈R 1×1024 respectively;

[0120] Step S3: Use the position encoding E pos to process the embedding vectors to generate corresponding input feature vectors, denoted as Z 0 cfp ∈R 1×1024 and Z 0 oct ∈R 1×1024 respectively;

[0121] Step S4: Input the input feature vectors into the Transformer blocks for processing. Each block contains a multi-head attention mechanism and a feed-forward neural network, and the output formula is expressed as:

[0122]

[0123] After being processed by 24 Transformer blocks, the high-dimensional feature vectors Z cfp ∈R 1×1024 and Z oct ∈R 1 ×1024 are obtained.

[0124] Furthermore, in the low-rank adaptation process, the high-dimensional feature vectors to be transformed can be further transformed to reduce the computational load during the model fine-tuning stage. In the embodiments of the present invention, in order to cope with the high computational requirements for cross-modal feature fusion of CFP and OCT images during fine-tuning, the low-rank adaptation (LoRA) technology is adopted to finely adjust the model parameters. As Figure 4 shown, it is a schematic diagram of the network structure of the low-rank adaptation technology, and its core idea is to optimize the self-attention layer of the Transformer through low-rank decomposition to effectively control the computational complexity.

[0125] Specifically, the input CFP image and OCT image are converted into input feature vectors Z after being processed by patch embedding and positional encoding. 0 cfp ∈R 1×1024 and Z 0 oct ∈R 1×1024 These vectors are then fed into their respective ViT-large encoders, each of which consists of multiple Transformer blocks containing the query matrix WQ, key matrix WK, and value projection matrix WV of the multi-head self-attention mechanism. The query matrix WQ and value projection matrix WV in the self-attention layer of the Transformer block are factorized into low ranks using the low-rank adaptation technique to optimize weight updates, and the formula is as follows:

[0126] ΔW = BA

[0127] where B ∈ R d×r and A ∈ R r×d are matrices generated by low-rank factorization, and the rank r is much smaller than the dimension d, aiming to reduce the computational load in the model fine-tuning stage.

[0128] After LoRA processing, the feature vectors h cfp ∈R 1×1024 and h oct ∈R 1×1024 of the CFP modality and OCT modality are generated by the following formulas respectively:

[0129] h = W0Z + ΔWZ = W0Z + BAZ

[0130] where W0 is the initial pre-trained weight; ΔW is the weight update term after low-rank factorization; Z is the high-dimensional feature vector to be transformed before low-rank adaptive fine-tuning.

[0131] It can be understood that by independently processing and updating the features of the CFP image and OCT image in this way, it ensures efficient and accurate feature representation of multi-modal data before fusion. In addition, by setting different ranks of the low-rank matrices, the balance between model complexity and computational efficiency can be adjusted. During model fine-tuning, only these low-rank matrices are adjusted, and other parameters remain unchanged. This strategy can effectively reduce training time and computational resources while maintaining model performance.

[0132] Step S104, construct a deep canonical correlation analysis module based on a neural network, and perform non-linear feature transformation on the OCT high-dimensional feature vector and CFP high-dimensional feature vector respectively to obtain the OCT target feature vector and CFP target feature vector with maximized correlation in the shared space;

[0133] In the embodiments of the present invention, deep canonical correlation analysis (DCCA) is introduced into the multi-modal AMD classification task, aiming to address the challenges of heterogeneity and non-linear feature alignment and correlation modeling for CFP and OCT modality data. The core idea of DCCA is to perform non-linear mapping on the input features of each modality through a deep neural network, transform them into a shared latent space, and then enhance the synergistic effect between modalities by maximizing the correlation of different modality features in this space, thereby improving the effect of multi-modal feature fusion. As Figure 5 shown, it is a schematic diagram of the network structure of the deep canonical correlation analysis module.

[0134] Specifically, the deep canonical correlation analysis module consists of two independent neural networks, which process the features of the CFP modality and the OCT modality respectively. Let and represent the instance feature matrices from the CFP modality and the OCT modality respectively. Among them, N is the number of instances, and d1 and d2 are the feature dimensions of the two modalities respectively. The input features of each modality are subjected to multi-layer non-linear transformation through their respective deep neural networks to generate two new feature representations O cfp ∈R N×d and O oct ∈R N×d :

[0135] O cfp =f cfp (H cfp ,W1), O oct =f oct (H oct ,W2)

[0136] where W1 and W2 are the parameters for non-linear transformation of the two deep neural networks respectively; d is the output dimension. In order to measure the correlation between the two modality features, we use the canonical correlation analysis (CCA) method. First, the output features of the two modalities are centered, and then the cross-modal covariance matrix is constructed:

[0137]

[0138] where, ∑ 11 and ∑ 22 represent the covariance matrix of the CFP modality and the covariance matrix of the OCT modality respectively; ∑ 12 is the cross-modal covariance matrix. The goal of DCCA is to jointly optimize the neural network parameters W1 and W2 to maximize the correlation of the cross-modal features O cfp and O oct in the shared space, which can be represented by the following formula:

[0139]

[0140] After being trained by two neural networks, the transformed feature O cfp and O oct The projection directions in the shared space are adjusted to maximize their correlation. According to the above construction process, DCCA brings the following advantages to multi-modal AMD recognition: By separately transforming different modalities, the changing features of the two modalities O cfp and O oct can be explicitly extracted to facilitate the examination of the characteristics and relationships of the two-modal transformations; Under the specified CCA constraints, the non-linear mappings (f cfp () and f oct ()) can be adjusted, and the model can retain the information related to AMD.

[0141] Step S105: Construct a multi-modal based feature fusion module to perform feature fusion on the OCT target feature vector and the CFP target feature vector, and output the target fusion feature.

[0142] When performing automatic diagnosis of AMD diseases, relying solely on single-modal data may lead to the loss of some key pathological information. In addition, the deep learning training process pursues the global optimal solution. However, the setting of hyperparameters (such as the learning rate) often causes the model to fall into the local optimal solution of certain specific points, thereby causing the model to stop optimizing. In the embodiments of the present invention, by constructing a multi-modal based feature fusion module FFM to integrate the features output by the model, the advantages of different modalities can be combined, providing a more comprehensive and detailed description of the lesions, significantly improving the comprehensive analysis ability of the model for key lesion information, and helping to solve the optimization dilemma brought by the local optimal solution. Specifically, the embodiments of the present invention adopt the method of moving average to fuse the feature outputs of the two-modal images, and the steps are as follows:

[0143] Step S1: Take the OCT target feature vector and the CFP target feature vector as input features;

[0144] Step S2: Set a moving average variable to record the fused features; Initialize the moving average variable y avg as a zero matrix (y avg = 0), with the size (H×W×C) being the same as the input features;

[0145] Step S3: For each moment t, update the moving average variable y avg , and the formula is as follows:

[0146] y avg = β·y avg + (1 - β)·y t

[0147] Among them, β is the sliding average decay coefficient, taking a value close to 1, and the sliding average decay coefficient determines the influence degree of historical information on the current fusion result; y t is the input feature at the current moment; y t is the result of element-wise multiplication of the OCT target feature vector and the CFP target feature vector at time t. Preferably, the sliding average decay coefficient can be 0.8;

[0148] Step S4: Output the fused feature y output = y avg ;

[0149] Step S5: Use y output as the input for the next moment t + 1;

[0150] Step S6: Select a specific time step t * as the termination, and output the final target fused feature.

[0151] Based on the sliding average process proposed in the embodiments of the present invention, the fused feature at each moment takes into account the information of historical moments, realizing the smooth fusion of multi-modal image features.

[0152] Step S106, construct a classifier based on a preset classification algorithm, use the target fused feature as the input of the classifier, take the output sub-type category as the prediction result, the AMD degeneration label as the actual result, train and verify the classifier, and obtain the target classifier with the minimized loss function.

[0153] In the embodiments of the present invention for the construction of the macular degeneration classification system, the obtained target fused feature is used as the input, and after passing through the activation function, global average pooling, dropout processing, and the fully connected layer, it enters the classifier constructed by using a preset classification algorithm, and the classifier is trained and optimized. The preset classification algorithm can be selected from support vector machine (SVM), random forest (RandomForest), K-nearest neighbor algorithm (K-Nearest Neighbors, KNN), multi-layer perceptron (MultilayerPerceptron, MLP), etc. The training objective of the classifier is to minimize the loss function, and the embodiments of the present invention use cross-entropy loss as the classification loss for optimization:

[0154]

[0155] Among them, N is the number of training samples; C is the number of categories; y i,j is the actual label of training sample i; is the probability that training sample i is predicted to be category j.

[0156] In addition to the classification loss, the embodiments of the present invention also introduce the DCCA loss function, which aims to enhance the correlation between multimodal features. To adapt to the gradient descent optimization process, the objective function of DCCA is converted into negative correlation, so that it can be optimized by the minimization method. Its loss function is defined as follows:

[0157]

[0158] In the embodiments of the present invention, the overall loss of the classifier is a weighted combination of the classification loss and the DCCA loss, that is:

[0159]

[0160] where λ is the loss weight used to balance the influence of the classification loss and the deep canonical correlation loss on the training of the classifier. The overall loss function comprehensively considers the correlation optimization of feature extraction and the accuracy of the classification task, and realizes the collaborative optimization of the two. Through this dual optimization strategy, the classifier can not only learn highly correlated multimodal features, but also show higher accuracy in the classification task, thus achieving better overall performance in the multimodal AMD classification task.

[0161] Step S107: Integrate the loose pairing technology, the feature extraction module, the deep canonical correlation analysis module, the feature fusion module, and the target classifier to obtain a macular degeneration classification system for performing AMD degeneration classification.

[0162] The macular degeneration classification system constructed in the embodiments of the present invention acquires OCT images and CFP images, sets the two-modal images to a preset size, and respectively processes them through the above-mentioned loose pairing technology, feature extraction module, deep canonical correlation analysis module, and feature fusion module. Then, the trained target classifier can be used to perform classification prediction on the input of the two-modal images. The final classification prediction result is one of the four AMD degeneration labels: normal retina (normal), dry AMD (dryAMD), wet AMD (wetAMD), and polypoidal choroidal vasculopathy (PCV).

[0163] The present invention proposes a method for constructing a macular degeneration classification system, aiming to solve the problems of insufficient accuracy and high computational complexity in existing AMD automated screening due to scarce high-quality labeled data, modal heterogeneity, etc. By finely tuning a large pre-trained single-modal retinal basic model (such as RETFound) and combining the multi-modal information of color fundus photography (CFP) and optical coherence tomography (OCT) images, the present invention realizes the effective fusion of multi-modal features. The RETFound model already has a powerful feature extraction ability in the single-modal case, and by expanding it to multi-modal applications, the present invention enables it to not only make full use of the surface texture information of CFP images but also effectively capture the deep structure features in OCT images, thus significantly improving the recognition ability and diagnostic accuracy for complex AMD lesions. To cope with the computational burden brought by multi-modal fusion, the present invention introduces the Low-Rank Adaptation (LoRA) technique. LoRA decomposes the full parameter matrix in the model into two low-rank matrices, so that only a small number of parameters need to be adjusted during the fine-tuning process, greatly reducing the computational and memory requirements. At the same time, to solve the heterogeneity problem between the CFP and OCT modalities, the deep canonical correlation analysis (DCCA) method is adopted. Through two independent neural networks, non-linear transformations are performed on the features of different modalities, and their linear correlation is maximized through regularization. This method can map the features of different modalities to a unified high-dimensional space, laying a foundation for the deep fusion of cross-modal information. Then, the feature outputs of the two-modal images are fused by the method of moving average, and the features between modalities are fused by element-wise multiplication, so that the multi-modal information can be fully utilized, avoiding the loss and redundancy of feature information. At the same time, the method of moving average can smooth the historical information, improving the stability and consistency of the fusion result. In addition, the moving average multi-modal image fusion method has real-time performance and can perform fusion while extracting features, without waiting for all features to be extracted before processing. This can save time and computational resources and improve the overall processing efficiency. The superiority of the present invention is mainly reflected in the following five aspects:

[0164] ① By finely tuning a large pre-trained single-modal retinal basic model, it is effectively extended to a model for multi-modal AMD recognition; while realizing multi-modal recognition, the powerful processing ability of the large single-modal model in a single modality is retained.

[0165] ② To cope with the challenge of increased computational complexity in the process of multi-modal data integration, the present invention introduces the LoRA method. By decomposing the model parameters into low-rank matrices, only a small number of parameters need to be optimized during the fine-tuning stage, significantly reducing the demand for computational resources;

[0166] ③ To address the heterogeneity issue among different modal features, the present invention adopts the DCCA method to perform non-linear transformation on CFP and OCT modal features through independent neural networks, thereby achieving highly correlated multi-modal feature learning;

[0167] ④ By using the moving average method to fuse the features of the two modal images, making full use of multi-modal information, avoiding feature loss and redundancy, and improving the stability and consistency of the fusion result. Meanwhile, it saves time and computing resources, improves the overall processing efficiency, and is of great value for the processing of the two modal image data and model training.

[0168] ⑤ The present invention can be used for AMD classification during primary screening.

[0169] Furthermore, the embodiment of the present invention also provides a macular degeneration classification system, which is constructed by the construction method of the macular degeneration classification system as shown in Figure 1 As shown, this system can at least include: a loose pairing module 610, a feature extraction module 620, a deep canonical correlation analysis module 630, a feature fusion module 640, and a target classifier 650. Figure 6

[0170] The loose pairing module 610 can be used to obtain the OCT image and CFP image to be classified, and use the loose pairing technology to pair the OCT image and CFP image with the same AMD degeneration label to obtain multiple groups of OCT-CFP training sample pairs;

[0171] The feature extraction module 620 can be used to extract features from the OCT-CFP training sample pairs to obtain the OCT high-dimensional feature vector and CFP high-dimensional feature vector after feature transformation based on the low-rank adaptation technology;

[0172] The deep canonical correlation analysis module 630 can be used to perform non-linear feature transformation on the OCT high-dimensional feature vector and CFP high-dimensional feature vector to obtain the OCT target feature vector and CFP target feature vector with maximized correlation in the shared space;

[0173] The feature fusion module 640 can be used to fuse the OCT target feature vector and CFP target feature vector to obtain the target fusion feature;

[0174] The target classifier 650 can be used to take the target fusion feature as input and output the AMD degeneration categories corresponding to the OCT image and CFP image to be classified;

[0175] Among them, the AMD degeneration categories include normal, dry, wet, and polypoidal choroidal vasculopathy.

[0176] ​It should be noted that for other corresponding descriptions of each functional module involved in the macular degeneration classification system provided in the embodiments of the present invention, reference can be made to Figure 1 the corresponding description of the method shown, which will not be elaborated here.

[0177] Those skilled in the art can clearly understand that the specific working processes of the above-described system, device, module, and unit can refer to the corresponding processes in the foregoing method embodiments. For the sake of brevity, they will not be described separately here.

[0178] In addition, each functional unit in various embodiments of the present invention can be physically independent of each other, or two or more functional units can be integrated together, or all functional units can be integrated in a processing unit. The above-mentioned integrated functional units can be implemented in the form of hardware, or in the form of software or firmware.

[0179] Those of ordinary skill in the art can understand that if the integrated functional unit is implemented in the form of software and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computing device (such as a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present invention when the instructions are run. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0180] Alternatively, all or part of the steps of implementing the foregoing method embodiments can be completed by hardware related to program instructions (such as a computing device such as a personal computer, a server, or a network device), and the program instructions can be stored in a computer-readable storage medium. When the program instructions are executed by the processor of the computing device, the computing device executes all or part of the steps of the methods described in the embodiments of the present invention.

[0181] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that within the spirit and principles of the present invention, they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the corresponding technical solutions to deviate from the protection scope of the present invention.

Claims

1. A method for constructing a macular degeneration classification system, characterized in that The method includes: Obtaining a plurality of training samples carrying different AMD degeneration labels to form a multi-modal AMD image data set; wherein each of the training samples includes an OCT image and a CFP image; the AMD degeneration labels include normal, dry, wet, and polypoidal choroidal vasculopathy; Using a loose pairing technique to pair the OCT images and CFP images in the multi-modal AMD image data set with the same AMD degeneration label to obtain multiple groups of OCT-CFP training sample pairs; Constructing a feature extraction module integrating the low-rank and adaptive technique, and using the feature extraction module to extract features from the OCT-CFP training sample pairs to obtain an OCT high-dimensional feature vector and a CFP high-dimensional feature vector; Constructing a deep canonical correlation analysis module based on a neural network, and respectively performing non-linear feature transformation on the OCT high-dimensional feature vector and the CFP high-dimensional feature vector to obtain an OCT target feature vector and a CFP target feature vector with maximized correlation in the shared space; Constructing a multi-modal based feature fusion module to perform feature fusion on the OCT target feature vector and the CFP target feature vector, and outputting a target fusion feature; Constructing a classifier based on a preset classification algorithm, using the target fusion feature as the input of the classifier, and using the output sub-type category as the prediction result and the AMD degeneration label as the actual result to train and verify the classifier to obtain a target classifier with minimized loss function; Integrating the loose pairing technique, the feature extraction module, the deep canonical correlation analysis module, the feature fusion module, and the target classifier to obtain a macular degeneration classification system for performing AMD degeneration classification.

2. The method according to claim 1, wherein Before using the loose pairing technique to pair the OCT images and CFP images in the multi-modal AMD image data set with the same AMD sub-type label, the method further includes: Performing image preprocessing on the multi-modal image data set to obtain a preprocessed target image data set; the target image data set includes a plurality of training samples carrying different AMD degeneration labels; each training sample includes a preprocessed OCT image and a CFP image; wherein the image preprocessing includes the following steps: S1: Deleting abnormal image data in the multi-modal image data set to obtain a first image data set; S2: Performing image normalization processing and image weighted enhancement processing on the first image data in the first image data set to obtain a second image data set; S3: Performing image enhancement processing on the second image data in the second image data set.

3. The method according to claim 2, wherein The image normalization processing includes at least one of normalizing the channel numerical range of image pixels, equalizing the means of image pixels, standardizing image pixels, and uniformly adjusting the image size; And / or The processing formula of the image weighted enhancement is expressed as: I weight = I org * α + I blur * β + γ I blur = I org * kernel h×w where, I weight is the second image data; I org is the first image data; I blur is the image after Gaussian blur of the first image data, and kernal h×w represents a Gaussian kernel with a size of h×w; α and β are weighting coefficients; γ is a constant; And / or The image enhancement includes at least one of image rotation, image translation, random flipping, applying contrast-limited adaptive histogram equalization to the CFP image, and applying median filtering method to the OCT image.

4. The method according to claim 2, wherein Using the loose pairing technique to pair the OCT images and CFP images in the multimodal AMD image dataset with the same AMD degeneration labels to obtain multiple groups of OCT-CFP training sample pairs, including: For each training sample i in the target image dataset I aug the CFP image I aug _f i and the OCT image I aug _o i are used as the original pair; Initializing an empty loose pairing list; For other training samples j in the target image dataset, where j ≠ i, if the AMD degeneration labels of the training sample i and the training sample j are the same, then add (I aug _f i , I aug _o j ) and (I aug _f j , I aug _o i ) to the loose pairing list; Merging the original pairs and the loose pairing list to obtain a multimodal AMD image set that expands multiple groups of OCT-CFP training sample pairs.

5. The method according to claim 1, wherein Constructing a feature extraction module integrating the low-rank adaptation technique, and using the feature extraction module to extract features from the OCT-CFP training sample pairs to obtain an OCT high-dimensional feature vector and a CFP high-dimensional feature vector, including: Constructing a feature extraction module including two independent ViT-large encoders; the ViT-large encoder includes a patch embedding layer, 24 Transformer blocks, and a 1024-dimensional embedding vector; Input the CFP image \(X\) cfp \(\in\mathbb{R}\) H×W×C and the OCT image \(X\) oct \(\in\mathbb{R}\) H×W×C into two independent ViT-large encoders respectively; Divided into patches of a fixed size through the patch embedding layer and mapped to 1024-dimensional embedding vectors, denoted as P cfp ∈R 1×1024 and P oct ∈R 1×1024 ; Process the embedding vector using positional encoding to generate corresponding input feature vectors, denoted as Z 0 cfp ∈R 1 ×1024 and Z 0 oct ∈R 1×1024 ; Inputting the input feature vector into the Transformer block pair for processing, and the output formula is expressed as: After being processed by 24 Transformer blocks, the high-dimensional feature vector Z to be converted is obtained cfp ∈R 1×1024 and Z oct ∈R 1×1024 .

6. The method according to claim 5, wherein The method further includes: Using the low-rank adaptation technique to perform low-rank decomposition on the query matrix WQ and the value projection matrix WV in the self-attention layer of the Transformer block to optimize weight update, and the formula is expressed as follows: ΔW = BA where B ∈ R d×r and A ∈ R r×d are matrices generated by low-rank decomposition, and the rank r is much smaller than the dimension d; Process the input feature vector using the optimized Transformer block with weight update to obtain the OCT high-dimensional feature vector h oct ∈R 1×1024 and the CFP high-dimensional feature vector h cfp ∈R 1×1024 , which is expressed by the following formula: h = W0Z + ΔWZ = W0Z + BAZ Where, W0 is the initial pre-trained weight; ΔW is the weight update term after low-rank decomposition; Z is the high-dimensional feature vector to be transformed.

7. The method according to claim 1, wherein Constructing a deep canonical correlation analysis module based on a neural network, and performing non-linear feature transformation on the OCT high-dimensional feature vector and the CFP high-dimensional feature vector respectively to obtain an OCT target feature vector and a CFP target feature vector with maximized correlation in the shared space; Construct a deep canonical correlation analysis module that includes two independent deep neural networks. Input the OCT high-dimensional feature vector and the CFP high-dimensional feature vector into the two independent deep neural networks respectively for multiple non-linear transformations to generate output features O cfp ∈R N×d and O oct ∈R N×d , which is expressed by the formula as follows: O cfp = f cfp (h cfp , W1), O oct = f oct (h oct , W2) Among them, are the CFP instance feature matrix and the OCT instance feature matrix respectively; N is the number of instances; d1 is the feature dimension of the CFP modality; d2 is the feature dimension of the OCT modality; W1 and W2 are the parameters for non-linear transformation by two deep neural networks respectively; d is the output dimension; Centering the output features and constructing a cross-modal covariance matrix: Among them, Σ 11 and Σ 22 represent the covariance matrix of the CFP mode and the covariance matrix of the OCT mode respectively; Σ 12 is the cross-modal covariance matrix; Maximize the output feature O using the following formula cfp and O oct The correlation in the shared space to obtain the OCT target feature vector and the CFP target feature vector, which is expressed by the formula as follows:

8. The method according to claim 1, characterized in that, Constructing a multimodal-based feature fusion module to fuse the OCT target feature vector and the CFP target feature vector and output a target fusion feature, including: Initialize the moving average variable y with the OCT target feature vector and the CFP target feature vector as input features. avg It is a zero matrix with the same size as the input features. Update the moving average variable y avg , which is expressed by the following formula: y avg = β·y avg + (1 - β)·y t where β is the sliding average decay coefficient; y t is the input feature at the current moment; y t is the result of element-wise multiplication of the OCT target feature vector and the CFP target feature vector at time t; Output the fused feature y output = y avg ; Take y output as the input for the next moment t + 1; Select a specific time step t * As the termination, output the final target fusion feature.

9. The method according to claim 7, wherein The preset classification algorithm includes at least one of support vector machine, random forest, K-nearest neighbor algorithm, and multi-layer perceptron; The loss function formula is expressed as: Among them, is the classification loss function; is the deep canonical correlation loss function; N is the number of training samples; C is the number of classes; y i,j is the actual result of training sample i; is the probability that the predicted result of training sample i is class j; λ is the loss weight used to balance the influence of classification loss and deep canonical correlation loss on the classifier training.

10. A macular degeneration classification system constructed by the construction method of the macular degeneration classification system according to any one of claims 1 to 9, characterized in that, The system includes: a loose pairing module, a feature extraction module, a deep canonical correlation analysis module, a feature fusion module, and a target classifier; The loose pairing module is used to obtain the OCT image and the CFP image to be classified, and use the loose pairing technique to pair the OCT image and the CFP image with the same AMD degeneration labels to obtain multiple groups of OCT-CFP training sample pairs; The feature extraction module is used to extract features from the OCT-CFP training sample pairs to obtain an OCT high-dimensional feature vector and a CFP high-dimensional feature vector after feature transformation based on the low-rank adaptation technique; The deep canonical correlation analysis module is used to perform non-linear feature transformation on the OCT high-dimensional feature vector and the CFP high-dimensional feature vector to obtain the OCT target feature vector and the CFP target feature vector with maximized correlation in the shared space; The feature fusion module is used to perform feature fusion on the OCT target feature vector and the CFP target feature vector to obtain the target fusion feature; The target classifier is used to take the target fusion feature as input and output the AMD degeneration category corresponding to the OCT image and the CFP image to be classified; Among them, the AMD degeneration categories include normal, dry, wet, and polypoidal choroidal vasculopathy.