Multi-modal data glaucoma classification method and system, terminal and medium
Through the Mamba model combining feature extraction and fusion of color fundus and OCT images, the problems of insufficient accuracy of glaucoma classification and large computing overhead in the prior art are solved, and efficient glaucoma classification is achieved.
Patent Information
- Application Number
- CN202510269726.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to effectively classify glaucoma with color fundus images and OCT images, resulting in insufficient diagnostic accuracy, and the existing multimodal image fusion method has a large calculation overhead, making it difficult to capture the global context.
The multimodal data glaucoma classification method based on the Mamba model is adopted, and the pattern-specific features of color fundus and OCT images are extracted respectively through the first and second feature extraction modules, and the feature fusion is used for the shallow and deep feature fusion modules, and the classification prediction is performed using multi-layer perceptrons.
It improves the accuracy of glaucoma classification, reduces the computational complexity, can effectively capture global context features, and achieve accurate glaucoma classification.
Smart Images

Figure CN120108027A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a multimodal data glaucoma classification method, system, terminal and medium. Background Art
[0002] Glaucoma is a chronic neurodegenerative disease and one of the leading causes of irreversible but preventable blindness in the world. Blindness is often caused by lack of timely detection and treatment. Therefore, early screening is essential for early treatment to preserve vision and maintain quality of life. Color fundus imaging and optical coherence tomography (OCT) are the two most cost-effective images for glaucoma screening. Both imaging modalities have significant biomarkers to indicate the suspicion of glaucoma, such as the vertical cup-to-disc ratio (vCDR) on fundus images and the retinal nerve fiber layer (RNFL) thickness on OCT images. In clinical practice, screening with these two images is usually recommended to obtain a more accurate and reliable diagnosis.
[0003] Although many algorithms based on color fundus images or OCT images have been proposed to automatically detect glaucoma, few methods combine the two modalities to achieve the goal, mainly due to the differences in the characteristics and dimensions of the two modalities. The task is technically challenging. According to reports, nearly 46.3% of glaucoma cases will be ignored if fundus images or OCT images are used alone. In addition, existing multimodal image fusion methods, such as those based on convolutional neural networks, have difficulty capturing global context due to their limited receptive fields, which makes it challenging to generate high-quality fused images. Secondly, although models based on Transformer or a combination of Transformer and other models show good performance in global modeling, the computational overhead increases significantly due to the self-attention mechanism.
[0004] Therefore, the prior art still has defects. Summary of the invention
[0005] The technical problem to be solved by the present invention is to provide a multimodal data glaucoma classification method and system in view of the above-mentioned defects of the prior art. The technical solution adopted by the present invention is as follows:
[0006] In a first aspect, the present invention provides a multimodal data glaucoma classification method, wherein the method comprises:
[0007] Acquire a color fundus image and an OCT image in different image modes, and perform high-level feature extraction on the color fundus image and the OCT image based on a first feature extraction module and a second feature extraction module, respectively, to obtain mode-specific features of the two modalities;
[0008] The mode-specific features of the two modalities are fused based on the shallow feature fusion module to obtain shallow fusion features;
[0009] Inputting the mode-specific features of the two modalities and the shallow fusion features into the deep feature fusion module for fusion to obtain the final fusion features;
[0010] The final fusion features are classified and predicted based on a multi-layer perceptron classifier to obtain classification results, which include: no glaucoma, early glaucoma or mid-to-late glaucoma.
[0011] In one implementation, the first feature extraction module consists of a convolutional layer and a plurality of stacked Mamba blocks, and the second feature extraction module consists of a convolutional layer and a plurality of stacked Vision Mamba blocks.
[0012] In one implementation, an algorithm for performing high-level feature extraction on the color fundus image and the OCT image based on the first feature extraction module and the second feature extraction module is expressed as:
[0013]
[0014] in, and They represent the feature sequences obtained after low-level feature extraction of color fundus images and OCT images, respectively, 1n Represents a feature sequence Processed by n Mamba blocks, φ 2n Represents a feature sequence Processed by n Vision Mamba blocks, and Represents the modality-specific features of the two modalities.
[0015] In one implementation, the mode-specific features of the two modalities are fused based on the shallow feature fusion module to obtain the algorithm of shallow fusion features as follows:
[0016]
[0017] Among them, the Represents shallow fusion features.
[0018] In one implementation, the deep feature fusion module is a multimodal Mamba module, which uses features of a specific modality to guide the generation of modality fusion features.
[0019] In one implementation, the mode-specific features of the two modalities and the shallow fusion features are input into the deep feature fusion module for fusion, and the algorithm for obtaining the final fusion features is expressed as follows:
[0020]
[0021] in, and The modality-specific features representing the two modalities serve as two additional input branches to the deep feature fusion module. Represents deep fusion features.
[0022] In a second aspect, an embodiment of the present invention further provides a multimodal data glaucoma classification system, the system is used to implement the steps of the multimodal data glaucoma classification method described in any one of the above schemes, the system comprising:
[0023] A feature extraction module, used for acquiring color fundus images and OCT images of different image modes, performing advanced feature extraction on the color fundus images and the OCT images respectively, to obtain mode-specific features of the two modalities;
[0024] The shallow feature fusion module is used to fuse the mode-specific features of the two modalities to obtain shallow fusion features;
[0025] A deep feature fusion module is used to fuse the mode-specific features of the two modalities and the shallow fusion features to obtain a final fusion feature;
[0026] The classification prediction module is used to perform classification prediction on the final fusion features based on a multi-layer perceptron classifier to obtain a classification result, wherein the classification result includes: no glaucoma, early glaucoma or mid-to-late glaucoma.
[0027] In one implementation, the feature extraction module includes a first feature extraction module and a second feature extraction module; the first feature extraction module consists of a convolutional layer and a plurality of stacked Mamba blocks, and the second feature extraction module consists of a convolutional layer and a plurality of stacked Vision Mamba blocks.
[0028] In a third aspect, an embodiment of the present invention further provides a terminal, wherein the terminal includes a memory, a processor, and a multimodal data glaucoma classification program stored in the memory and executable on the processor, and when the processor executes the multimodal data glaucoma classification program, the steps of the multimodal data glaucoma classification method of any one of the above-mentioned schemes are implemented.
[0029] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein a multimodal data glaucoma classification program is stored on the computer-readable storage medium, and when the multimodal data glaucoma classification program is executed by a processor, the steps of the multimodal data glaucoma classification method described in any one of the above schemes are implemented.
[0030] Beneficial effects: Compared with the prior art, the present invention provides a multimodal data glaucoma classification method. The present invention first obtains color fundus images and OCT images of different image modes, and performs high-level feature extraction on the color fundus images and the OCT images based on the first feature extraction module and the second feature extraction module, respectively, to obtain mode-specific features of the two modalities. Then, the mode-specific features of the two modalities are fused based on the shallow feature fusion module to obtain shallow fusion features. Next, the mode-specific features of the two modalities and the shallow fusion features are input into the deep feature fusion module for fusion to obtain the final fusion features. Then, the final fusion features are classified and predicted based on the multilayer perceptron classifier to obtain classification results, which include: no glaucoma, early glaucoma, or mid-to-late glaucoma. The present invention is conducive to accurate glaucoma classification through two-level feature extraction, two-stage feature fusion and classifier. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 A flowchart of a preferred embodiment of the multimodal data glaucoma classification method provided in an embodiment of the present invention.
[0032] Figure 2 This is a network flow chart of a multimodal data glaucoma classification method according to an embodiment of the present invention.
[0033] Figure 3 Schematic diagram of a shallow feature fusion module in a multimodal data glaucoma classification method according to an embodiment of the present invention.
[0034] Figure 4 Schematic diagram of a deep feature fusion module in a multimodal data glaucoma classification method according to an embodiment of the present invention.
[0035] Figure 5 A schematic diagram of the architecture of a multimodal data glaucoma classification system provided in an embodiment of the present invention.
[0036] Figure 6 A functional block diagram of a terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0037] In order to make the purpose, technical solution and effect of the present invention clearer and more specific, the present invention is further described in detail with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0038] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations or steps, nor must they be executed in the order described. For example, some operations or steps may also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.
[0039] It should be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0040] It should be understood that, in order to clearly describe the technical solutions of the embodiments of the present invention, in the embodiments of the present invention, words such as "first" and "second" are used to distinguish between identical or similar items with substantially identical functions and effects. For example, the first control information and the second control information are only used to distinguish different control information, and their order is not limited.
[0041] Those skilled in the art can understand that the words "first", "second", etc. do not limit the quantity and execution order, and the words "first", "second", etc. do not necessarily limit the differences.
[0042] It should be further understood that the term “and / or” used in the present specification and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0043] Based on the defects of the prior art, this embodiment provides a multimodal data glaucoma classification method, which can be applied to a terminal, which can be a computer, a smart TV, a mobile phone and other intelligent terminal products. Figure 1 As shown in , the multimodal data glaucoma classification method of this embodiment includes the following steps:
[0044] Step S100, obtaining color fundus images and OCT images in different image modes, and performing high-level feature extraction on the color fundus images and the OCT images based on the first feature extraction module and the second feature extraction module, respectively, to obtain mode-specific features of the two modalities.
[0045] Step S200: fusing the mode-specific features of the two modalities based on the shallow feature fusion module to obtain shallow fusion features;
[0046] Step S300: inputting the mode-specific features of the two modalities and the shallow fusion features into a deep feature fusion module for fusion to obtain a final fusion feature;
[0047] Step S400: classify and predict the final fusion features based on a multi-layer perceptron classifier to obtain a classification result, wherein the classification result includes: no glaucoma, early glaucoma, or mid-to-late glaucoma.
[0048] Combination Figure 2 As shown, this embodiment designs a dual-branch fusion network based on Mamba for glaucoma classification of multimodal data, which includes dual-branch feature extraction, two-stage feature fusion and a multi-layer perceptron classifier. The dual-branch feature extraction is a feature extraction branch for color fundus images and a feature extraction branch for OCT images. Figure 2 As shown, two different image modes of color fundus images I are input. 1 and OCT images I 2 , use I f Represents the color fundus image I 1 and OCT images I 2 The feature extraction stage includes low-level and high-level feature extraction. In low-level feature extraction, I 1 and I 2 Through the convolutional layer Projected into a shared feature space, low-level feature extraction includes two convolutional layers, using Leaky ReLU activation function, each with a kernel size of 3×3 and a stride of 1. However, the convolutional layer may not be able to capture global features due to its limited receptive field. Therefore, after the image patch is embedded, the color fundus image I 1 Using N Mamba blocks for high-level feature extraction, OCT image I 2 High-level feature extraction using N Vision Mamba blocks to obtain pattern-specific features and Among them, in the Mamba block, an input feature sequence F t-1 ∈R B×N×C , first apply layer normalization to get F t ' -1 , then, on two separate branches, F t ' -1 Two multi-layer perceptrons (MLPs) are used to project onto x and z. In the first branch, x is convolved and activated with SiLU to obtain x′. Subsequently, x′ is used to calculate f through the state-space model. In the other branch, z uses a SiLU activation function as the gating factor of f to obtain f′. Finally, after the MLP layer and residual connection, the high-level feature F is obtained. tAs output. The color fundus image branch can extract features related to the optic disc and cup area, which is a biomarker associated with the occurrence of glaucoma in clinical practice. In the Vision Mamba block (Vim), the Vim model first converts the OCT image into a series of image patches, and then maps these patches into vectors with additional position encoding. In addition, a category tag is included to represent the entire sequence of patches. These vectors are input into the Vim encoder, and the final output of the category tag is passed to the subsequent module. The model is able to process both forward and backward information in the sequence, thereby capturing the comprehensive global context of data dependence. It compresses the visual representation through a bidirectional selection state space model without relying on the self-attention mechanism commonly used in ViT. This method overcomes the quadratic computational complexity associated with the Transformer. By introducing position embedding, Vim can better understand the complex spatial relationships in the visual data. Specifically, the Vim module normalizes the sequence, linearly projects it into two components x and z, and then processes x in the forward and reverse directions. For each direction, a one-dimensional convolution is first applied, and then processed in a state space module (SSM), and the resulting forward and reverse outputs y_forward y_backward are gated by z and combined to obtain the final output of the Vim module. The OCT image branch can extract features related to the thickness of the retinal nerve fiber layer, which is often used in clinical practice to determine the probability of glaucoma. In this embodiment, the algorithm for high-level feature extraction of the color fundus image and the OCT image is expressed as:
[0049]
[0050] in, and They represent the feature sequences obtained after low-level feature extraction of color fundus images and OCT images, respectively, 1n Represents a feature sequence Processed by n Mamba blocks, φ 2n Represents a feature sequence Processed by n Vision Mamba blocks, and Represents the modality-specific features of the two modalities.
[0051] Then, in the shallow feature fusion module, the fusion strategy of the shallow fusion rule is adopted to obtain For local detail features, an improved Mamba block is used to construct a deep fusion module, and the deep fusion features are derived using multimodal features as a guide. Specifically, in this embodiment, the shallow feature fusion module fuses the mode-specific features of the two modalities to obtain the algorithm of the shallow fusion feature as follows:
[0052]
[0053] Among them, the Represents shallow fusion features.
[0054] Combination Figure 3 As shown in , a channel exchange method is used in the shallow feature fusion module. and To process, and The extracted two modal features are used as inputs to the shallow feature fusion module, which do not require additional parameters or computational operations, thus achieving lightweight exchange of features from multiple modalities. The exchanged features are then processed by their respective Mamba blocks. By repeating the above steps, the modality-specific features and Able to integrate features from another modality and then perform a fusion operation to obtain a shallow fusion feature The fusion operation can be an addition operation or an L1 regularization operation. The fusion process can be expressed as follows:
[0055]
[0056] Among them, M(B, N, C) refers to the mask for channel swapping, which consists of 1 and 0, where 0 means no swap and 1 means swap. B and N represent the size of the feature map, which is B×N, and C represents the number of channels.
[0057] Since the shallow feature fusion module can only process global features and cannot process deep texture detail features, it is necessary to further perform local feature fusion operations. Since the current Mamba architecture cannot directly process multimodal image information because it lacks a mechanism similar to cross attention, in order to improve this situation, a multimodal Mamba module is used, which uses the features of a specific modality to guide the generation of deep fusion features, such as Figure 4As shown. The input is the shallow fusion feature from the shallow fusion module. In addition, two additional input branches are introduced. Each input branch takes the features of different modalities as input. In this embodiment, the mode-specific features of the two modalities can be used as two additional input branches of the deep feature fusion module. Similarly, these branches are processed by layer normalization, convolution, SiLU activation and parameter discretization, and the output y is obtained through SSM. After modulation by the gating factor, it is added to the output of the original branch to obtain the final fusion feature. The deep feature fusion module is composed of M improved Mamba blocks. Specifically, in this embodiment, the mode-specific features of the two modalities and the shallow fusion features are input into the deep feature fusion module for fusion, and the algorithm for obtaining the final fusion feature is expressed as:
[0058]
[0059] in, and The modality-specific features representing the two modalities serve as two additional input branches to the deep feature fusion module. Represents deep fusion features.
[0060] Finally, this embodiment performs classification prediction on the final fusion features based on a multi-layer perceptron classifier to obtain classification results, which include: no glaucoma, early glaucoma, or mid-to-late glaucoma.
[0061] It can be seen that in the glaucoma classification method of this embodiment, the Mamba model is used for feature extraction and fusion. Through the parameterized selection mechanism, specific data can be selected or ignored in a targeted manner according to the characteristics of the input data, thereby capturing global information while maintaining linear complexity, which is both effective and efficient. At the same time, the Mamba model also has excellent performance in multimodal data. It performs well in aligning mixed source data and achieving linear complexity expansion of sequence length, enabling it to provide very valuable and complementary information. Most of the existing glaucoma classification methods for multimodal data are based on convolutional neural networks. Due to their limited receptive fields, it is difficult to capture global context features. The Mamba model can effectively overcome this limitation.
[0062] Based on the above embodiment, the present invention further provides a multimodal data glaucoma classification system, which is used to implement the steps of the multimodal data glaucoma classification method in the above embodiment, such as Figure 5As shown in , the system includes: a feature extraction module 10, a shallow feature fusion module 20, a deep feature fusion module 30 and a classification prediction module 40. Specifically, the feature extraction module 10 is used to obtain color fundus images and OCT images of different image modes, and perform high-level feature extraction on the color fundus images and the OCT images respectively to obtain mode-specific features of the two modalities. The shallow feature fusion module 20 is used to fuse the mode-specific features of the two modalities to obtain shallow fusion features. The deep feature fusion module 20 is used to fuse the mode-specific features of the two modalities and the shallow fusion features to obtain final fusion features. The classification prediction module 30 is used to perform classification prediction on the final fusion features based on a multi-layer perceptron classifier to obtain a classification result, and the classification result includes: no glaucoma, early glaucoma or mid-to-late glaucoma.
[0063] In one implementation, the feature extraction module includes a first feature extraction module and a second feature extraction module; the first feature extraction module consists of a convolutional layer and a plurality of stacked Mamba blocks, and the second feature extraction module consists of a convolutional layer and a plurality of stacked Vision Mamba blocks.
[0064] Based on the above embodiment, the present invention further provides a multimodal data glaucoma classification system, which is used to implement the steps of the multimodal data glaucoma classification method in the above embodiment. Figure 5 As shown in, the system includes: a data acquisition module 10, a signal processing module 20, a feature screening module 30, a model training module 40 and a result output module 50. Specifically, the data acquisition module 10 is used to collect multimodal physiological signals of the user when performing a virtual task, and the multimodal physiological signals include: electromyographic signals, electrocardiographic signals and electrodermal signals. The signal processing module 20 is used to perform signal processing on the multimodal physiological signals to obtain processed physiological signals. The feature screening module 30 is used to perform feature screening on the processed physiological signals to obtain several effective features that are highly correlated with the VR user experience, and combine the independent variables to obtain a feature vector. The model training module 40 is used to train multiple machine learning algorithms based on the feature vector, construct multiple prediction models, and select the optimal prediction model. The result output module 50 is used to output prediction results for different dimensions of VR experience and comprehensive results of user experience based on the optimal prediction model.
[0065] The working principles of each module in the multimodal data glaucoma classification system of this embodiment are the same as the principles of each step in the above method embodiment, and will not be repeated here.
[0066] Each module in the above multimodal data glaucoma classification system can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of the processor in the terminal in the form of hardware, or can be stored in the memory in the terminal in the form of software, so that the processor can call and execute the operations corresponding to each module above.
[0067] Based on the above embodiment, the present invention further provides a terminal, the principle block diagram of the terminal can be as follows: Figure 6 The terminal may include one or more processors 100 ( Figure 6 Only one is shown in the figure), a memory 101 and a computer program 102 stored in the memory 101 and executable on one or more processors 100. For example, a multimodal data glaucoma classification program. When one or more processors 100 execute the computer program 102, the various steps in the embodiment of the multimodal data glaucoma classification method can be implemented. Alternatively, when one or more processors 100 execute the computer program 102, the functions of each module / unit in the embodiment of the multimodal data glaucoma classification system can be implemented, which is not limited here.
[0068] In one embodiment, the processor 100 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0069] In one embodiment, the memory 101 may be an internal storage unit of an electronic device, such as a hard disk or memory of the electronic device. The memory 101 may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the memory 101 may also include both an internal storage unit of the electronic device and an external storage device. The memory 101 is used to store computer programs and other programs and data required by the terminal. The memory 101 may also be used to temporarily store data that has been output or is to be output.
[0070] Those skilled in the art will understand that Figure 6 The principle block diagram shown in the figure is only a block diagram of a partial structure related to the scheme of the present invention, and does not constitute a limitation on the terminal to which the scheme of the present invention is applied. The specific terminal may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0071] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, operating database or other media used in the embodiments provided by the present invention may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double operational data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0072] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multimodal data glaucoma classification method, characterized in that: The method comprises: Acquire color fundus images and OCT images in different image modes, and perform advanced feature extraction on the color fundus images and the OCT images based on the first feature extraction module and the second feature extraction module, respectively, to obtain mode-specific features of the two modalities; The mode-specific features of the two modalities are fused based on the shallow feature fusion module to obtain shallow fusion features; Inputting the mode-specific features of the two modalities and the shallow fusion features into the deep feature fusion module for fusion to obtain the final fusion features; The final fusion features are classified and predicted based on a multi-layer perceptron classifier to obtain classification results, which include: no glaucoma, early glaucoma or mid-to-late glaucoma.
2. The multimodal data glaucoma classification method according to claim 1, characterized in that: The first feature extraction module consists of a convolutional layer and a plurality of stacked Mamba blocks, and the second feature extraction module consists of a convolutional layer and a plurality of stacked Vision Mamba blocks.
3. The multimodal data glaucoma classification method according to claim 1, characterized in that: The algorithm for performing high-level feature extraction on the color fundus image and the OCT image based on the first feature extraction module and the second feature extraction module is expressed as follows: in, and They represent the feature sequences obtained after low-level feature extraction of color fundus images and OCT images, respectively, 1n Represents a feature sequence Processed by n Mamba blocks, φ 2n Represents a feature sequence Processed by n VisionMamba blocks, and Represents the modality-specific features of the two modalities.
4. The multimodal data glaucoma classification method according to claim 1, characterized in that: Based on the shallow feature fusion module, the mode-specific features of the two modalities are fused to obtain the algorithm of shallow fusion features as follows: Among them, the Represents shallow fusion features.
5. The multimodal data glaucoma classification method according to claim 4, characterized in that: The deep feature fusion module is a multimodal Mamba module, which uses features of a specific modality to guide the generation of modality fusion features.
6. The multimodal data glaucoma classification method according to claim 5, characterized in that: The mode-specific features of the two modalities and the shallow fusion features are input into the deep feature fusion module for fusion, and the algorithm for obtaining the final fusion features is expressed as: in, and The modality-specific features representing the two modalities serve as two additional input branches to the deep feature fusion module. Represents deep fusion features.
7. A multimodal data glaucoma classification system, characterized in that: The system is used to implement the steps of the multimodal data glaucoma classification method according to any one of claims 1 to 6, and the system comprises: A feature extraction module, used for acquiring color fundus images and OCT images of different image modes, performing advanced feature extraction on the color fundus images and the OCT images respectively, to obtain mode-specific features of the two modalities; The shallow feature fusion module is used to fuse the mode-specific features of the two modalities to obtain shallow fusion features; A deep feature fusion module is used to fuse the mode-specific features of the two modalities and the shallow fusion features to obtain a final fusion feature; The classification prediction module is used to perform classification prediction on the final fusion features based on a multi-layer perceptron classifier to obtain a classification result, wherein the classification result includes: no glaucoma, early glaucoma or mid-to-late glaucoma.
8. The multimodal data glaucoma classification system according to claim 7, characterized in that: The feature extraction module includes a first feature extraction module and a second feature extraction module; the first feature extraction module is composed of a convolutional layer and a plurality of stacked Mamba blocks, and the second feature extraction module is composed of a convolutional layer and a plurality of stacked Vision Mamba blocks.
9. A terminal, characterized in that: The terminal includes a memory, a processor, and a multimodal data glaucoma classification program stored in the memory and executable on the processor. When the processor executes the multimodal data glaucoma classification program, the steps of the multimodal data glaucoma classification method according to any one of claims 1 to 6 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a multimodal data glaucoma classification program, and when the multimodal data glaucoma classification program is executed by a processor, the steps of the multimodal data glaucoma classification method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Glaucoma detection method, medium and system based on feature fusion
CN121304661A
Glaucoma detection method, medium and system based on feature fusion
CN121304661B
Multi-mode glaucoma detection method, medium and system
CN121304676A
Multimodal glaucoma detection method, medium, and system
CN121304676B