Visual conversion system, method and equipment for auxiliary diagnosis of eye diseases and storage medium

Through multimodal preprocessing, efficient feature extraction and dynamic expert scheduling of the visual conversion system, the existing AI system for ophthalmology diagnosis is solved, and the multi-task parallel output and efficient diagnosis are realized to meet the needs of clinical multimodal diagnosis.

CN120496813APending Publication Date: 2025-08-15南通诺瞳奕目医疗科技有限公司 +1

Patent Information

Application Number
CN202510786258.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing AI system for ophthalmology diagnosis has single functions, high data dependence, large computing resources demand, and weak interpretability, which cannot meet the clinical needs of multimodal and multitask diagnosis.

Method used

The visual transformation system is adopted, including the input processing module, the visual transformer backbone, the hybrid expert layer and the decoder head, and the ophthalmic images are processed through image block extraction and embedding, multi-head self-attention and feedforward network, combined with the sparse hybrid expert layer and the top K-gated network, and dynamically select the expert neural network to achieve multimodal and multitasking diagnosis.

Benefits of technology

It realizes efficient processing of multimodal eye disease image data, reduces computing resource requirements, improves diagnostic accuracy and efficiency, supports multi-task parallel output, adapts to new clinical needs, and improves the long-term practicality and adaptability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496813A_ABST
    Figure CN120496813A_ABST
Patent Text Reader

Abstract

The invention discloses a visual conversion system, method and equipment for eye disease auxiliary diagnosis and a storage medium, and solves the problems of single function, high data dependence, large computing resource demand and weak interpretability in the field of eye disease auxiliary diagnosis in the prior art. The system comprises an input processing module, a visual transformer trunk, a hybrid expert layer, a decoder head and an output generation module. According to the visual conversion system for eye disease auxiliary diagnosis, various ophthalmology imaging modes can be processed, multiple tasks can be executed, and the input processing module, the visual transformer trunk, the mixed expert layer, the decoder head and the output generation module in the system cooperate to reduce calculation requirements and maintain accuracy at the same time; the expert is dynamically selected according to the complexity of the input image, it is ensured that different inputs can obtain optimal feature processing, accuracy and efficiency are both considered, a decoder structure can be customized according to the features of different tasks, and the output precision of each task is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an artificial intelligence system for medical image analysis, and in particular to a visual conversion system, method, device and storage medium for auxiliary diagnosis of eye diseases. Background Art

[0002] Visual impairment and eye diseases are major issues affecting public health worldwide. According to statistics, millions of people around the world are at risk of vision impairment or even blindness due to various eye diseases, which not only seriously reduces the quality of life of patients, but also places a huge burden on the medical system. Early detection and accurate diagnosis are the key to blocking the progression of eye diseases and preserving visual function. However, the complexity of eye diseases (such as the coexistence of multiple types and overlapping clinical manifestations) and the uneven distribution of medical resources (especially the shortage of ophthalmologists) make it difficult for traditional manual diagnosis models to meet efficient and accurate clinical needs. The innovation of medical imaging technology has provided important support for the diagnosis of eye diseases. The current mainstream ophthalmic imaging modes include fundus photography, optical coherence tomography (OCT, providing high-resolution images of the retinal layered structure), fluorescein angiography (FFA, showing abnormal leakage or blockage of retinal blood vessels), etc. These technologies can intuitively present subtle lesions in intraocular tissues.

[0003] The rise of artificial intelligence (AI) and deep learning technology has provided a new path for the automation of eye disease diagnosis. Deep learning models (such as convolutional neural networks (CNNs)) have shown potential in single-modality medical image analysis. For example, they can be trained to automatically identify and grade the characteristics of diabetic retinopathy in fundus images. However, existing AI systems for eye disease diagnosis still have significant limitations: they are single-functional, with most systems targeting only a single imaging mode or a single task, making it difficult to meet the actual needs of multimodal data fusion analysis or joint diagnosis of multiple diseases in clinical practice; they are highly data-dependent, with the acquisition of high-quality annotated data being expensive and time-consuming, limiting the model's generalization capabilities in rare eye diseases or multi-center scenarios.

[0004] In recent years, the emergence of transformer architectures has driven rapid technological development. Visual transformers (ViTs) have demonstrated excellent performance in image analysis tasks and have become a powerful tool for processing image data. However, when applying ViTs to the specific scenario of multimodal, multi-task eye disease diagnosis, many technical bottlenecks have been encountered. For example, a prominent challenge in developing AI systems for eye disease diagnosis is obtaining large amounts of labeled training data, which is often costly, time-consuming, and labor-intensive in the medical field. At the same time, the huge computing resources required to run complex AI models have become a major obstacle in resource-constrained medical environments.

[0005] In summary, existing technologies in the field of auxiliary diagnosis of eye diseases suffer from limitations such as limited functionality, high data dependency, high computational resource requirements, and weak interpretability, failing to meet clinical demands for efficient, accurate, and universal diagnostic tools. Therefore, there is an urgent need to develop a visual conversion system and method that can integrate multimodal eye disease imaging data, support multi-task analysis, lower data and computational barriers, and provide reliable interpretability to improve the accessibility and accuracy of auxiliary diagnosis of eye diseases. Summary of the Invention

[0006] (1) Technical problems solved

[0007] In response to the shortcomings of the existing technology, the present invention provides a visual conversion system, method, device and storage medium for auxiliary diagnosis of eye diseases, which solves the problems of the existing technology in the field of auxiliary diagnosis of eye diseases, such as single function, high data dependence, large computing resource requirements and weak interpretability.

[0008] (2) Technical solution

[0009] To achieve the above objectives, the present invention is implemented through the following technical solutions:

[0010] In one aspect, the present invention provides a visual conversion system for aiding diagnosis of eye diseases, the system comprising:

[0011] an input processing module for receiving and preprocessing input ophthalmic images;

[0012] A visual transformer backbone that processes the prepared input ophthalmic images and generates a set of image labels;

[0013] A mixture of experts layer that receives image tags generated by the visual transformer backbone and selectively activates an expert neural network;

[0014] at least one decoder head configured to perform a specific eye disease diagnosis task using outputs of the selectively activated expert neural networks in the hybrid expert layer;

[0015] Output generation module, which processes the results from the task-specific decoder head and generates the final diagnostic output.

[0016] The input processing module includes:

[0017] An image block extractor, which is used to segment the input ophthalmic image into multiple blocks;

[0018] Block Embedding Module,The block embedding module is used to convert 16x16 image blocks into 768-dimensional,embeddings.

[0019] The visual transformer backbone consists of multi-head self-attention and feed-forward networks.

[0020] The hybrid expert layer includes:

[0021] An expert pool, wherein the expert pool includes a plurality of expert neural networks;

[0022] Router,expert pool contains 16 expert neural networks.

[0023] The decoder head includes a disease classification head, a prognosis head and a biomarker detection head.

[0024] At least one router is configured to activate a first K subset of the expert neural networks, where K is less than a total number of expert neural networks in the expert pool.

[0025] In another aspect, the present invention further provides a visual conversion method for auxiliary diagnosis of eye diseases, the method comprising:

[0026] S1, receiving an input ophthalmic image. It should be noted that after receiving the input ophthalmic image, the input processing module performs initial preprocessing on the received image;

[0027] S2, generating a set of image tags from the input ophthalmic image using a visual transformer backbone, wherein when the visual transformer backbone processes the input image to generate the set of image tags, the image tags represent key features of the ophthalmic image;

[0028] S3, selectively activating a subset of expert neural networks in the hybrid expert layer based on the image label by at least one router;

[0029] S4, performing a specific eye disease diagnosis task using the output of the selectively activated expert neural network by at least one task-specific decoder head.

[0030] The method for receiving an input ophthalmic image comprises:

[0031] S101, loading input ophthalmic image;

[0032] S102, performing block embedding, the image block extractor divides the input image into blocks, and the block embedding module converts the divided blocks into high-dimensional embeddings;

[0033] S103, the visual transformer backbone processes the embedding block, where multi-head self-attention and feed-forward network components work together to convert the embedding block into a set of image tags;

[0034] S104 , routing the image tags through the hybrid expert layer, where the router selectively activates a subset of expert neural networks from the expert pool based on features of the image tags.

[0035] On the other hand, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the visual conversion method for auxiliary diagnosis of eye diseases are implemented.

[0036] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the visual conversion method for auxiliary diagnosis of eye diseases.

[0037] (3) Beneficial effects

[0038] The present invention provides a visual conversion system, method, device, and storage medium for auxiliary diagnosis of eye diseases. Compared with the existing technology, it has the following advantages:

[0039] 1. In an embodiment of the present invention, a visual conversion system for auxiliary diagnosis of eye diseases can process various ophthalmic imaging modalities and perform multiple tasks. The input processing module, visual transformer backbone, hybrid expert layer, decoder head, and output generation module in the system work together to reduce computing requirements while maintaining accuracy. Experts are dynamically selected based on the complexity of the input image to ensure that different inputs can obtain optimal feature processing, taking into account both accuracy and efficiency. The decoder structure can be customized according to the characteristics of different tasks to ensure the output accuracy of each task.

[0040] 2. In an embodiment of the present invention, a visual conversion system for auxiliary diagnosis of eye diseases further provides an input processing module, which is composed of an image block extractor and a block embedding module. The setting of the image block extractor makes the image processing unit smaller, which helps to perform more detailed and accurate processing of the image later, and can capture smaller details and features in the image, laying the foundation for subsequent accurate diagnosis of eye diseases. Compared with directly processing the entire image, it can reduce the complexity of image processing, improve processing efficiency, enable the system to pre-process the image faster, reduce the patient's waiting time, and improve diagnostic efficiency. Through the embedding conversion process of the block embedding module, the original pixel information of the image block is converted into a higher-dimensional embedding representation, which can more fully explore and express the potential features in the image block, provide a richer information basis for the subsequent extraction of relevant features from the input ophthalmic image, and help improve the accuracy of eye disease diagnosis.

[0041] 3. In the embodiments of the present invention, the visual conversion system for auxiliary diagnosis of eye diseases has the advantages of multimodal data compatibility, efficient multi-task processing, low computing resource requirements, and strong clinical scalability. It ultimately solves the pain points of "single data, fragmented tasks, limited resources, and weak interpretability" in traditional eye disease diagnosis, and provides a highly accessible and accurate AI-assisted diagnostic tool for resource-poor areas.

[0042] 4. In this embodiment of the present invention, the system supports the input of multiple ophthalmic imaging modalities, including fundus images, OCT, and FFA, overcoming the limitation of traditional diagnostic systems that can only process single-modality data. By using a visual transformer backbone to extract unified features from multimodal data, it enables fusion analysis of cross-modal information, providing doctors with a more comprehensive basis for lesion assessment and avoiding missed diagnoses or misdiagnoses that may result from a single modality.

[0043] 5. In an embodiment of the present invention, the system adopts a sparse mixed expert layer and a top K-gated network design to activate only the subset of experts that are most relevant to the current input or task. This mechanism significantly reduces the amount of redundant computation (e.g., simple cases only require a small number of experts to process), while maintaining diagnostic accuracy, significantly reducing energy consumption and processing time, and providing two variants: multi-gate MoE-ViT (independent router for each task + shared expert pool) and task-conditional MoE-ViT (shared router + task embedding driver). There is no need to add additional routers for new tasks, and functional expansion can be achieved by simply expanding the expert network or adjusting the task embedding. This design significantly reduces the cost of system upgrades and maintenance, enabling it to quickly adapt to new clinical eye disease diagnosis needs (such as rare diseases or multimodal joint diagnosis tasks), and improving the long-term practicality and adaptability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0045] Figure 1 The diagram shows the architecture of a visual conversion system for aiding diagnosis of eye diseases.

[0046] Figure 2 Shown is an architectural diagram of the input processing module.

[0047] Figure 3 Figure 2 shows the architecture diagram of the multi-gate visual transformer.

[0048] Figure 4 A flowchart of the implementation of a visual conversion method for auxiliary diagnosis of eye diseases is shown.

[0049] Figure 5 A schematic diagram of the implementation flow of the method for receiving input ophthalmic images is shown.

[0050] Figure 6 A diagram showing the key steps of the visual conversion system process for auxiliary diagnosis of eye diseases. DETAILED DESCRIPTION

[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0052] The embodiments of the present application solve the problems of the existing technology in the field of auxiliary diagnosis of eye diseases, such as single function, high data dependence, large computing resource requirements, and weak interpretability, by providing a visual conversion system, method, device and storage medium for auxiliary diagnosis of eye diseases. It realizes dynamic selection of experts according to the complexity of the input image, ensures that different inputs can obtain optimal feature processing, takes into account both accuracy and efficiency, and can customize the decoder structure according to the characteristics of different tasks to ensure the output accuracy of each task.

[0053] The technical solution in the embodiments of the present application is to solve the above technical problems, and the overall idea is as follows:

[0054] This application achieves efficient, accurate and flexible conversion from input to diagnostic results through a full-process design of "multimodal preprocessing, efficient feature extraction, dynamic expert scheduling, and multi-task parallel output", which solves the pain points of "single data, fragmented tasks, and limited resources" in traditional eye disease diagnosis.

[0055] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0056] Example 1:

[0057] In recent years, the emergence of transformer architectures has driven rapid technological development. Visual transformers (ViTs) have demonstrated excellent performance in image analysis tasks and have become a powerful tool for processing image data. However, when applying ViTs to the specific scenario of multimodal, multi-task eye disease diagnosis, many technical bottlenecks have been encountered. For example, a prominent challenge in developing AI systems for eye disease diagnosis is obtaining large amounts of labeled training data, which is often costly, time-consuming, and labor-intensive in the medical field. At the same time, the huge computing resources required to run complex AI models have become a major obstacle in resource-constrained medical environments.

[0058] In the first aspect, the present invention first proposes a visual conversion system for auxiliary diagnosis of eye diseases. Figure 1 The following is an architecture diagram of a visual conversion system for aiding diagnosis of eye diseases, comprising:

[0059] An input processing module 100 is configured to receive and pre-process an input ophthalmic image. In some cases, the input ophthalmic image may be a fundus image, an optical coherence tomography (OCT) image, or a fluorescein angiography (FFA) image. The input processing module 100 may prepare the image data for further processing by subsequent components of a visual conversion system for aiding ophthalmic disease diagnosis;

[0060] A visual transformer backbone 200, which is used to process the prepared input ophthalmic image and generate a set of image tags. The visual transformer backbone 200 can apply a self-attention mechanism and other transformations to extract relevant features from the input image;

[0061] A hybrid expert layer 300 comprises multiple expert neural networks and at least one router 320. The hybrid expert layer 300 is configured to receive image labels generated by the visual transformer backbone 200. In some cases, the routers within the hybrid expert layer 300 can selectively activate a subset of the expert neural networks based on the received image labels. This selective activation may enable efficient processing of image data by utilizing only the most relevant expert neural networks for a given input.

[0062] At least one decoder head 400 includes a plurality of specific tasks, and each task-specific decoder head 400 is configured to perform a specific eye disease diagnosis task using the output of the selectively activated expert neural network in the hybrid expert layer 300. These specific tasks may include, but are not limited to, disease classification, prognosis determination, or biomarker detection.

[0063] Output generation module 500, which is used to process the results from the decoder head 400 for a specific task and generate a final diagnosis output. The output generation module 500 may format and present the diagnosis results in a manner suitable for interpretation or further analysis by a medical professional.

[0064] It should be noted that the components of the visual transformation system for auxiliary diagnosis of eye diseases can work together to process the input ophthalmic images through various analysis stages, and use the advantages of visual transformers and hybrid expert architecture to efficiently perform multiple eye disease diagnosis tasks. The visual transformation system for auxiliary diagnosis of eye diseases can be briefly described as the EyeMOE system architecture.

[0065] In an embodiment of the present invention, the visual conversion system for auxiliary diagnosis of eye diseases can process various ophthalmic imaging modalities and perform multiple tasks. The input processing module 100, the visual transformer backbone 200, the hybrid expert layer 300, the decoder head 400, and the output generation module 500 in the system work together to reduce computing requirements while maintaining accuracy. Experts are dynamically selected according to the complexity of the input image to ensure that different inputs can obtain optimal feature processing, taking into account both accuracy and efficiency, and can customize the decoder structure according to the characteristics of different tasks to ensure the output accuracy of each task.

[0066] In an embodiment of the present invention, an input processing module 100 is provided. Figure 2 Schematic diagram of the input processing module 100 is shown. The input processing module 100 includes:

[0067] An image block extractor 110 , the image block extractor 110 is used to divide the input ophthalmic image into a plurality of blocks, and the image block extractor 110 may have a block dimension of 16×16 pixels;

[0068] After the image block extractor 110 segments the input ophthalmic image into blocks, the block embedding module 120 processes these blocks. The block embedding module 120 is used to convert the 16x16 image blocks into a 768-dimensional embedding. This conversion process may transform the spatial information of the image blocks into a format suitable for further processing by subsequent components of the system.

[0069] The embedding blocks generated by the block embedding module 120 can be passed to the visual transformer backbone 200 for further processing. The visual transformer backbone 200 then applies various transformations to the embedding blocks to extract relevant features from the input ophthalmic image. By pre-processing the input ophthalmic image through block extraction and embedding, the input processing module 100 may prepare the image data for efficient analysis by subsequent components of the system (including the mixture of experts layer 300 and the task-specific decoder head 400).

[0070] In embodiments of the present invention, the visual transformer backbone 200 may be configured to process the embedding blocks generated by the block embedding module 120. In some cases, the visual transformer backbone 200 may be a data-efficient visual transformer (DeiT). The visual transformer backbone 200 may be designed to efficiently extract relevant features from the input ophthalmic image while minimizing computational requirements.

[0071] In an embodiment of the present invention, an input processing module 100 is provided. The input processing module 100 is composed of an image block extractor 110 and a block embedding module 120. The setting of the image block extractor 110 makes the image processing unit smaller, which helps to perform more detailed and accurate processing of the image later, and can capture smaller details and features in the image, laying the foundation for the subsequent accurate diagnosis of eye diseases. Compared with directly processing the entire image, it can reduce the complexity of image processing, improve processing efficiency, enable the system to pre-process the image faster, reduce the patient's waiting time, and improve diagnostic efficiency. Through the embedding conversion process of the block embedding module 120, the original pixel information of the image block is converted into a higher-dimensional embedding representation, which can more fully explore and express the potential features in the image block, provide a richer information basis for the subsequent extraction of relevant features from the input ophthalmic image, and help improve the accuracy of eye disease diagnosis.

[0072] The visual transformer backbone 200 includes multiple components that work together to process the embedding block. The visual transformer backbone 200 includes a multi-head self-attention 210 and a feed-forward network 220. These components may be arranged in a series of encoder iterations to gradually transform the input data.

[0073] It should be noted that the visual transformer backbone 200 may use 12 encoder iterations. Each encoder iteration may involve processing data through multi-head self-attention 210 and feed-forward network 220. This iterative process may allow the visual transformer backbone 200 to capture complex relationships and patterns in the input ophthalmic images.

[0074] Multi-head self-attention 210 can be configured to analyze the relationships between different parts of the input data. By applying multiple attention mechanisms in parallel, multi-head self-attention 210 is able to capture various types of dependencies and patterns within the embedding block.

[0075] After the multi-head self-attention 210, the attention-processed data is further processed by the feed-forward network 220. The feed-forward network 220 is able to apply nonlinear transformations to the data, allowing the visual transformer backbone 200 to model complex functions and extract high-level features from the input.

[0076] By repeatedly applying multi-head self-attention 210 and feed-forward networks 220 in encoder iterations, the visual transformer backbone 200 is able to generate a set of image tags that may represent a high-level encoding of the input ophthalmic image, capturing relevant features and patterns that can be used for subsequent diagnostic tasks.

[0077] The image labels generated by the visual transformer backbone 200 can then be passed to the mixture of experts layer 300 for further processing. By efficiently converting the input ophthalmic images into a set of information-rich image labels, the visual transformer backbone 200 plays a key role in preparing data for subsequent stages of the system, including the task-specific diagnosis process performed by the task-specific decoder head 400.

[0078] In an embodiment of the present invention, a visual transformer backbone 200 is provided. The visual transformer backbone 200 can process embedded blocks and generate image tags through the collaborative work of multiple components. These image tags represent high-level encodings of the input ophthalmic images, capturing relevant features and patterns that can be used for subsequent diagnostic tasks, and passing these image tags to the hybrid expert layer 300 for further processing. After the multi-head self-attention 210, the feedforward network 220 can further process the attention-processed data and apply nonlinear transformations to the data, allowing the visual transformer backbone 200 to model complex functions and extract high-level features from the input, which helps to convert the captured original features into more representative and discriminative high-level features, providing more powerful information support for subsequent diagnostic tasks.

[0079] In an embodiment of the present invention, multi-head self-attention 210 is added to the hybrid expert layer 300 to construct a non-pipelined multimodal + multi-task system rather than a traditional pipeline architecture. Its core advantage lies in the design of "dynamic collaboration + global perception", which breaks through the limitations of traditional architectures in modal fusion, task collaboration and interpretability, while significantly broadening the application boundaries of the visual transformer backbone 200 in medical AI. The sparse activation mechanism of MOE (only activating relevant experts) combined with the focus characteristics of attention (only focusing on key tokens) avoids the problem of repeated calculations in each stage of the pipeline architecture. For example, when processing multimodal inputs, the pipeline may need to perform full-image feature extraction for each modality separately, while the MOE+attention function processes multiple modalities simultaneously through a unified attention mechanism, and allocates computing resources only in key areas (such as lesions), significantly reducing inference delays and meeting clinical real-time requirements.

[0080] In the embodiment of the present invention, a hybrid expert layer 300 is further provided. The hybrid expert layer 300 includes:

[0081] An expert pool 310, wherein the expert pool 310 includes a plurality of expert neural networks, each of which is used to process a feature or pattern of a type in image labeling;

[0082] Router 320, expert pool 310 contains 16 expert neural networks. This configuration may allow for diverse specialized processing capabilities within the hybrid expert layer 300;

[0083] In a Mixture of Experts (MOE) system, expert load imbalance and sudden failures (such as performance degradation and parameter collapse caused by overload) are core issues affecting system robustness. Traditional MOE routers typically dynamically select experts based solely on input features (e.g., Top-k routing), but lack active monitoring of expert load status and performance levels. This can easily lead to some experts "crashing" due to overload (e.g., latency surges and accuracy plummets). Meanwhile, low-load experts are underutilized, ultimately reducing overall system efficiency and reliability. Router 320 can be configured to selectively activate a subset of expert neural networks from the expert pool 310 based on input image tags. This selective activation enables sparse activation of expert neural networks, allowing for efficient utilization of computing resources. Furthermore, in this embodiment, to achieve expert fault handling and load balancing, expert status indicators (load indicators, performance indicators), threshold triggering, and dynamic response mechanisms are predefined. Combined with a dynamic routing probability adjustment strategy and a hierarchical fault response mechanism, the system's stability and reliability are significantly improved. This design not only avoids system failures caused by expert overload and collapse, but also optimizes computing resource utilization through load balancing.

[0084] The routers 320 can activate the top K subsets of the expert neural networks, and at least one router 320 is configured to activate the top K subsets of the expert neural networks, where K is less than the total number of expert neural networks in the expert pool 310. This approach may allow the hybrid expert layer 300 to focus on the most relevant expert neural networks for a given input, potentially improving both efficiency and accuracy.

[0085] In embodiments of the present invention, the selection count K can be set to different values based on the specific requirements of the system. In some cases, the selection count K can be 4. For 16 experts and a selection count of 4, router 320 may only activate 25% of the available expert neural networks for each input, thereby reducing computational overhead while maintaining diagnostic accuracy. Specifically, modality experts are defined during expert neural network training. By defining specific modality experts, such as fundus focus and OCT focus, each expert neural network can conduct in-depth analysis of image features specific to a particular modality, achieving accurate diagnosis. For example, fundus-focused experts can meticulously capture subtle vascular changes and retinal color abnormalities in fundus images, providing a reliable basis for fundus disease diagnosis. Furthermore, the expert neural network incorporates a self-organizing network (SON) mechanism, allowing experts to autonomously optimize their processing capabilities for specific features through network training, further enhancing the system's long-term practicality and generalization. The SON mechanism monitors the distribution of features labeled in input images (e.g., newly identified pathological features of rare eye diseases) and automatically adjusts the expert's network parameters (e.g., convolution kernel weights and attention head orientation) to more accurately capture emerging feature patterns (e.g., specific lesions of rare hereditary eye diseases) without the need for manual redesign of the expert structure. Through performance evaluation during training (e.g., classification accuracy and task loss), the SON mechanism identifies and eliminates inefficient experts (e.g., redundant experts with minimal contribution to the current dataset) while strengthening the parameters of highly contributing experts (e.g., increasing their routing weights), achieving "self-streamlining" of the expert pool and reducing computational redundancy.

[0086] It should be noted that the expert neural network also includes a cache network. The expert neural network pre-trains the visual transformer backbone and hybrid expert layer 300 on a large-scale, unlabeled ophthalmology dataset (self-supervision) to reduce dependency on labeled data. Furthermore, each network routing path is defined for a specific modality expert, thereby reducing latency.

[0087] When training the expert neural network, the load-importance balance loss (LIB Loss) weight is defined as 0.01, wherein LIB Loss (weight = 0.01) is the key training mechanism. By mathematically constraining the expert load distribution and task matching, it directly solves the core problems of expert collapse and load imbalance. It should be noted that the load-importance balance loss (LIB Loss) is the core component of VisionMOE. By constraining the expert load distribution and task importance, it achieves fair expert utilization (avoiding some experts from being idle or overloaded) and stable multi-task training (preventing low-performance experts from outputting erroneous results). This embodiment follows this mechanism and optimizes parameters for ophthalmic scenarios to overcome the problems of uneven expert load, performance collapse caused by expert overload, and insufficient stability of multi-task training in the prior art multi-modal multi-task.

[0088] The process of routing image labels through the mixture of experts layer 300 may involve several steps. First, the router 320 may analyze the incoming image labels to determine which expert neural networks are best suited to process the current input. Based on this analysis, the router 320 may then selectively activate a subset of expert neural networks from the expert pool 310.

[0089] The activated expert neural networks can then process the image tags, applying their expertise to extract relevant features or patterns. The outputs from these activated expert neural networks may be combined or aggregated in some way and then passed to the task-specific decoder head 400 for further processing.

[0090] By implementing sparse activation and selective routing, the mixture of experts layer 300 enables the system to efficiently process a variety of ophthalmic images while adapting to the specific characteristics of each input. This approach may enable the system to perform multiple eye disease diagnosis tasks with improved efficiency and accuracy.

[0091] It should be noted that, in the embodiment of the present invention, the task-specific decoder head 400 includes a disease classification head 410, a prognosis head 420, and a biomarker detection head 430. Each head can be designed to perform a specific type of eye disease diagnosis task.

[0092] In embodiments of the present invention, disease classification head 410 can classify various eye diseases based on the processed image data. Disease classification head 410 analyzes the output from mixture of experts layer 300 to identify patterns or features indicative of specific eye diseases. In some cases, disease classification head 410 can output a probability distribution of multiple possible eye diseases or conditions.

[0093] Prognostic header 420 is used to provide a disease prognosis or progression prediction. Prognostic header 420 is used to analyze processed image data to estimate the likelihood of disease progression or predict a patient's potential outcome. In some cases, prognostic header 420 can also generate a time-based prediction or risk assessment for various eye diseases.

[0094] Biomarker detection head 430 is configured to detect systemic biomarkers from ophthalmic images. Biomarker detection head 430 is used to analyze the output from hybrid expert layer 300 to identify indicators of systemic health conditions that may be visible in the eye. In some cases, biomarker detection head 430 is capable of detecting markers associated with diabetes, hypertension, or other systemic diseases that may affect the eye.

[0095] In this embodiment, each task-specific decoder head 400 may generate a diagnostic output based on the output of the selectively activated expert neural network in the mixture of experts layer 300. The diagnostic output may include disease classification, prognosis prediction, or detected biomarkers, depending on the specific task performed by each decoder head.

[0096] In some cases, task-specific decoder heads 400 may run in parallel, allowing the system to perform multiple diagnostic tasks simultaneously on a single input ophthalmic image. This parallel processing capability may enable comprehensive eye disease diagnosis from a single image analysis.

[0097] The output from the task-specific decoder head 400 may be passed to the output generation module 500 for further processing and presentation. By utilizing specialized decoder heads 400 for different diagnostic tasks, the system may provide comprehensive analysis of ophthalmic images, potentially improving the accuracy and efficiency of eye disease diagnosis.

[0098] The system may include multiple variants to accommodate different processing requirements and mission configurations. Two such variants are multi-gate variants and mission-conditional variants, each offering unique advantages in terms of flexibility and efficiency.

[0099] In this embodiment of the present invention, the system supports the input of multiple ophthalmic imaging modalities, including fundus images, OCT, and FFA, overcoming the limitation of traditional diagnostic systems that can only process data from a single modality. By using a visual transformer backbone to extract unified features from multimodal data, it enables fusion analysis of cross-modal information, providing doctors with a more comprehensive basis for lesion assessment and avoiding missed diagnoses or misdiagnoses that may result from a single modality.

[0100] The system uses a sparse mixed expert layer and a top K-gated network design to activate only the subset of experts most relevant to the current input or task. This mechanism significantly reduces redundant computation (e.g., simple cases only require a small number of experts to process), while maintaining diagnostic accuracy while significantly reducing energy consumption and processing time. It also provides two variants: multi-gate MoE-ViT (independent router for each task + shared expert pool) and task-conditional MoE-ViT (shared router + task embedding drive). There is no need to add additional routers for new tasks, and functionality can be expanded by simply expanding the expert network or adjusting the task embedding. This design significantly reduces the cost of system upgrades and maintenance, enabling it to quickly adapt to new clinical eye disease diagnosis needs (such as rare diseases or multimodal joint diagnosis tasks), improving the long-term practicality and adaptability of the system.

[0101] In the embodiment of the present invention, the visual transformation system for auxiliary diagnosis of eye diseases further includes a multi-door visual transformer 600, Figure 3 : shows an architectural diagram of a multi-gate vision transformer 600, which represents one implementation of a multi-gate variant.

[0102] The multi-gate vision transformer 600 may include a shared vision tag 602 that serves as an input to multiple parallel routers.

[0103] In a multi-gate variant, the system employs multiple parallel routers, each dedicated to a specific task, and the multi-gate visual transformer 600 includes a first task router 604, a second task router 606, and a third task router 608. Each task router is configured to process the shared visual mark 602 independently.

[0104] The multi-gate variant includes a shared expert pool 610 of multiple expert modules. In some cases, the shared expert pool 610 includes an initial expert module 612 and a terminal expert module 616, as well as other expert modules, such as expert module 614. Each parallel task router may select the top K expert subsets from the shared expert pool 610 based on the input labels and specific task requirements.

[0105] After expert selection and processing, the output may be directed to a task decoder 622 for a specific task. The multi-gate visual transformer 600 includes a first task decoder 618 associated with the first task router 604, a second task decoder 620 associated with the second task router 606, and a third task decoder 622 associated with the third task router 608. Each task decoder is capable of performing a specific eye disease diagnosis task using the output from the selected expert module.

[0106] The multi-gate variant can provide advantages in situations where different tasks require different routing strategies. By employing separate parallel routers for each task, the system may optimize the expert selection process for each specific diagnostic task, potentially improving accuracy and efficiency.

[0107] In addition to the multi-gate variant, the system may also implement a task-conditional variant. In the task-conditional variant, the system may utilize a single shared router that receives both the concatenated label and the task embedding.

[0108] The task-conditional variant may employ a lightweight 64-dimensional task embedding. This task embedding can be concatenated with the visual labels generated by the visual transformer backbone 200. The combined labels and task embeddings can then be processed by a shared router to determine which expert modules to activate.

[0109] By using task embeddings, task-conditional variants allow for scalable expansion of system capabilities without adding additional routers. The shared router is able to learn to interpret task embeddings and select the most appropriate expert module for each specific task.

[0110] Task-conditional variants offer advantages in terms of efficiency and flexibility. Using a single shared router may reduce computational overhead compared to multi-gate variants, while the task embedding approach may allow new diagnostic tasks to be easily added without major architectural changes.

[0111] Both multi-gate and task-conditional variants of the system can utilize the hybrid expert layer 300 to dynamically route computation to the most relevant expert module for each input and task. This approach has the potential to enable efficient processing of ophthalmic images across multiple diagnostic tasks, potentially improving the overall performance and adaptability of the system.

[0112] Example 2:

[0113] In a second aspect, the present invention also provides a visual conversion method for auxiliary diagnosis of eye diseases. Figure 4 A flowchart of a visual conversion method for aiding diagnosis of eye diseases is shown, the method comprising:

[0114] S1, receiving an input ophthalmic image. It should be noted that after receiving the input ophthalmic image, the input processing module 310 can perform initial pre-processing on the received image;

[0115] S2, generating a set of image tags from the input ophthalmic image using the visual transformer backbone 200, wherein when the visual transformer backbone 200 processes the input image to generate the set of image tags, the image tags represent key features of the ophthalmic image;

[0116] S3, at least one router 320 selectively activates a subset of the expert neural networks in the hybrid expert layer 300 based on the image tags. It should be noted that the routers 320 in the hybrid expert layer 300 can analyze the image tags and determine which expert neural networks are most suitable for processing the current input;

[0117] S4, at least one task-specific decoder head 400 uses the output of the selectively activated expert neural network to perform a specific eye disease diagnosis task.

[0118] In an embodiment of the present invention, a visual conversion method for auxiliary diagnosis of eye diseases is provided. The input preprocessing module supports the input of multiple ophthalmic imaging modalities such as fundus images, OCT, FFA, etc., and standardized preprocessing eliminates data differences between different devices and formats to ensure that the subsequent visual transformer backbone can stably extract features, which solves the limitation of traditional systems that only support a single modality. The visual transformer backbone can capture the global correlation of ophthalmic images and focus on local details through a multi-head self-attention mechanism and multi-round encoder iterations. The generated image tags representing key features contain multi-dimensional semantic information, providing high-quality input for the dynamic calculation of the subsequent hybrid expert layer. The multiple specific task decoder heads of the conversion method steps can be executed in parallel and output multi-dimensional diagnostic results at the same time, rather than the serial output of the traditional single-task model, realizing efficient, accurate and flexible conversion from input to diagnostic results, which solves the pain points of "single data, fragmented tasks, and limited resources" in traditional eye disease diagnosis.

[0119] It should be noted that the embodiment of the present invention also provides a method for receiving an input ophthalmic image. Figure 5 A schematic diagram of a process flow for implementing a method for receiving an input ophthalmic image is shown. The method for receiving an input ophthalmic image comprises:

[0120] S101, loading input ophthalmic image;

[0121] S102, performing block embedding, the image block extractor 110 divides the input image into blocks, and the block embedding module 120 converts the divided blocks into high-dimensional embeddings;

[0122] S103, the visual transformer backbone 200 processes the embedding block, and the multi-head self-attention 210 and feedforward network 220 components work together to convert the embedding block into a set of image tags;

[0123] S104 , routing the image tags through the hybrid expert layer 300 , the router 320 selectively activates a subset of expert neural networks from the expert pool 310 based on the features of the image tags.

[0124] In the embodiment of the present invention, Figure 6The key steps of the visual conversion system process for auxiliary diagnosis of eye diseases are shown. After receiving the input ophthalmic image in step 1, the system generates image tags in step 2 and sends these tags for further processing in step 3. Figure 6 Step 4 in FIG. 4 illustrates the selective activation of the expert neural network, corresponding to the operation of the hybrid expert layer 300. In step 5, the selected expert neural network is used to process the image labeling. The final stage of the method operation is in Figure 6 This is represented in steps 7 and 8 of . In step 7, the task-specific decoder head 400 can use the output from the selectively activated expert neural network to perform a specific eye disease diagnosis task. This step may involve the disease classification head 410, the prognosis head 420, and the biomarker detection head 430 working in parallel to analyze the processed data. Finally, in step 8, the diagnostic results for multiple eye disease tasks are output. The output generation module 500 compiles and formats these results for presentation to medical professionals or further analysis.

[0125] The systems and methods described in embodiments of the present invention may be implemented using computer hardware and software components, including a processor that executes instructions stored on a non-transitory computer-readable storage medium. When executed, these instructions may cause the processor to perform the various steps of the multi-task eye disease diagnosis method.

[0126] This approach may allow for more efficient and accurate diagnosis of multiple eye conditions from a single input image, potentially improving patient care and resource utilization in ophthalmology practices. The described system may be particularly useful in settings with limited specialist care resources, as it may be able to perform comprehensive analysis of ophthalmic images across multiple diagnostic tasks.

[0127] It can be understood that the visual conversion method for auxiliary diagnosis of eye diseases provided in the embodiment of the present invention corresponds to the above-mentioned visual conversion system for auxiliary diagnosis of eye diseases. The explanations, examples, beneficial effects, etc. of the relevant contents can refer to the corresponding contents in the visual conversion system for auxiliary diagnosis of eye diseases, and will not be repeated here.

[0128] Example 3:

[0129] In a third aspect, the present invention further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the computer program implements the steps of the visual conversion method for auxiliary diagnosis of eye diseases described in any one of the above embodiments and preferred embodiments, wherein the method mainly comprises:

[0130] S1, receiving an input ophthalmic image. It should be noted that after receiving the input ophthalmic image, the input processing module 310 can perform initial pre-processing on the received image;

[0131] S2, generating a set of image tags from the input ophthalmic image using the visual transformer backbone 200, wherein when the visual transformer backbone 200 processes the input image to generate the set of image tags, the image tags represent key features of the ophthalmic image;

[0132] S3, at least one router 320 selectively activates a subset of the expert neural networks in the hybrid expert layer 300 based on the image tags. It should be noted that the routers 320 in the hybrid expert layer 300 can analyze the image tags and determine which expert neural networks are most suitable for processing the current input;

[0133] S4, at least one task-specific decoder head 400 uses the output of the selectively activated expert neural network to perform a specific eye disease diagnosis task.

[0134] It is understandable that the computer device provided in the embodiment of the present invention corresponds to the above-mentioned visual conversion system and method for auxiliary diagnosis of eye diseases. The explanation, examples, beneficial effects, etc. of the relevant contents can refer to the corresponding contents in the visual conversion system and method for auxiliary diagnosis of eye diseases, and will not be repeated here.

[0135] Example 4:

[0136] In a fourth aspect, the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to perform the steps of the visual conversion method for auxiliary diagnosis of eye diseases described in any of the above embodiments and preferred embodiments, wherein the method mainly includes:

[0137] S1, receiving an input ophthalmic image. It should be noted that after receiving the input ophthalmic image, the input processing module 310 can perform initial pre-processing on the received image;

[0138] S2, generating a set of image tags from the input ophthalmic image using the visual transformer backbone 200, wherein when the visual transformer backbone 200 processes the input image to generate the set of image tags, the image tags represent key features of the ophthalmic image;

[0139] S3, at least one router 320 selectively activates a subset of the expert neural networks in the hybrid expert layer 300 based on the image tags. It should be noted that the routers 320 in the hybrid expert layer 300 can analyze the image tags and determine which expert neural networks are most suitable for processing the current input;

[0140] S4, at least one task-specific decoder head 400 uses the output of the selectively activated expert neural network to perform a specific eye disease diagnosis task.

[0141] It is understandable that the computer device provided in the embodiment of the present invention corresponds to the above-mentioned visual conversion system, method, and device for auxiliary diagnosis of eye diseases. The explanation, examples, beneficial effects, etc. of the relevant contents can refer to the corresponding contents in the visual conversion system, method, and device for auxiliary diagnosis of eye diseases, and will not be repeated here.

[0142] In summary, compared with the existing technology, the present invention has the following beneficial effects:

[0143] 1. In an embodiment of the present invention, a visual conversion system for auxiliary diagnosis of eye diseases can process various ophthalmic imaging modalities and perform multiple tasks. The input processing module 100, visual transformer backbone 200, hybrid expert layer 300, decoder head 400, and output generation module 500 in the system work together to reduce computing requirements while maintaining accuracy. Experts are dynamically selected according to the complexity of the input image to ensure that different inputs can obtain optimal feature processing, taking into account both accuracy and efficiency. The decoder structure can be customized according to the characteristics of different tasks to ensure the output accuracy of each task.

[0144] 2. In an embodiment of the present invention, an input processing module 100 is provided. The input processing module 100 is composed of an image block extractor 110 and a block embedding module 120. The setting of the image block extractor 110 makes the image processing unit smaller, which helps to perform more detailed and accurate processing of the image later. It can capture smaller details and features in the image, laying the foundation for the subsequent accurate diagnosis of eye diseases. Compared with directly processing the entire image, it can reduce the complexity of image processing, improve processing efficiency, enable the system to pre-process the image faster, reduce the patient's waiting time, and improve diagnostic efficiency. Through the embedding conversion process of the block embedding module 120, the original pixel information of the image block is converted into a higher-dimensional embedding representation, which can more fully explore and express the potential features in the image block, provide a richer information basis for the subsequent extraction of relevant features from the input ophthalmic image, and help improve the accuracy of eye disease diagnosis.

[0145] 3. In an embodiment of the present invention, a visual transformer backbone 200 is provided. The visual transformer backbone 200 can process embedded blocks and generate image tags through the collaborative work of multiple components. These image tags represent high-level encodings of the input ophthalmic images, capturing relevant features and patterns that can be used for subsequent diagnostic tasks, and passing these image tags to the hybrid expert layer 300 for further processing. After the multi-head self-attention 210, the feedforward network 220 can further process the attention-processed data and apply nonlinear transformations to the data, allowing the visual transformer backbone 200 to model complex functions and extract high-level features from the input, which helps to convert the captured original features into more representative and discriminative high-level features, providing more powerful information support for subsequent diagnostic tasks.

[0146] 4. In an embodiment of the present invention, a visual conversion method for auxiliary diagnosis of eye diseases is provided. The input preprocessing module supports the input of multiple ophthalmic imaging modalities such as fundus images, OCT, FFA, etc., and standardized preprocessing eliminates data differences between different devices and formats to ensure that the subsequent visual transformer backbone can stably extract features, solving the limitation of traditional systems that only support a single modality. The visual transformer backbone can capture the global correlation of ophthalmic images and focus on local details through a multi-head self-attention mechanism and multi-round encoder iterations. The generated image tags representing key features contain multi-dimensional semantic information, providing high-quality input for the dynamic calculation of the subsequent hybrid expert layer. The multiple task-specific decoder heads of the conversion method steps can be executed in parallel and output multi-dimensional diagnostic results at the same time, rather than the serial output of the traditional single-task model, realizing efficient, accurate and flexible conversion from input to diagnostic results, which solves the pain points of "single data, fragmented tasks, and limited resources" in traditional eye disease diagnosis.

[0147] 5. When training the expert neural network, the load-importance balance loss (LIB Loss) weight is defined as 0.01, where LIB Loss (weight = 0.01) is the key training mechanism. By mathematically constraining the expert load distribution and task matching, it directly solves the core problems of expert collapse and uneven load. It should be noted that the load-importance balance loss (LIB Loss) is the core component of VisionMOE. By constraining the expert load distribution and task importance, it achieves fair expert utilization (avoiding some experts from being idle or overloaded) and stable multi-task training (preventing low-performance experts from outputting erroneous results). This embodiment follows this mechanism and optimizes parameters for ophthalmic scenarios to overcome the problems of uneven expert load, performance collapse caused by expert overload, and insufficient stability of multi-task training in the prior art multi-modal multi-task.

[0148] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element.

[0149] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A visual conversion system for auxiliary diagnosis of eye diseases, characterized in that: The system includes: an input processing module for receiving and preprocessing input ophthalmic images; A visual transformer backbone that processes the prepared input ophthalmic images and generates a set of image labels; A mixture of experts layer that receives image tags generated by the visual transformer backbone and selectively activates an expert neural network; at least one decoder head configured to perform a specific eye disease diagnosis task using outputs of the selectively activated expert neural networks in the hybrid expert layer; Output generation module, which processes the results from the task-specific decoder head and generates the final diagnostic output.

2. The visual conversion system for auxiliary diagnosis of eye diseases according to claim 1, characterized in that: The input processing module includes: An image block extractor, which is used to segment the input ophthalmic image into multiple blocks; Block Embedding Module,The block embedding module is used to convert 16x16 image blocks into 768-dimensional,embeddings.

3. The visual conversion system for auxiliary diagnosis of eye diseases according to claim 1, characterized in that: The visual transformer backbone consists of multi-head self-attention and feed-forward networks.

4. The visual conversion system for auxiliary diagnosis of eye diseases according to claim 3, characterized in that: The hybrid expert layer includes: An expert pool, wherein the expert pool includes a plurality of expert neural networks; Router,expert pool contains 16 expert neural networks.

5. The visual conversion system for auxiliary diagnosis of eye diseases according to claim 3, characterized in that: The decoder head includes a disease classification head, a prognosis head and a biomarker detection head.

6. The visual conversion system for auxiliary diagnosis of eye diseases according to claim 1, characterized in that: At least one router is configured to activate a first K subset of the expert neural networks, where K is less than a total number of expert neural networks in the expert pool.

7. A visual conversion method for auxiliary diagnosis of eye diseases, characterized in that: The method comprises: S1, receiving an input ophthalmic image. It should be noted that after receiving the input ophthalmic image, the input processing module performs initial preprocessing on the received image; S2, generating a set of image tags from the input ophthalmic image using a visual transformer backbone, wherein when the visual transformer backbone processes the input image to generate the set of image tags, the image tags represent key features of the ophthalmic image; S3, selectively activating a subset of expert neural networks in the hybrid expert layer based on the image label by at least one router; S4, performing a specific eye disease diagnosis task using the output of the selectively activated expert neural network by at least one task-specific decoder head.

8. The visual conversion method for auxiliary diagnosis of eye diseases according to claim 7, characterized in that: The method for receiving an input ophthalmic image comprises: S101, loading input ophthalmic image; S102, performing block embedding, the image block extractor divides the input image into blocks, and the block embedding module converts the divided blocks into high-dimensional embeddings; S103, the visual transformer backbone processes the embedding block, where multi-head self-attention and feed-forward network components work together to convert the embedding block into a set of image tags; S104 , routing the image tags through the hybrid expert layer, where the router selectively activates a subset of expert neural networks from the expert pool based on features of the image tags.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the visual conversion method for auxiliary diagnosis of eye diseases as described in claim 7 or 8 are implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the visual conversion method for auxiliary diagnosis of eye diseases as described in claim 7 or 8 are implemented.

Citation Information

Patent Citations

  • Color fundus image micro-hemangioma hyperfine segmentation method based on Transform

    CN118537552A

  • Ophthalmic disease intelligent diagnosis and treatment management method and system based on large model

    CN119181495A

  • Image processing method and device, equipment and storage medium

    CN119579912A

  • Auxiliary diagnostic system, method, program product and device for heart valve para-marginal repair

    CN119601212A

  • Large-model hybrid expert system, construction method thereof and expert selection method

    CN119962579A

Cited By

  • Multi-layer lookup table image mapping method based on conditional awareness and hybrid expert mechanism

    CN121746538A

  • Multi-layer lookup table image mapping method based on conditional perception and hybrid expert mechanism

    CN121746538B