Transform-based ophthalmology image anomaly detection method and system
By using an ophthalmic image anomaly detection system based on the Transformer architecture, combined with EfficientNet-B3 feature extraction and multi-scale embedding layers, the system addresses the shortcomings of existing technologies in multi-scale feature recognition and model generalization ability, achieving high-precision anomaly detection and early lesion discovery in ophthalmic images.
Patent Information
- Application Number
- CN202511106451.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies have limitations in processing the multi-scale characteristics of ophthalmic images. They cannot effectively identify macroscopic lesions to microscopic cellular lesions, and their models have insufficient generalization ability, failing to adapt to changes in the feature distribution of multiple types of samples.
An ophthalmic image anomaly detection system based on the Transformer architecture is adopted. It combines the EfficientNet-B3 feature extraction module with a multi-scale embedding layer, establishes global pixel-level correlation through a self-attention mechanism, realizes in-depth mining of image features and multi-scale feature fusion, and uses a neighborhood mask attention mechanism for feature reconstruction to identify abnormal regions.
It significantly improves the accuracy and robustness of ophthalmic image anomaly detection, can accurately identify abnormal regions in multiple categories of samples, adapts to the needs of deep feature modeling of complex medical images, and improves the early lesion detection rate.
Smart Images

Figure CN120953236A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent monitoring technology for power systems, specifically relating to a method and system for detecting anomalies in ophthalmic images based on transformers. Background Technology
[0002] Detection of ophthalmic imaging abnormalities is crucial for the early diagnosis of lesions, especially in cases such as infectious keratitis and diabetic keratopathy, where the identification of changes in nerve fibers and cellular structures is necessary. However, current technologies suffer from systemic limitations in handling the multi-scale characteristics of ophthalmic images (from macroscopic lesions to microscopic cellular changes): 1. Limitations of Ophthalmology-Specific Methods: The invention patent publication number CN116703915B, "A Consultation Method and System for Diabetic Retinopathy Based on Multiple Image Fusion," has the following problems: It relies solely on color features: Detecting abnormalities using an RGB color benchmark cannot process single-channel grayscale images such as OCT and FFA, and ignores the shape / texture features of subtle lesions such as microaneurysms; Limited application scenarios: It is only applicable to ultra-wide-angle fundus color photographs and fails to work with high-resolution microscopic images (such as corneal in vivo confocal microscopy images). The invention patent publication number CN116703915B, "A Method for Detecting Abnormalities in Fundus Images of Premature Infants Combining Self-Supervised and Supervised Methods," has the following problems: Labeling dependency bottleneck: Although there is self-supervised pre-training, the supervised stage still requires a large number of labeled rare lesion samples (such as retinoblastoma), making practical deployment difficult; Multi-category blind spots: Modeling only a single category of "normal fundus" cannot simultaneously process multiple normal subtypes (such as abnormal cell morphologies at multiple corneal layers), leading to a high false positive rate.
[0003] 2. Limitations in processing microscopic images: For example, the invention patent publication number CN114155237A, "A Method and Terminal for Detecting Anomalies in Medical Images Based on Unsupervised Learning," uses an intensity range encoding method, which mainly focuses on the distribution of voxel intensity and ignores more complex features in the image (such as texture, shape, and structure), resulting in insufficient feature representation ability. Secondly, the intensity range-based partitioning method relies on predefined cluster sets, lacks adaptability, and cannot dynamically adjust to adapt to the feature distribution of different categories of samples. This static cluster set partitioning limits the model's adaptability and generalization ability, making it unable to effectively distinguish category features when faced with multiple categories and diverse normal samples, resulting in insufficient generalization ability. For example, the invention patent publication number CN117934477A, "A Brain Tumor Image Detection Method Based on Unsupervised Learning", proposes a brain tumor image detection method based on unsupervised learning. The base model of this method is the convolutional neural network model ResNet18. According to recent studies, the convolutional neural network architecture is weaker than the transformer architecture in representing image features. Moreover, this method only uses the ResNet18 network with a small number of parameters, which may make it difficult to accurately represent the lesion features of microscopic images. Summary of the Invention
[0004] To overcome the aforementioned shortcomings of existing technologies, this invention proposes a method and system for anomaly detection in ophthalmic images based on the Transformer architecture. The algorithm is centered on the Transformer architecture and deeply integrates the EfficientNet-B3 feature extraction module with multi-scale embedding layers. By designing a large number of parameters, it achieves deep feature mining of images. This algorithm can utilize the original pixel features of microscopic images, handle diverse and multi-class samples, and exhibits strong generalization ability. Based on image reconstruction tasks, and leveraging the powerful sequence generation capabilities and contextual modeling advantages of Transformer, global pixel-level associations are established through a self-attention mechanism. This enables the accurate capture of the potential distribution patterns of image data, thereby significantly improving the accuracy and robustness of anomaly detection tasks and better meeting the deep feature modeling needs of complex medical images.
[0005] The technical solution adopted by the present invention to solve its technical problem is: an ophthalmic image abnormality detection system based on transformer, including an image acquisition module, a preprocessing module, a data labeling and storage management module, an abnormality detection module, and a doctor's judgment module; The image acquisition module is used to acquire high-quality ophthalmic image data and supports connection to various types of ophthalmic image acquisition devices; The preprocessing module takes the acquired raw images into the preprocessing flow; it supports image cropping and ROI selection functions, allowing users to manually or automatically select the analysis area and eliminate invalid background interference. The data tagging and storage management module is used to automatically associate each acquired image with metadata to ensure image traceability; images and their metadata are uniformly stored in the database; and images are pre-labeled with image classification tags, which include "normal", "suspicious", and "abnormal". The anomaly detection module consists of an anomaly detection model, which includes a multi-scale feature extraction stage, a multi-scale feature fusion stage, and a feature reconstruction stage. The multi-scale feature extraction stage uses a pre-trained EfficientNet-B3 as the backbone network for feature extraction. Feature extraction is performed from the 2nd to the 5th stages of EfficientNet-B3, extracting multi-scale features from low to high levels respectively. The feature map output by each stage contains different receptive fields and semantic depths, which can cover local texture information and overall structural information. The multi-scale feature fusion stage inputs the feature maps from the four stages output from the multi-scale feature extraction module into four embedding layers respectively, mapping them to a unified dimension; these mapped features are then input into a sub-network consisting of three fully connected layers to achieve cross-scale feature fusion; a hierarchical structure is introduced during the fusion process to support the interaction and combination of features at multiple levels; the output is a fused feature map of a unified dimension, providing input for the feature reconstruction stage; The feature reconstruction stage employs a Transformer architecture based on a self-attention mechanism. It reconstructs input features and calculates the reconstruction error. Based on the magnitude of the reconstruction error, it identifies abnormal regions, thus achieving anomaly detection. The feature reconstruction stage consists of multiple neighborhood mask encoders and hierarchical query decoders stacked alternately. Each neighborhood mask encoder follows a standard Transformer encoder structure, consisting of an attention module and a feedforward network. The attention module uses a neighborhood mask attention mechanism. The hierarchical query decoder consists of two parallel neighborhood mask encoder submodules and a feedforward network. In the decoding stage, the hierarchical query decoder receives the output of the previous hierarchical query decoder and the encoded vector from the neighborhood mask encoder submodule as input, performing layer-by-layer feature refinement to achieve progressive reconstruction. The output of the last hierarchical query decoder is considered the reconstructed features of the image. By comparing them with the original image features, a reconstruction error map is calculated. If the reconstruction error exceeds a set threshold, the system determines the image to be an abnormal image, and the area with significant error is the location of potential lesions. The doctor's judgment module includes an abnormal area visualization unit, a human-computer interaction annotation unit, and a comprehensive judgment and report generation unit; The abnormal region visualization unit receives the reconstruction error map output by the abnormality detection module and marks the abnormal region on the original image by overlaying a heat map; it supports multiple visualization modes to improve doctors' interpretation efficiency; the visualization results are used to guide doctors to focus on potential abnormal regions and improve the early lesion detection rate. The human-computer interaction annotation unit provides a graphical user interface, allowing doctors to interact with the visualized results. These interactions include: confirming abnormalities, adjusting regions, and adding diagnostic notes. The annotation data can be used for subsequent model fine-tuning and continuous learning, enabling adaptive optimization of model performance. The system supports doctors in confirming and rejecting AI diagnostic results and records their operation process and judgment criteria. The integrated judgment and report generation unit is used to automatically summarize abnormal area information, doctor annotation results, and model confidence scores; based on preset rules or auxiliary knowledge bases, it generates a preliminary structured diagnostic report, including lesion indications, abnormality type suggestions, and further examination suggestions; the report supports standard format export and can be connected to medical information platforms.
[0006] Preferably, ophthalmic image acquisition equipment includes in vivo confocal microscopes, slit-lamp microscope cameras, and optical coherence tomography (OCT).
[0007] Preferably, the preprocessing steps include grayscale normalization, noise reduction, contrast enhancement, and resolution unification.
[0008] Preferably, the preprocessing module incorporates an image quality assessment module to automatically score the acquired images, and prompts for re-acquisition if the score does not reach the threshold.
[0009] Preferably, the visualization modes supported by the abnormal area visualization unit include threshold highlighting, area masking, and outline tracing.
[0010] Preferably, the abnormal area information automatically summarized by the comprehensive judgment and report generation unit includes location, area, and morphological features.
[0011] Preferably, the training method for the anomaly detection model includes the following steps: Construction of the training set: After the image samples used for training are collected, they are independently evaluated by two or more experts to determine their image classification labels; images with conflicting evaluation results or unclear evaluations are excluded, while images with consistent expert evaluations are assigned clear true classification labels; samples with the true classification label determined to be "normal" are included in the training set; During training, the objective function is defined as the original feature map. With reconstructed feature maps Mean square error loss between; loss The calculation method is as follows: ; Where H and W are the dimensions of the feature map; For the input image to be predicted, the model computes the original feature map. With reconstructed feature maps The L2 norm of the differences between each pixel is used to process the new image, and then these values are averaged over all pixels to obtain the reconstruction error of the image. ; When the reconstruction error S of a given image exceeds a predefined threshold, the image will be classified as an anomalous image; this predefined threshold is determined based on the reconstruction error values of the training set samples.
[0012] A transformer-based method for detecting anomalies in ophthalmic images, based on the aforementioned system, includes the following steps: Starting with the image acquisition module, which serves as the front-end input, the module connects to the ophthalmic image acquisition equipment via a standard interface protocol. The acquired raw images enter the preprocessing process. The processed images are automatically associated with the patient's basic information, acquisition time, and other metadata, and are stored in a unified database. The database supports local and cloud synchronization and can also be integrated with the hospital's existing data management system. The standardized image input anomaly detection module, processed by the image acquisition module, identifies abnormal regions based on the Transformer reconstruction mechanism: It utilizes a pre-trained EfficientNet-B3 as the backbone network, extracting multi-scale features from stages 2 to 5, covering local texture and overall structural information, and freezing pre-trained weights to ensure feature extraction consistency; the feature maps from each stage are mapped to a unified dimension through embedding layers, and input into a fully connected sub-network to achieve cross-scale fusion, resulting in a unified fused feature map; feature reconstruction is performed using a Transformer architecture composed of alternating stacks of neighborhood mask encoders and hierarchical query decoders, introducing learnable positional encoding to preserve spatial location information; the reconstructed features output from the last layer are compared with the original features, and a reconstruction error map is calculated. If the error exceeds a set threshold, the image is judged as abnormal, and areas with significant reconstruction errors are identified as potential lesions. The reconstruction error map generated by the anomaly detection module is transmitted to the doctor's judgment module. The anomaly region visualization unit marks the anomaly regions on the original image, guiding doctors to focus on potential anomalies. The human-computer interaction annotation unit provides a graphical interface for doctors to confirm anomalies, adjust regions, and add notes. The annotation data is used for fine-tuning the system model. Doctors can confirm or reject AI diagnostic results and record the basis. The comprehensive judgment and report generation unit summarizes the anomaly region information, doctor annotation results, and model confidence scores, and generates a structured preliminary diagnostic report based on preset rules or a knowledge base. It supports standard format export and integration with the hospital information system.
[0013] Compared with the prior art, the beneficial effects of the present invention are: 1. The algorithm proposed in this invention uses a Transformer architecture based on a self-attention mechanism. By reconstructing the input features and calculating the reconstruction error, the abnormal region is determined based on the magnitude of the reconstruction error, thereby achieving accurate anomaly detection for ophthalmic microscopic images.
[0014] 2. This invention integrates a feature extraction module, a feature fusion module, and a feature reconstruction module, using three network structures to fully utilize the advantages of each network structure and significantly improve the accuracy and robustness of ophthalmic image anomaly detection tasks: EfficientNet-B3 can fully extract the edge and texture features of the image, multi-scale embedding layers are used to fuse image features of different scales, and the Transformer network is used for image reconstruction.
[0015] 3. The attention module of this invention employs a "Neighborhood Masked Attention" (NMA) mechanism. Neighborhood masks effectively prevent information leakage between neighboring areas, maintaining the integrity of the spatial structure. For reconstruction tasks, without NMA, the model tends to directly generate the original image, which fails to achieve anomaly detection; therefore, NMA is necessary. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below.
[0017] Figure 1 This is an architecture diagram of an anomaly detection model for an ophthalmic image anomaly detection system based on transformer, according to an embodiment of the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0019] In the description of this invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.
[0020] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances. Furthermore, the technical features involved in the different embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0021] This invention discloses an ophthalmic image anomaly detection system based on transformer, comprising an image acquisition module, a preprocessing module, a data labeling and storage management module, an anomaly detection module, and a doctor's decision module.
[0022] The image acquisition module is used to acquire high-quality ophthalmic image data and supports connection to various types of ophthalmic imaging acquisition devices. As the front-end input part of the system, the image acquisition module is primarily responsible for acquiring high-quality ophthalmic image data, providing a reliable data foundation for subsequent anomaly detection and physician interpretation. Supported ophthalmic imaging devices include in vivo confocal microscopy (IVCM), optical coherence tomography (OCT), and slit-lamp microscopy. It interfaces with hardware devices through standard interface protocols (such as DICOM, USB video stream, and image file import), exhibiting good adaptability. For IVCM images, the system supports the acquisition of continuous image sequences, preserving image temporal and depth location information.
[0023] The preprocessing module takes the acquired raw images into the preprocessing flow; it supports image cropping and ROI selection functions, allowing users to manually or automatically select the analysis area and eliminate invalid background interference. The preprocessing flow includes grayscale normalization, noise reduction, contrast enhancement, and resolution unification. The preprocessing module also incorporates an image quality assessment module, which automatically scores the acquired images; if the score does not meet the threshold, it prompts for re-acquisition.
[0024] The data tagging and storage management module automatically associates metadata with each acquired image to ensure image traceability; images and their metadata are uniformly stored in the database; images are pre-labeled with image classification tags, including "normal," "suspicious," and "abnormal," to facilitate subsequent screening and analysis; images and their metadata are uniformly stored in the database, supporting both local storage and cloud synchronization modes to ensure data security and scalability; it can seamlessly integrate with existing hospital data management systems (such as PACS) to support centralized management and remote access to image data.
[0025] The anomaly detection module stores an anomaly detection model, which includes a multi-scale feature extraction stage, a multi-scale feature fusion stage, and a feature reconstruction stage. The core of the ophthalmic image anomaly detection system proposed in this embodiment lies in the anomaly detection module. This module, based on the Transformer reconstruction mechanism, identifies abnormal regions in the image by modeling the reconstruction capability of normal images. The architecture diagram of the anomaly detection model is shown below. Figure 1 As shown.
[0026] The multi-scale feature extraction stage uses a pre-trained EfficientNet-B3 as the backbone network. This network has good feature representation capabilities and can extract rich spatial and semantic information while maintaining a lightweight model. Feature extraction is performed from stage 2 to stage 5 of EfficientNet-B3, extracting multi-scale features from low to high levels respectively. The feature maps output by each stage contain different receptive fields and semantic depths, covering both local texture information and overall structural information. Since EfficientNet-B3 has fixed parameters (i.e., frozen pre-trained weights), the consistency and generalization of feature extraction can be guaranteed, providing high-quality input for subsequent fusion and reconstruction.
[0027] The multi-scale feature fusion stage inputs the feature maps from the four stages (stage-2 to stage-5) output from the multi-scale feature extraction module into four embedding layers (ELs) respectively, mapping them to a unified dimension. These mapped features are then input into a sub-network (Fully Connected Network, FCN) consisting of three fully connected layers to achieve cross-scale feature fusion. A hierarchical structure is introduced during the fusion process to support the interaction and combination of features at multiple levels, effectively improving the ability to perceive minute abnormal structures in corneal images. The output is a fused feature map of a unified dimension, providing input for the feature reconstruction stage. The feature reconstruction stage adopts a Transformer architecture based on a self-attention mechanism. It reconstructs the input features and calculates the reconstruction error. Based on the magnitude of the reconstruction error, it identifies abnormal regions and achieves anomaly detection. The feature reconstruction stage consists of multiple Neighbor Masked Encoders (NMEs) and Layer-wise Query Decoders (LQDs) stacked alternately. Each Neighbor Masked Encoder follows the standard Transformer encoder structure and consists of an attention module and a feedforward network (FFN). The attention module adopts a Neighborhood Masked Attention (NMA) mechanism. The hierarchical query decoder consists of two parallel neighborhood mask encoder submodules and a feedforward network. During the decoding stage, the hierarchical query decoder receives the output of the previous hierarchical query decoder and the encoded vector from the neighborhood mask encoder submodule as input, performs layer-by-layer feature refinement, and achieves step-by-step reconstruction. The output of the last hierarchical query decoder is regarded as the reconstructed features of the image. By comparing it with the features of the original image, the reconstruction error map is calculated. If the reconstruction error exceeds a set threshold, the system determines that the image is an abnormal image, and the area with significant error is the location of potential lesions. The anomaly detection module leverages the global modeling capabilities and powerful feature representation capabilities of Transformer to effectively improve the sensitivity and accuracy of anomaly detection while maintaining the integrity of the image structure. It is particularly suitable for automatic identification of corneal image anomalies in label-free environments.
[0028] The physician judgment module includes an abnormal area visualization unit, a human-computer interaction annotation unit, and a comprehensive judgment and report generation unit. As the final judgment stage in the system of this invention, the physician judgment module acts as a bridge between artificial intelligence-assisted diagnosis and clinical decision support. Based on the result image generated by the abnormality detection module and combined with the physician's professional knowledge, this module outputs the final diagnostic judgment and treatment recommendations for ophthalmic images.
[0029] The abnormal region visualization unit receives the reconstruction error map output by the abnormality detection module and marks the abnormal region on the original image by overlaying a heatmap. It supports multiple visualization modes to improve doctors' interpretation efficiency. The visualization results are used to guide doctors to focus on potential abnormal regions and improve the early detection rate of lesions. The visualization modes supported by the abnormal region visualization unit include threshold highlighting, region masking, and contour outlining.
[0030] The human-computer interaction annotation unit provides a graphical user interface, allowing doctors to interact with the visualized results. These interactions include: confirming abnormalities, adjusting regions, and adding diagnostic notes. The annotated data can be used for subsequent model fine-tuning and continuous learning, enabling adaptive optimization of model performance. The system supports doctors in confirming and rejecting AI diagnostic results and records their operation process and judgment criteria. The integrated assessment and report generation unit automatically summarizes information on abnormal areas, doctor annotations, and model confidence scores. Based on preset rules or an auxiliary knowledge base, it generates a structured preliminary diagnostic report, including lesion indications, suggested abnormality types (such as keratitis, corneal degeneration, etc.), and recommendations for further examinations. The report supports export in a standard format and can be integrated with medical information platforms, including Hospital Information Systems (HIS) and Electronic Medical Records (EMR). The abnormal area information automatically summarized by the integrated assessment and report generation unit includes location, area, and morphological characteristics.
[0031] The training method for anomaly detection models includes the following steps: Construction of the training set: After the image samples used for training are collected, they are independently evaluated by two or more experts to determine their image classification labels; images with conflicting or unclear evaluation results from different experts are excluded, while images with consistent expert evaluations are assigned clear true classification labels, determining which category they belong to: "normal", "suspicious", or "abnormal"; samples with the true classification label determined to be "normal" are included in the training set. During training, the objective function is defined as the original feature map. With reconstructed feature maps Mean square error loss between; loss The calculation method is as follows: ; Where H and W are the dimensions of the feature map; For the input image to be predicted, the model computes the original feature map. With reconstructed feature maps The L2 norm of the differences between each pixel is used to process the new image, and then these values are averaged over all pixels to obtain the reconstruction error of the image. ; When the reconstruction error S of a given image exceeds a predefined threshold, the image will be classified as an anomalous image; this predefined threshold is determined based on the reconstruction error values of the training set samples.
[0032] A transformer-based method for detecting anomalies in ophthalmic images, based on the aforementioned system, includes the following steps: Starting with the image acquisition module, which serves as the front-end input, the module connects to the ophthalmic image acquisition equipment via a standard interface protocol. The acquired raw images enter the preprocessing process. The processed images are automatically associated with the patient's basic information, acquisition time, and other metadata, and are stored in a unified database. The database supports local and cloud synchronization and can also be integrated with the hospital's existing data management system. The standardized image input anomaly detection module, processed by the image acquisition module, identifies abnormal regions based on the Transformer reconstruction mechanism: It utilizes a pre-trained EfficientNet-B3 as the backbone network, extracting multi-scale features from stages 2 to 5, covering local texture and overall structural information, and freezing pre-trained weights to ensure feature extraction consistency; the feature maps from each stage are mapped to a unified dimension through embedding layers, and input into a fully connected sub-network to achieve cross-scale fusion, resulting in a unified fused feature map; feature reconstruction is performed using a Transformer architecture composed of alternating stacks of neighborhood mask encoders and hierarchical query decoders, introducing learnable positional encoding to preserve spatial location information; the reconstructed features output from the last layer are compared with the original features, and a reconstruction error map is calculated. If the error exceeds a set threshold, the image is judged as abnormal, and areas with significant reconstruction errors are identified as potential lesions. The reconstruction error map generated by the anomaly detection module is transmitted to the doctor's judgment module. The anomaly region visualization unit marks the anomaly regions on the original image, guiding doctors to focus on potential anomalies. The human-computer interaction annotation unit provides a graphical interface for doctors to confirm anomalies, adjust regions, and add notes. The annotation data is used for fine-tuning the system model. Doctors can confirm or reject AI diagnostic results and record the basis. The comprehensive judgment and report generation unit summarizes the anomaly region information, doctor annotation results, and model confidence scores, and generates a structured preliminary diagnostic report based on preset rules or a knowledge base. It supports standard format export and integration with the hospital information system.
[0033] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A transformer-based ophthalmic image anomaly detection system, characterized in that, It includes an image acquisition module, a preprocessing module, a data labeling and storage management module, an anomaly detection module, and a doctor's decision module; The image acquisition module is used to acquire high-quality ophthalmic image data and supports connection to various types of ophthalmic image acquisition devices; The preprocessing module takes the acquired raw images into the preprocessing flow; it supports image cropping and ROI selection functions, allowing users to manually or automatically select the analysis area and eliminate invalid background interference. The data tagging and storage management module is used to automatically associate each captured image with metadata to ensure image traceability. Images and their metadata are stored uniformly in the database; images are pre-labeled with image classification tags, which include "normal", "suspicious", and "abnormal"; The anomaly detection module consists of an anomaly detection model, which includes a multi-scale feature extraction stage, a multi-scale feature fusion stage, and a feature reconstruction stage. The multi-scale feature extraction stage uses a pre-trained EfficientNet-B3 as the backbone network for feature extraction. Feature extraction is performed from the 2nd to the 5th stages of EfficientNet-B3, extracting multi-scale features from low to high levels respectively. The feature map output by each stage contains different receptive fields and semantic depths, which can cover local texture information and overall structural information. The multi-scale feature fusion stage inputs the feature maps from the four stages output from the multi-scale feature extraction module into four embedding layers respectively, mapping them to a unified dimension; these mapped features are then input into a sub-network consisting of three fully connected layers to achieve cross-scale feature fusion; a hierarchical structure is introduced during the fusion process to support the interaction and combination of features at multiple levels; the output is a fused feature map of a unified dimension, providing input for the feature reconstruction stage; The feature reconstruction stage adopts a Transformer architecture based on a self-attention mechanism. It reconstructs the input features and calculates the reconstruction error. Based on the magnitude of the reconstruction error, it identifies abnormal regions and achieves anomaly detection. The feature reconstruction stage consists of multiple neighborhood mask encoders and hierarchical query decoders stacked alternately. Each neighborhood mask encoder follows the standard Transformer encoder structure and consists of an attention module and a feedforward network. The attention module employs a neighborhood mask attention mechanism; The hierarchical query decoder consists of two parallel neighborhood mask encoder submodules and a feedforward network. During the decoding stage, the hierarchical query decoder receives the output of the previous hierarchical query decoder and the encoded vector from the neighborhood mask encoder submodule as input, performs layer-by-layer feature refinement, and achieves step-by-step reconstruction. The output of the last hierarchical query decoder is regarded as the reconstructed features of the image. By comparing it with the features of the original image, a reconstruction error map is calculated. If the reconstruction error exceeds a set threshold, the system determines that the image is an abnormal image, and the area with significant error is the location of potential lesions. The doctor's judgment module includes an abnormal area visualization unit, a human-computer interaction annotation unit, and a comprehensive judgment and report generation unit; The abnormal region visualization unit receives the reconstruction error map output by the abnormality detection module and marks the abnormal region on the original image by overlaying a heat map. It supports multiple visualization modes to improve doctors' interpretation efficiency; Visualization results are used to guide doctors to focus on potentially abnormal areas and improve the early detection rate of lesions; The human-computer interaction annotation unit provides a graphical user interface, allowing doctors to interact with the visualized results. These interactions include: confirming abnormalities, adjusting regions, and adding diagnostic notes. The annotation data can be used for subsequent model fine-tuning and continuous learning, enabling adaptive optimization of model performance. The system supports doctors in confirming and rejecting AI diagnostic results and records their operation process and judgment criteria. The integrated judgment and report generation unit is used to automatically summarize abnormal area information, doctor annotation results, and model confidence scores; based on preset rules or auxiliary knowledge bases, it generates a preliminary structured diagnostic report, including lesion indications, abnormality type suggestions, and further examination suggestions; the report supports standard format export and can be connected to medical information platforms.
2. The ophthalmic image abnormality detection system based on transformer according to claim 1, characterized in that, Ophthalmic imaging equipment includes in vivo confocal microscopes, slit-lamp microscopes, and optical coherence tomography (OCT).
3. The ophthalmic image abnormality detection system based on transformer according to claim 1, characterized in that, The preprocessing workflow includes grayscale normalization, noise reduction, contrast enhancement, and resolution unification.
4. The ophthalmic image abnormality detection system based on transformer according to claim 1, characterized in that, The preprocessing module incorporates an image quality assessment module to automatically score the acquired images. If the score does not reach the threshold, the system prompts the user to re-acquire the images.
5. The ophthalmic image abnormality detection system based on transformer according to claim 1, characterized in that, The visualization modes supported by the abnormal area visualization unit include threshold highlighting, area masking, and outline tracing.
6. The ophthalmic image abnormality detection system based on transformer according to claim 1, characterized in that, The abnormal area information automatically summarized by the comprehensive judgment and report generation unit includes location, area, and morphological characteristics.
7. The ophthalmic image abnormality detection system based on transformer according to claim 1, characterized in that, The training method for the anomaly detection model includes the following steps: Construction of the training set: After the image samples used for training are collected, they are independently evaluated by two or more experts to determine their image classification labels, which are divided into "normal", "suspicious" and "abnormal". Images with conflicting evaluation results or unclear evaluations by different experts are excluded, while images with consistent evaluation opinions are assigned clear true classification labels. Samples with the true classification label determined to be "normal" are included in the training set. During training, the objective function is defined as the original feature map. With reconstructed feature maps Mean square error loss between; loss The calculation method is as follows: ; Where H and W are the dimensions of the feature map; For the input image to be predicted, the model computes the original feature map. With reconstructed feature maps The L2 norm of the differences between each pixel is used to process the new image, and then these values are averaged over all pixels to obtain the reconstruction error of the image. ; When the reconstruction error S of a given image exceeds a predefined threshold, the image will be classified as an anomalous image; this predefined threshold is determined based on the reconstruction error values of the training set samples.
8. A method for detecting anomalies in ophthalmic images based on transformer, using any one of the systems described in claims 1-7, comprising the following steps: Starting with the image acquisition module, which serves as the front-end input, the module connects to the ophthalmic image acquisition equipment via a standard interface protocol. The acquired raw images enter the preprocessing process. The processed images are automatically associated with the patient's basic information, acquisition time, and other metadata, and are stored in a unified database. The database supports local and cloud synchronization and can also be integrated with the hospital's existing data management system. The standardized image input anomaly detection module, processed by the image acquisition module, identifies abnormal regions based on the Transformer reconstruction mechanism: It utilizes a pre-trained EfficientNet-B3 as the backbone network, extracting multi-scale features from stages 2 to 5, covering local texture and overall structural information, and freezing pre-trained weights to ensure feature extraction consistency; the feature maps from each stage are mapped to a unified dimension through embedding layers, and input into a fully connected sub-network to achieve cross-scale fusion, resulting in a unified fused feature map; feature reconstruction is performed using a Transformer architecture composed of alternating stacks of neighborhood mask encoders and hierarchical query decoders, introducing learnable positional encoding to preserve spatial location information; the reconstructed features output from the last layer are compared with the original features, and a reconstruction error map is calculated. If the error exceeds a set threshold, the image is judged as abnormal, and areas with significant reconstruction errors are identified as potential lesions. The reconstruction error map generated by the anomaly detection module is transmitted to the doctor's judgment module. The anomaly region visualization unit marks the anomaly regions on the original image, guiding doctors to focus on potential anomalies. The human-computer interaction annotation unit provides a graphical interface for doctors to confirm anomalies, adjust regions, and add notes. The annotation data is used for fine-tuning the system model. Doctors can confirm or reject AI diagnostic results and record the basis. The comprehensive judgment and report generation unit summarizes the anomaly region information, doctor annotation results, and model confidence scores, and generates a structured preliminary diagnostic report based on preset rules or a knowledge base. It supports standard format export and integration with the hospital information system.
Citation Information
Patent Citations
Medical image anomaly detection method and terminal based on unsupervised learning
CN114155237A
Diabetic retinopathy consultation method and system based on multiple image fusion
CN116703915B
Brain tumor image detection method based on unsupervised learning
CN117934477A