An optical coherence tomography-based eye index prediction method and device
By using a deep learning model based on optical coherence tomography, the limitations of existing technologies in predicting the light sensitivity threshold and mode deviation probability map are overcome, generating a more accurate mode deviation probability map and enabling early detection and more efficient diagnosis of glaucoma.
Patent Information
- Application Number
- CN202411942299.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-12-27
AI Technical Summary
Existing technologies have limitations in predicting visual field light sensitivity thresholds and pattern bias probability maps, particularly due to limited sample diversity, lack of external validation, and difficulty in accurately mapping the relationship between structural damage and functional loss. Consequently, they cannot effectively generate pattern bias probability maps, which affects the early detection and diagnosis of glaucoma.
By constructing an ocular index prediction method based on optical coherence tomography (OCT), utilizing 3D OCT scan volume data, and combining it with a deep learning neural network model, this method employs data augmentation, fusion, and sampling strategies, and introduces oversampling and cost-sensitive learning methods to generate photosensitivity threshold maps and mode bias probability maps, thereby addressing the dataset imbalance problem.
It enables more accurate prediction of visual field light sensitivity threshold and pattern deviation probability map, provides more detailed structural and functional analysis, and can detect small but clinically significant lesions at an early stage, thus improving the accuracy and efficiency of glaucoma diagnosis.
Smart Images

Figure CN119745312B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, and in particular to an eye index prediction method and device based on optical coherence tomography. BACKGROUND
[0002] The diagnosis of glaucoma mainly relies on the combination of structural and functional tests of optical coherence tomography (OCT) and visual field examination. Structural-functional mapping is crucial in the diagnosis of glaucoma, as it elucidates the relationship between structural damage and functional loss, allowing early structural changes to be detected before functional defects become apparent in visual field examination. This early detection is crucial for timely intervention and possibly better control of glaucoma. OCT is used to detect structural changes in the optic nerve head (ONH) and macula, while visual field examination is used to assess visual field function. Structural-functional mapping is key to the diagnosis of glaucoma, as it reveals the relationship between structural damage and functional loss. However, existing techniques have limitations in predicting visual field light sensitivity thresholds and pattern deviation probability maps (PDP).
[0003] Humphrey visual field (HVF) is a medical test used to assess and document visual field (VF), developed by Humphrey Instruments, and is an integral part of modern ophthalmic clinical practice, providing eye doctors with an important tool for assessing and understanding the visual field status of patients. This test is mainly used to diagnose and monitor glaucoma and other ophthalmic diseases that can cause visual field defects. The HVF test assesses the light sensitivity threshold of the visual field by measuring the response of the retina to light stimuli, helping doctors understand potential damage in the patient's visual field. A key feature of the HVF test is that the test results are presented in the form of visual field maps, showing the light sensitivity threshold in different areas of the patient's visual field. These maps usually include total deviation (TD) and pattern deviation (PD) maps, as well as corresponding probability maps. The pattern deviation value map is a map that highlights local light sensitivity threshold reduction by filtering out overall light sensitivity threshold reduction, thus highlighting local light sensitivity threshold reduction. The pattern deviation probability map is a probability map that describes the pattern deviation result, reflecting local visual field defects, and can ignore normal range changes to highlight small but clinically significant lesions, and can be used to observe early visual field defects.
[0004] Previous studies have demonstrated a strong association between OCT and visual field (VF) defects, suggesting the possibility of predicting the light sensitivity threshold of VF and generating a pattern deviation probability map (PDP) from structural OCT scans.
[0005] OCT and VF results are not static, but evolve over time. Accurately temporally linking OCT and VF results is crucial to more precisely understand the relationship between them.
[0006] Furthermore, glaucoma's regional variability in structural and functional changes at the optic nerve head and HVF has not been fully addressed. Most studies rely on cross-sectional or circular scans, rather than volumetric scans, thus neglecting information from other retinal structures outside the scanned area.
[0007] In summary, the drawbacks of the prior art include:
[0008] These studies have limitations in limited sample diversity, lack of broad external validation, and difficulty in accurately predicting the visual field light sensitivity threshold and pattern deviation probability map (PDP). Specifically:
[0009] (1) Temporally inconsistent OCT and visual field test results make it difficult to accurately map the relationship between structural damage and functional loss.
[0010] (2) Limited sample size, lack of external medical center validation, limits the universality and reliability of the results.
[0011] (3) Only the light sensitivity numerical view in VF data is predicted, and the pattern deviation probability map cannot be predicted, such as Figure 1 As shown, the light sensitivity threshold map, although providing part of the information of retinal function, is difficult to highlight local minor but clinically significant lesions due to the influence of factors that generally reduce retinal sensitivity, such as age, media opacity, cataract, etc., and is not sufficient to comprehensively assess the risk of glaucoma disease. In addition, the readability of the light sensitivity threshold map is poor, and it is difficult for doctors to efficiently understand its content in actual clinical practice.
[0012] In clinical practice, the pattern deviation probability map (PDP) is often used with the light visual field sensitivity threshold Figure 1 to provide a more comprehensive analysis of the visual field. In fact, the PDP is one of the views that glaucoma doctors pay the most attention to, which provides accurate structural-function relationships to clinicians in an easy-to-understand way, helps to distinguish visual field damage due to pathological conditions such as glaucoma from normal variations, and can be used to observe early visual field defects. SUMMARY
[0013] The present application aims to overcome the defects of the prior art that can only generate a photosensitivity threshold map and cannot provide a pattern deviation probability map, and provides an optical coherence tomography-based eye index prediction method and device.
[0014] The object of the present application can be achieved by the following technical solutions:
[0015] An optical coherence tomography-based eye index prediction method comprises the following steps:
[0016] Obtain the optical coherence tomography volume data and the corresponding test report, and construct a model training data set;
[0017] Data preprocessing is performed on the volume data in the model training data set;
[0018] An eye index prediction model is constructed, which comprises a visual encoder, a regression head and a classification head. The visual encoder is used to extract features from the volume data. The regression head is used to regress and calculate the photosensitivity threshold of each test point according to the extracted features, thereby obtaining a photosensitivity threshold map. The classification head is used to predict whether each test point is abnormal and the abnormality level according to the extracted features, thereby obtaining a pattern deviation probability map.
[0019] The eye index prediction model is trained based on the model training data set, and the trained eye index prediction model is used to predict the eye index of the optical coherence tomography data of the to-be-tested object.
[0020] Further, the data preprocessing process comprises normalizing, scaling, adding random noise and data flipping to the volume data for data enhancement.
[0021] Further, the data flipping comprises sequentially flipping each slice image of the left eye or right eye volume data of the same object, and horizontally mirroring each slice image.
[0022] Further, the method further comprises obtaining multiple types of optical coherence tomography volume data, and fusing various types of optical coherence tomography volume data of the same object, thereby constructing the model training data set.
[0023] The fusion method comprises horizontally splicing different types of volume data, or sampling and merging each type of slice image.
[0024] The types of optical coherence tomography volume data comprise macular volume data, optic disc volume data and wide field volume data obtained based on optical coherence tomography.
[0025] Further, the method further comprises: expanding the model training data set by copying samples including more abnormal test points, according to the abnormal test point data contained in each sample in the model training data set.
[0026] Further, the calculation process of the loss function in the training process of the eye index prediction model comprises:
[0027] The actual number of normal test points and abnormal test points in the current training process is counted, the loss of the normal test points and the abnormal test points is calculated respectively, and the final loss result is obtained by dividing the loss of the normal test points and the abnormal test points by the actual number of the normal test points and the abnormal test points respectively.
[0028] Further, in the regression calculation process of the light sensitivity threshold, the calculated light sensitivity threshold is normalized and limited within a new range.
[0029] Further, the visual encoder is a 2D encoder or a 3D encoder, the input of the 2D encoder is a plurality of slice images in the volume data, the features corresponding to each slice image output by the 2D encoder are merged to form the features corresponding to the volume data.
[0030] The input of the 3D encoder is the volume data, and the features corresponding to the volume data are directly obtained.
[0031] Further, the method further comprises: using a region random sampling strategy to select slice images of each volume data, and inputting the selected slice images into the eye index prediction model for prediction.
[0032] The application also provides an eye index prediction device based on optical coherence tomography, comprising a memory and a processor, the memory stores a computer program, and the processor calls the computer program to execute the steps of the method as described above.
[0033] Compared with the prior art, the application has the following advantages:
[0034] (1) In order to realize the prediction of the mode deviation probability map, the original 3D OCT scan volume data is directly used for processing, the original 3D OCT scan volume data more accurately predicts the 52 visual field light sensitivity thresholds and the mode deviation probability map of VF, and provides a more useful tool for the diagnosis of glaucoma.
[0035] Compared with 2D images, the OCT scan body data contains more key information about the retina, but the processing and analysis of the OCT image involve the processing of high-dimensional data, feature extraction, computational efficiency and other technical difficulties. In contrast, the present application proposes a variety of data enhancement, fusion and sampling methods, and proposes a neural network model as a basic framework, which combines regression and classification prediction heads, and can point-by-point predict the light sensitivity threshold and pattern deviation probability points in the visual field test; the pattern deviation probability map can be more detailed in spatial analysis than the light sensitivity threshold, providing a closer link between structure and function, and can highlight local small but clinically significant lesions.
[0036] (2) The present application uses a variety of OCT scans (macula, optic disc, wide field and their fusion) for more detailed structural analysis, and the deep learning-based system can more accurately predict the light sensitivity threshold.
[0037] (3) Compared with generating a light sensitivity threshold map, generating a pattern deviation probability map is a more difficult task, and one of the main challenges is the imbalance of the data set. To this end, the present application first uses oversampling technology to count the abnormal test point data contained in each sample in the training data set, and increases the representativeness of the sample including more abnormal test points in the training set by copying it;
[0038] In addition, the method of cost-sensitive learning is also introduced, which dynamically adjusts the weight of the model in each iteration to allow the model to adjust the importance of its learning according to the actual number of samples of each class. Through this method, the learning process of the model can be adjusted without changing the data set, adapting to the inherent imbalance of the data and paying more attention to those samples that are more critical to clinical decision-making. BRIEF DESCRIPTION OF DRAWINGS
[0039] Figure 1 A visual field report example graph provided in the background art of the present application.
[0040] Figure 2 A flowchart of an optical coherence tomography-based eye index prediction method provided in an embodiment of the present application;
[0041] Figure 3 A data processing flowchart of an optical coherence tomography-based eye index prediction method provided in an embodiment of the present application;
[0042] Figure 4 A structure block diagram of an optical coherence tomography-based eye index prediction device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0043] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the following will be combined with the drawings in the embodiments of the present application to make a clear and complete description of the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings can be arranged and designed in various different configurations.
[0044] Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative labor are within the scope of protection of the present application.
[0045] It should be noted that: similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0046] Embodiment 1
[0047] As shown in Figure 2 and Figure 3 , the present embodiment provides an optical coherence tomography-based eye index prediction method, comprising the following steps:
[0048] S1: obtaining the volume data of optical coherence tomography and the corresponding test report, and constructing a model training data set;
[0049] S2: data preprocessing is performed on the volume data in the model training data set;
[0050] S3: constructing an eye index prediction model, the eye index prediction model comprising a visual encoder, a regression head and a classification head, the visual encoder being used to extract features of the volume data; the regression head being used to regress and calculate the light sensitivity threshold of each test point according to the extracted features, so as to obtain a light sensitivity threshold map; the classification head being used to predict whether each test point is abnormal and the abnormal level according to the extracted features, so as to obtain a mode deviation probability map;
[0051] S4: training the eye index prediction model based on the model training data set, and using the trained eye index prediction model to predict the eye index of the optical coherence tomography data of the to-be-tested object.
[0052] The following will specifically describe each aspect of the method.
[0053] I. Model training data set
[0054] For the above model training data set, optionally, the optical coherence tomography (OCT) volume data includes optic disc OCT, macular OCT, wide field OCT and the like. The test report can be a VF test report, which is obtained by using a Humphrey perimeter manufactured by Zeiss company.
[0055] II. Process of data preprocessing
[0056] The process of data preprocessing includes the following aspects:
[0057] 2.1. Data flipping
[0058] The order of each slice image of the same object left eye or right eye volume data is flipped, and each slice image is horizontally mirrored and flipped.
[0059] Specifically, three-dimensional OCT volume data is used as data, and one volume data contains 256 slice images. In order to align the left eye and right eye data, the right eye OCT data is flipped in this embodiment, and the specific implementation includes the order of 256 flipped slice images (flipped from 0, 1…255 to 255, 254…0) and horizontal mirror flipping of each slice image.
[0060] 2.2. Data augmentation
[0061] In the training stage, various image preprocessing methods are used to enhance the OCT data, including normalization, scaling, adding random noise and the like.
[0062] 2.3. Multiple OCT fusion
[0063] Preferably, the method further comprises obtaining multiple types of optical coherence tomography volume data, and fusing various types of optical coherence tomography volume data of the same object, thereby constructing a model training data set;
[0064] The fusion method includes horizontal splicing of different types of volume data, or sampling and merging of slice images of each type of volume data;
[0065] The types of optical coherence tomography volume data include macular volume data, optic disc volume data and wide field volume data obtained based on optical coherence tomography.
[0066] For example: fusion of optic disc OCT and macular OCT for prediction, one implementation scheme is to horizontally splice the optic disc slice image and the macular slice image, and another implementation scheme is to sample S / 2 slice images for each OCT, and obtain S slice images after merging.
[0067] 2.4. Region random sampling strategy
[0068] Preferably, the method of the present invention further includes selecting slice images of each individual data by using a regional random sampling strategy, and inputting the selected slice images into the eye index prediction model for prediction.
[0069] Specifically, training with OCT data including 256 slices requires extremely large computing resources and training time. Given the small OCT scan interval, in this invention, data with different numbers of slices were compared for training and testing, including 8, 16, 32, … 256, etc. For this reason, a regional random sampling strategy was designed to select slice images. Taking the training with 16 slice images as an example, first, the 256 slice images are divided into 16 groups. For example, the first group is [0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15]. Then, each time during the training iteration, a slice image is randomly selected from the first group as a sample of the first group. Similarly, the same sampling strategy is adopted for the other 15 groups, and finally, 16 groups of slice images are obtained and used as input data.
[0070] III. Data processing process of the eye index prediction model
[0071] The data processing process for the above eye index prediction model includes:
[0072] 3.1. The visual encoder is a 2D encoder or a 3D encoder. The input of the 2D encoder is multiple slice images in the volume data, and the features corresponding to each slice image output by the 2D encoder are merged to form the features corresponding to the volume data;
[0073] The input of the 3D encoder is the volume data, and the features corresponding to the volume data are directly obtained.
[0074] Here, taking the Transformer model as an example: When using a 2D encoder, the basic skeletons of the Swin-Transformer, ViT, and VIT-base series models can be directly used as the encoder. The OCT data is sent into the network in the form of 2D slice images. S slice images of an OCT data are encoded simultaneously. The size of each slice image is H*W. Finally, S slice image embeddings are obtained. By merging the S slice image embeddings, the OCT embedding is finally obtained. When using a 3D encoder, the feature embedding, position encoding, etc. of the Transformer encoder are adjusted according to the 3D data processing requirements. The OCT data is sent into the network in the form of 3D volume data with a size of H*W*S, where H and W are the height and width of the slice image respectively, and S is the number of slice images used, and the OCT embedding is obtained through encoding operations.
[0075] 3.2. The goal of the present invention is to predict the light sensitivity threshold and pattern deviation probability at each point of the VF test.
[0076] For the prediction of the light sensitivity threshold task, it is considered as a continuous regression task, and the light sensitivity threshold of each test point is calculated by regression according to the extracted features, to obtain a light sensitivity threshold map;
[0077] For the prediction of the pattern deviation probability task, it is considered as a classification task. If it is predicted whether a test point is abnormal, it is a 2-classification task. If it is necessary to predict the abnormal level at the same time, it is a 5-classification task.
[0078] Therefore, the present application designs two prediction heads related to the task. The first prediction head is a regression head, which is used to regress the light sensitivity threshold of each test point. The second prediction head is a classification head, which is used to determine whether each test point of the pattern deviation probability map is abnormal and the abnormal level.
[0079] Preferably, for the calculation of the light sensitivity threshold map, the light sensitivity threshold interval is usually [0, 40]; in order to train stability and efficiency, the network prediction target is converted to limit in a specific range and normalized to a new range, for example, adjusted to [-0.5, 0.5].
[0080] In the traditional weighting method, the identification of abnormal test points often depends on the preset threshold range, which may lead to insufficient sensitivity in identifying abnormal conditions, especially when the threshold range is wide. However, in the diagnosis of ophthalmic diseases such as glaucoma, the identification of abnormal test points is crucial, as these test points may indicate early development of the disease. The method of the present application can more accurately identify these abnormal conditions by in-depth analysis of the light sensitivity threshold of abnormal test points. This method not only improves the identification ability of abnormal test points, but also more accurately reflects the visual field function status, thereby providing more valuable reference for the diagnosis and treatment of diseases.
[0081] Optionally, the above-mentioned eye index prediction model uses a neural network model as the basic framework, combining regression and classification prediction heads, and can point-by-point predict the light sensitivity threshold and pattern deviation probability points in the visual field test.
[0082] The neural network model referred to above includes network structures such as Transformer and CNN. One implementation case is to use the EVA02 model as the basic framework to extract the representation of OCT data, and use the multi-layer perceptron MLP as the regression and classification prediction head network, and the two tasks use the prediction head network of the exclusive task respectively.
[0083] Four, training process of the eye index prediction model
[0084] For the training process of the above-mentioned eye index prediction model:
[0085] The present application improves the accuracy of the light sensitivity threshold map and the pattern deviation probability map through advanced image processing and machine learning techniques, which helps to more accurately diagnose and monitor eye diseases. Compared to generating a light sensitivity threshold map, generating a pattern deviation probability map is a more difficult task, one of the main challenges is the imbalance of the dataset: in the dataset, the number of samples of normal test points far exceeds that of abnormal test points. Specifically, normal test points account for more than 80% of the total samples, while abnormal test points are scattered in the remaining minority categories. This imbalance makes it easier for the model to learn the characteristics of normal test points during training, but performs poorly in identifying abnormal test points.
[0086] The long-tail problem is a common problem in the field of machine learning, especially in medical image analysis. It is caused by the uneven distribution of real-world data, making it difficult for the model to capture the key features of the minority class. Through case analysis, the long-tail problem reduces the accuracy of the model in identifying abnormal test points of early and intermediate glaucoma, thereby affecting the diagnosis and treatment decisions for patients.
[0087] The strategy for solving the long-tail problem of the present application is as follows:
[0088] 4.1、First, the present embodiment uses oversampling technology to count the abnormal test point data contained in each sample in the training dataset, and increases the representation of samples containing more abnormal test points in the training set by duplicating them.
[0089] However, when training a deep learning model, it is observed that even with oversampling methods, the problem of class imbalance between normal test points and abnormal test points in a training batch cannot be solved, because in early and intermediate glaucoma patients, abnormal test points are relatively few, while normal test points account for the majority. Even with oversampling methods, the problem of class imbalance between normal test points and abnormal test points in a training batch cannot be solved, and in addition, it is observed that in this extreme case, oversampling can easily lead to overfitting of the model, resulting in poor performance on the external test set.
[0090] 4.2、To address this issue, the present embodiment introduces a method of cost-sensitive learning. In the diagnosis of glaucoma, the cost of misclassifying an abnormal test point as normal (false negative) can be much higher than misclassifying a normal test point as abnormal (false positive). Therefore, a strategy is needed to adjust the loss function so that the model pays more attention to those errors with higher cost during training. The present embodiment designs a specific loss function that dynamically adjusts the weight of each class during training to reflect its relative cost, to punish the model for errors in predicting abnormal classes and improve its performance in identifying abnormal test points. Specifically, in each iteration, the actual number of normal test points and abnormal test points in the current minibatch is counted. Then, the loss of normal test points and abnormal test points is calculated respectively, and these losses are divided by the corresponding number of test points.
[0091] Dynamic adjustment of weights in each iteration allows the model to adjust the importance of learning according to the actual number of samples in each class. By this method, the learning process of the model can be adjusted without changing the data set, adapting to the imbalance of the data and paying more attention to those samples that are more critical to clinical decision-making. This method not only improves the accuracy of the model in identifying abnormal test points, but also reduces the risk of overfitting and improves the generalization ability of the model on the external test set.
[0092] Five, evaluation and verification of the model
[0093] The experimental results of the present invention show that the performance of the model in handling the long-tail problem has been significantly improved by adopting the above strategy. This method not only avoids the problem of overfitting, but also improves the generalization ability of the model on the external test set.
[0094] The present invention significantly improves the accuracy of generating light sensitivity threshold by simultaneously optimizing two tasks: generating light sensitivity threshold (regression) and generating mode deviation probability map (classification). For the task of generating mode deviation map, the model not only accurately identifies normal test points, but also effectively identifies abnormal test points, reducing the occurrence of false negatives, which is particularly important in the field of medical diagnosis, because false diagnosis can lead to delayed treatment or inappropriate treatment, which can have a serious impact on the health and well-being of patients.
[0095] In summary, the present invention not only significantly improves the accuracy of generating light sensitivity threshold and improves the accuracy of mode deviation probability map, but also shows great potential in clinical application value. These results strongly support the innovation and practicality of the present invention, providing important technical support for the diagnosis and treatment of glaucoma and other ophthalmic diseases.
[0096] Six, clinical significance
[0097] The most crucial aspect of this invention lies in generating a pattern bias probability map; however, the ultimate value of this invention must be verified through rigorous clinical research.
[0098] In clinical practice, compared to photosensitivity threshold maps, pattern deviation probability maps (PDPs) are among the most important views for glaucoma physicians. These maps provide clinicians with accurate structure-function relationships in an easy-to-understand manner, offering excellent readability and aligning with clinical practice and habits. By interpreting PDPs, physicians can diagnose and monitor diseases such as glaucoma more accurately and efficiently, enabling them to develop more appropriate treatment plans for patients.
[0099] This invention underwent a clinical comparative study: The first group of clinicians used the pattern deviation map predicted by this invention and OCT reports to diagnose glaucoma, while the other group used the pattern deviation probability map from real clinical VF reports and OCT reports to diagnose glaucoma. The diagnostic conclusions of the two groups showed a high degree of consistency, indicating that using the pattern deviation probability map predicted by this invention for auxiliary diagnosis is a promising method in the absence of VF testing or when reliable VF test reports are difficult to obtain. Furthermore, external validation of this invention demonstrated its reliability and applicability in different clinical settings.
[0100] These clinical studies not only demonstrate the clinical effectiveness of pattern bias probability maps but also provide a powerful and user-friendly tool for glaucoma diagnosis, helping physicians make more accurate and timely treatment decisions. This is expected to improve treatment outcomes and quality of life for glaucoma patients, while also offering new perspectives and methods for the management of eye diseases.
[0101] In summary, the present invention has the following characteristics:
[0102] 1. Technical Complexity: 1) Compared to 2D images, OCT volumetric data contains more crucial information about the retina, but the processing and analysis of OCT images involves several technical challenges, including high-dimensional data processing, feature extraction, and computational efficiency. 2) The proportion of normal and abnormal test points in the dataset differs significantly, exhibiting a severe long-tail problem. This imbalance causes the model to perform poorly on a minority of classes during learning and prediction, necessitating the use of cost-sensitive learning and other methods to address this issue.
[0103] 2. Understanding of Clinical Needs: Recognizing the importance of pattern bias probability maps for glaucoma diagnosis is a clinical insight, not a direct technical derivation. This deep understanding of clinical needs led to the development of methods for generating pattern bias probability maps using OCT.
[0104] 3. Strategies for solving the long-tail problem: This invention adopts strategies such as sampling and cost-sensitive learning to solve the long-tail problem of the dataset. These strategies require a deep understanding of the long-tail problem and the development of effective solutions.
[0105] 4. Task Co-optimization: For the two sub-tasks of predicting the photosensitivity threshold and predicting the pattern bias probability map, this invention optimizes both tasks simultaneously and employs the same loss weighting strategy. This method ensures that the loss for each category (classification) or numerical range (regression) receives appropriate attention, thereby improving the model's classification and regression performance in identifying anomalous test points and providing more valuable references for disease diagnosis and treatment. This demonstrates the innovation and practicality of this invention in processing photosensitivity threshold maps and pattern bias probability maps.
[0106] 5. Requirements for Experimental Validation: Generating pattern bias probability maps and validating their clinical value requires a large amount of experimental data and clinical research. The collection and analysis of this data is a time-consuming process, requiring careful experimental design and accurate interpretation of the results.
[0107] Professional knowledge and experience: To ensure this technology can realize its potential in real-world medical settings, close collaboration with clinicians is required to design experiments, collect data, and statistically analyze the results to demonstrate the clinical effectiveness of the pattern bias probability plot. This collaborative effort is essential for the realization of this invention. This interdisciplinary collaboration not only promotes the optimization of the technology but also strengthens clinicians' trust and acceptance of it.
[0108] In addition to the basic technical solution, this invention can also be varied by changing the model architecture, data preprocessing methods, or training strategies. Furthermore, this invention can be applied to the diagnosis of other types of ophthalmic diseases, as long as these diseases are related to structural and visual field function assessment.
[0109] The present invention also provides an ocular index prediction device based on optical coherence tomography, comprising a memory and a processor, wherein the memory stores a computer program, and the processor calls the computer program to execute the steps of the ocular index prediction method based on optical coherence tomography as described above.
[0110] Specifically, this embodiment also provides Figure 4 This is a schematic diagram of an ocular index prediction device based on optical coherence tomography. Figure 4 At the hardware level, the ocular index prediction device based on optical coherence tomography includes a processor, an internal bus, a network interface, memory, and non-volatile memory, and may also include other hardware required for other operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it to achieve the above-mentioned functions. Figure 2The method. Of course, in addition to the software implementation, the present application does not exclude other implementations, such as logic devices or a combination of software and hardware, and the like, that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0111] The preferred embodiments of the present application have been described in detail. It should be understood that modifications and variations can be made by those of ordinary skill in the art without departing from the spirit and scope of the application. Therefore, the technical solutions obtained by logical analysis, reasoning or limited experiments based on the concept of the present application in the prior art should be within the protection scope defined by the claims.
Claims
1. An optical coherence tomography-based method of predicting an eye indicator, the method comprising: obtaining a plurality of images of a subject's eye; determining a plurality of parameters of the subject's eye based on the plurality of images; and predicting the eye indicator based on the plurality of parameters. The method comprises the following steps: obtaining optical coherence tomography volume data and corresponding test reports, and constructing a model training data set; performing data preprocessing on the volume data in the model training data set; constructing an eye index prediction model, which comprises a visual encoder, a regression head and a classification head, performing feature extraction on the volume data through the visual encoder, calculating the light sensitivity threshold of each test point through the regression head according to the extracted features, obtaining a light sensitivity threshold map, and predicting whether each test point is abnormal and the abnormality level through the classification head according to the extracted features, obtaining a pattern deviation probability map; training the eye index prediction model based on the model training data set, and predicting the eye index of the optical coherence tomography data of the to-be-tested object by using the trained eye index prediction model. 2.The method of claim 1, wherein, The data preprocessing process comprises normalizing, scaling, adding random noise and data flipping to the volume data for data enhancement. 3.The method of claim 2, wherein, The data flipping comprises sequentially flipping each slice image of the volume data of the left eye or the right eye of the same object, and horizontally mirroring each slice image. 4.The method of claim 1, wherein, The method further comprises obtaining multiple types of optical coherence tomography volume data, and fusing the volume data of various types of optical coherence tomography of the same object, thereby constructing the model training data set. The fusion method comprises horizontally splicing different types of volume data, or sampling and merging each type of volume data after slicing. The types of optical coherence tomography volume data comprise macular volume data, optic disc volume data and wide field volume data obtained based on optical coherence tomography. 5.The method of claim 1, wherein, The method further comprises counting the abnormal test point data contained in each sample in the model training data set, and expanding the model training data set by copying the samples containing more abnormal test points. 6.The method of claim 1, wherein, The loss function calculation process in the training process of the eye index prediction model comprises: counting the actual number of normal test points and abnormal test points in the current training process, calculating the loss of normal test points and abnormal test points respectively, and dividing the loss of normal test points and abnormal test points by the actual number of normal test points and abnormal test points respectively to obtain the final loss result. 7.The method of claim 1, wherein, In the regression calculation process of the light sensitivity threshold, the calculated light sensitivity threshold is normalized to be limited within a new range. 8.The method of claim 1, wherein, The visual encoder is a 2D encoder or a 3D encoder, the input of the 2D encoder is a plurality of slice images in the volume data, the features corresponding to each slice image output by the 2D encoder are merged to form the features corresponding to the volume data; The input of the 3D encoder is the volume data, and the features corresponding to the volume data are directly obtained. 9.The method of claim 1, wherein, The method further comprises selecting slice images of each volume data by using a regional random sampling strategy, and inputting the selected slice images into the eye index prediction model for prediction.
10. An optical coherence tomography-based eye index prediction device, characterized by, The device comprises a memory and a processor, the memory stores a computer program, and the processor calls the computer program to execute the steps of the method according to any one of claims 1-9.
Citation Information
Patent Citations
A method and system for monitoring and / or assessing pupillary responses
CN107205637A
Data processing method, device and equipment based on visual field inspection report
CN110619332A