A system and method for identifying and referring high-risk patients to ophthalmology specialists.

JP7923071B2Active Publication Date: 2026-09-17JOHNSON & JOHNSON VISION CARE INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022104246
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-06-30
Filing Date
2022-06-29
Publication Date
2026-09-17
Estimated Expiration
2042-06-29

AI Technical Summary

Benefits of technology

【0004】 一次診療医(PCP)から眼科治療専門家に対する視力喪失のリスク患者の特定及び照会は、問題を抱えたままである。2010年の研究では、PCP診療所環境内の眼科スクリーニングへのアクセス不足を含む多くの障壁が特定された。地域によっては、緑内障及び糖尿病性網膜症のリスクがある患者の治療効率を改善するための努力がなされてきた。しかしながら、既存のイニシアチブが、少数の人口統計学的及び同時罹患率パラメータのみに基づいて患者をトリアージする一方で、AMD、白内障、糖尿病性網膜症、緑内障、及びOSDについては、多くの全身関連性が特定されている。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007923071000018
    Figure 0007923071000018
  • Figure 0007923071000019
    Figure 0007923071000019
  • Figure 0007923071000020
    Figure 0007923071000020
Patent Text Reader

Abstract

To provide a computer-implemented method.SOLUTION: A computer-implemented method for identifying one or more patients at risk of having an undetected ophthalmic disease is described. The method may comprise using non-ophthalmic data and pre-processing the data to generate a culled dataset. A model may be trained and tested based on separate portions of the culled dataset. Finally, the model may output, based on the analysis of the data, an indication of the existence or non-existence of one or more ophthalmic diseases.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Traditional methods for identifying and referring at-risk patients from primary care physicians (PCPs) to ophthalmic specialists remain problematic. Many people suffer vision loss as a result of undiagnosed or untreated ophthalmic conditions.

[0002] In the United States alone, for example, an estimated 1.9 million people suffer from vision loss as a result of undiagnosed or untreated ophthalmic conditions. In the case of an estimated 1.2 million of these individuals, the cause is cataracts, and vision can be restored with appropriate consultation with an ophthalmologist. However, in the case of 700,000 Americans, vision loss is due to undiagnosed or untreated age-related macular degeneration (AMD), glaucoma, or diabetic retinopathy, and in the majority of these patients, the vision loss remains irreversible. The impact of poor vision is evident in the exacerbation of comorbidities, particularly the increased risk of disability in patients with cognitive impairment. [Overview of the Initiative] [Problems that the invention aims to solve]

[0003] Improvements are needed. [Means for solving the problem]

[0004] Identifying and referring patients at risk of vision loss from primary care physicians (PCPs) to ophthalmic specialists remains problematic. A 2010 study identified numerous barriers, including a lack of access to ophthalmic screening within PCP clinic settings. In some regions, efforts have been made to improve treatment outcomes for patients at risk of glaucoma and diabetic retinopathy. However, while existing initiatives triage patients based on only a few demographic and comorbidity parameters, many systemic connections have been identified for AMD, cataracts, diabetic retinopathy, glaucoma, and OSD.

[0005] Artificial intelligence (AI) modeling techniques are becoming increasingly important, particularly in ophthalmology and in medicine in general. In ophthalmology, AI is used to calculate intraocular lens (IOL) power, predict the progression of glaucoma, recognize diabetic retinopathy, and classify eye tumors. To our knowledge, AI has not yet been employed to triage primary care patients for ophthalmic referral. This specification reports the development, validation, and testing of multiple predictive AI models for five vision-threatening eye diseases (i.e., patient, AMD, cataract, diabetic retinopathy, glaucoma, and OSD) that may be employed by PCPs to triage patients for referral to ophthalmic treatment specialists.

[0006] This disclosure relates to the identification and referral of at-risk patients from primary care physicians (PCPs) to ophthalmic treatment specialists. As an example, the methods described herein may include a computer-running method for identifying one or more at-risk patients with undetected ophthalmic conditions. A computer or system may receive non-ophthalmic data and preprocess it to generate a selected dataset containing a subset of the non-ophthalmic data. An AI system or model may be trained on at least a first portion of the selected dataset. The model may be tested on at least a second portion of the selected dataset, distinct from the first portion. The model may receive non-ophthalmic patient data and analyze it to determine the presence or absence of one or more ophthalmic conditions. Based on the analysis of the non-ophthalmic patient data, the model may output an index of the presence or absence of one or more ophthalmic conditions.

[0007] One or more computer systems can be configured to perform a specific operation or action by having software, firmware, hardware, or a combination thereof installed on the system that causes the system to perform an action while it is running. One or more computer programs can be configured to perform a specific operation or action by including instructions that cause the device to perform an action when executed by a data processing device. One general embodiment includes a computer execution method for identifying one or more at-risk patients with undetected ophthalmic diseases. The computer execution method also includes receiving non-ophthalmic data; preprocessing the non-ophthalmic data to generate a selected dataset which may include a subset of the non-ophthalmic data; training a model based on at least a first portion of the selected dataset; testing the model based on at least a second portion of the selected dataset which differs from the first portion; receiving non-ophthalmic patient data; using the model to analyze the non-ophthalmic patient data to determine the presence or absence of one or more ophthalmic diseases; and outputting an index of the presence or absence of one or more ophthalmic diseases based on the analysis of the non-ophthalmic patient data. Other embodiments of this model include a corresponding computer system, apparatus, and a computer program stored in one or more computer storage devices, each configured to perform the operation of the method.

[0008] One general embodiment includes a digital health tool for identifying patients at higher risk for the presence of eye disease. The digital health tool also includes a user interface configured to receive patient data which may include non-ophthalmic data, and one or more processors configured to select a model, use the model to analyze non-ophthalmic patient data to determine the presence or absence of one or more eye diseases, and to output an index of the presence or absence of one or more eye diseases. Other embodiments of this embodiment include a corresponding computer system, apparatus, and computer program recorded in one or more computer storage devices, each configured to perform the operation of the method.

[0009] One general embodiment includes a computer execution method for identifying one or more at-risk patients for the presence of an eye disease. The computer execution method also includes selecting a model, using the model to analyze non-ophthalmic patient data to determine the presence or absence of an eye disease, and outputting an index of the presence or absence of an eye disease. Other embodiments of this embodiment include a corresponding computer system, apparatus, and computer program stored in one or more computer storage devices, each configured to perform the operation of the method. [Brief explanation of the drawing]

[0010] The following drawings are illustrative examples, but not limiting, of the various embodiments considered in this disclosure. Brief Description of the Drawings [Figure 1] This section shows the model accuracy for several machine learning algorithms for each disease state. [Figure 2] A box plot showing the most important features of exudative AMD is presented. [Figure 3] A box plot showing the most important features of non-exudative AMD is presented. [Figure 4] A box plot showing the most important features of cataracts is shown. [Figure 5]Shows box plots of the most important features of OSD. [Figure 6] Shows box plots of the most important features of glaucoma. [Figure 7] Shows box plots of the most important features of type 1 PDR. [Figure 8] Shows box plots of the most important features of type 1 NPDR. [Figure 9] Shows box plots of the most important features of type 2 PDR. [Figure 10] Shows box plots of the most important features of type 2 NPDR. [Figure 11] Shows the receiver operating characteristic (ROC) curve for exudative AMD. [Figure 12] Shows the ROC curve for non-exudative AMD. [Figure 13] Shows the ROC curve for cataract. [Figure 14] Shows the ROC curve for OSD. [Figure 15] Shows the ROC curve for glaucoma. [Figure 16] Shows the ROC curve for type 1 PDR. [Figure 17] Shows the ROC curve for type 1 NPDR. [Figure 18] Shows the ROC curve for type 2 PDR. [Figure 19] Shows the ROC curve for type 2 NPDR. [Figure 20] Shows a flow diagram. [Figure 21] Shows a flow diagram. [Figure 22] Shows a flow diagram. [Figure 23] Shows a flow diagram. MODE FOR CARRYING OUT THE INVENTION

[0011] This disclosure relates to a computer execution method for identifying one or more at-risk patients with undetected ophthalmic diseases. A computer or system may receive non-ophthalmic data and preprocess the non-ophthalmic data to generate a selected dataset containing a subset of the non-ophthalmic data. An artificial intelligence (AI) system or model may be trained on at least a first portion of the selected dataset. The model may be tested on at least a second portion of the selected dataset, distinct from the first portion. The model may receive non-ophthalmic patient data and analyze the data to determine the presence or absence of one or more ophthalmic diseases. Finally, the model may output an index of the presence or absence of one or more ophthalmic diseases based on the analysis of the non-ophthalmic patient data.

[0012] This disclosure relates to a digital health tool for identifying high-risk patients for the presence of eye disease. The digital health tool comprises a user interface configured to receive patient data, including non-ophthalmic data. The digital health tool also includes one or more processes that can select a model, analyze non-ophthalmic patient data to determine whether a patient may have one or more eye diseases, and output the analysis.

[0013] This disclosure relates to a computer execution method for identifying patients at risk of eye disease using non-ophthalmic patient data. This method includes selecting a model, using the model to analyze the non-ophthalmic data to determine whether a patient is likely to have an eye disease, and outputting the results to a user.

[0014] Statistical techniques such as ANOVA can provide insights into the relationships between several clinical parameters, but the incorporation of risk stratification and multiple demographic, pharmacological, and comorbidity attributes is well-suited to AI modeling. AI is generally divided into two broad categories, but more categories exist. Machine learning (ML), including decision tree models, organizes parameters (i.e., attributes or features) into layers to predict outcomes. ML is particularly useful for elucidating relationships between clinical parameters. Deep learning (DL) techniques, primarily consisting of neural networks including convolutional neural networks (CNNs), recurrent neural networks (RNNs), and perceptrons, often improve the predictive performance of ML, but at the cost of opacity and interpretability regarding how those predictions are made.

[0015] By building and comparing multiple artificial intelligence (AI) strategies, we obtain models that can be adopted by PCPs to triage patients for referral to ophthalmic treatment specialists.

[0016] Figure 20 shows a diagram of the method. Data based on one or more subjects may be collected in, for example, a database 202. This data may be screened to remove useless or sparse data and to limit outliers in the preprocessing step 204. Before identifying specific models for training and testing, the preprocessed data may be split into two groups: training data 206 and test data 208 (sometimes called validation data). The first dataset, training data 206, may be used to train at least one, and possibly several, models 210. After being trained using the training data 206, the models 210 are tested with the test data 208. The models can output their analyses, and their results can be compared in the analysis comparison step 214. Figure 21 in another diagram shows how the training data 206 may be fed into an untrained model 216 to create a trained model 218. The test data 208 may then be fed into the trained model 218 to create an analysis 212 and a prediction or likelihood that a patient has an eye disease.

[0017] Figure 22 shows one embodiment of cross-validation, which may be a modified form for creating and adjusting the model. In this case, the preprocessed data 204 may be duplicated, and in this embodiment, it may be duplicated five times, from 402 to 410. In each instance of data duplication, the data may be further divided or split into a number of partitions. In this case, the partitions are labeled A, B, C, D, and E. For each duplication of data 402, 404, 406, 408, and 410, the data may be an additional partition 420 (shaded for distinction) that can be used as training data 206 to train an untrained model 216 to create a trained model 218. The trained model 218 is then tested using the remaining partitions. Thus, each model may be trained at least once on each partition. This procedure helps prevent overfitting of the model to the data.

[0018] One or more computer systems can be configured to perform a specific operation or action by having software, firmware, hardware, or a combination thereof installed on the system that causes the system to perform an action while it is running. One or more computer programs can be configured to perform a specific operation or action by including instructions that cause the device to perform an action when executed by a data processing device. Figure 23 shows a flowchart of a computer execution method 2300 for identifying one or more at-risk patients with undetected ophthalmic diseases. The computer execution method includes receiving non-ophthalmic data in 2302. The non-ophthalmic data may be based on one or more subjects. For example, the non-ophthalmic data may be historical patient data collected across multiple subjects. The non-ophthalmic data may be preprocessed to generate a screening dataset which may contain a subset of the non-ophthalmic data. In 2304, the method may include training a model based on at least a first portion of the screening dataset. In 2306, the method may include testing the model based on at least a second portion of the screening dataset which differs from the first portion. In 2308, non-ophthalmic patient data may be received. Non-ophthalmic patient data may be obtained based on target patients. For example, non-ophthalmic data used to train or test a model may be based on one or more subjects that are different from or excluding the target patients. As a further example, a model may be trained and tested on data not associated with target patients. However, other data may also be used. In 2310, the method may include using a model to analyze non-ophthalmic patient data to determine the presence or absence of one or more ophthalmic diseases. In 2312, the method may include outputting an index of the presence or absence of one or more ophthalmic diseases based on the analysis of non-ophthalmic patient data. The model may be updated using additional data. For example, the model may be retrained or retested with new data, and the updated model may be used in the same or similar form as described herein.Other embodiments of this model include a corresponding computer system, apparatus, and a computer program stored in one or more computer storage devices, each configured to perform the operation of the method.

[0019] AI technology generally involves a "training" process, where the importance (i.e., weights) of attributes or median values ​​are adjusted based on a set of data called the training set. The model's performance can then be evaluated against another dataset called the test set. Similar model performance on the training and test sets demonstrates the model's generalizability. The emergence of large clinical databases has enabled the construction and training of both ML and neural network AI models. For this purpose, we employ large, commercially available electronic health record (EHR) databases containing demographic, diagnostic, and treatment data to create and manage an ophthalmology-focused dataset from which predictive models for numerous ophthalmic diseases can be constructed. After comparing several different AI approaches, we selected the creation of a model that could be used by PCPs to triage patients for referral to ophthalmic treatment specialists. The model thus created uses non-ophthalmic clinical and demographic data to assess relative risk scores for AMD, cataracts, OSD, glaucoma, and diabetic retinopathy.

[0020] Abbreviation: AI = Artificial Intelligence, AMD = Age-Related Macular Degeneration, AUC = Area Under the Curve, BMI = Body Mass Index, CNN = Convolutional Neural Network, DL = Deep Learning, EHR = Electronic Health Record, EQUALITY = Improving the Quality and Accessibility of Ophthalmic Care in the Community, GLM = Generalized Linear Model, ICD-10 = International Classification of Diseases, 10th Revision, IOL = Intraocular Lens, ML = Machine Learning, NLP = Natural Language Processing, NPDR = Non-Proliferative Diabetic Retinopathy, OR = Odds Ratio, OSD = Ocular Surface Disease, PCP = Primary Care Physician, PDR = Proliferative Diabetic Retinopathy, ROC = Receiver Operating Characteristics, RNN = Recurrent Neural Network, ECP = Ophthalmic Care Specialist.

[0021] method Data source In one embodiment, a case-control study used data from Optum's Pan-Therapetic EHR (Optum PanTher EHR) database. The Optum PanTher EHR consists primarily of data from the United States, representing clinical information from over 80 million patients, including at least 7 million patients in each US census area. Data from multiple EHR platforms, including Cerner, Epic, GE, and McKesson, is analyzed using natural language processing (NLP) to extract information on diagnosis, biometrics, experimental outcomes, treatments, and medications. The Optum PanTher EHR relies on a network of over 140,000 providers in over 700 hospitals and over 7,000 clinics.

[0022] Evaluation items In this example, the method attempts to predict the diagnosis of five major eye diseases: AMD, cataracts, diabetes mellitus, OSD, glaucoma, and retinopathy. AMD is classified based on the Optum PanTher EHR International Classification of Diseases, 10th Revision (ICD-10) code for Disease, subdivided into non-exudative (H35.31%) and exudative (H35.32%) groups, with "%" representing a wildcard. Cataract classification required a more restrictive definition than simply H25%. Because the ICD-10 code distinguishes between cataracts with a significant impact on vision and those with a less significant impact on vision, the use of cataract surgery was chosen for cataract surgery with a significant impact on vision. In this study, cataracts were defined by the cataract surgery CPT code 66982 or 66984, rather than the ICD-10 code. Diabetic retinopathy was classified based on the Optum PanTher EHR ICD-10 code into type 1 NPDR (H10.31%-H10.34%), type 1 PDR (H10.35%), type 2 NPDR (H11.31%-H11.34%), and type 2 PDR (H11.35%). Glaucoma was defined by the presence of one or more of three criteria: the ICD-10 code H40.1% (open-angle glaucoma), the prescription of glaucoma medication, or the presence of a CPT code indicating glaucoma surgery. This definition was developed to capture not only patients with a glaucoma diagnosis record, but also patients being treated for glaucoma or high-risk ocular hypertension whose glaucoma diagnosis is not recorded in the Optum EHR. Table 1 lists the inclusion criteria for glaucoma. Similar to cataracts, these codes do not distinguish between OSD requiring treatment and milder symptoms; therefore, OSD required a narrower definition than simply H04.1% and H02.88%. In this study, OSD was defined rather restrictively as patients receiving cyclosporine ophthalmic emulsion 0.05%, cyclosporine ophthalmic emulsion 0.09%, or Lifitegrast eye drops 5%.

[0023] [Table 1]

[0024] Machine Learning (ML) Several distinct ML approaches can be used to model the above results. In this example, the approaches included generalized linear models (GLMs), L1 regularized logistic regression, random forests, XGBoost, and J-48 decision trees.

[0025] Example of data preprocessing The Optum PanTher EHR data consisted of 380 attributes, including demographic information, diagnoses, biometrics, experimental outcomes, treatments, and medications. Because some of these attributes, particularly some laboratory tests, may be represented only sparsely, the data may be filtered to eliminate attributes (i.e., ML “features”) with more than 20% missing values. Missing values ​​may be substituted with the median for continuous variables (e.g., BMI), the “missing” group for categorical variables (e.g., smoking or drinking), and the mode for binary variables (e.g., level of laboratory test outcomes). Data winding may be performed by replacing values ​​below the 0.1 percentile with the 0.1 percentile and values ​​above the 99.9 percentile with the 99.9 percentile. Further feature engineering may be performed to eliminate or combine highly correlated features, such as “rheumatoid arthritis / collagen vascular disease” and its highly correlated homogeneous “connective tissue disease.” These feature engineering steps can be performed individually for each case-control dataset for each sub-symptom. In this embodiment, the resulting dataset exhibited 142–182 features after the above selection. Each of the nine feature-excluded datasets for the nine sub-symptoms in this embodiment was modeled using each of five distinct modeling strategies to generate a total of 45 individual ML models. Other machine learning models can also be used in this manner.

[0026] Exemplary Model Strategy The GLM family's "binary" link "logit" or logistic regression can be used to fit the model using maximum likelihood optimization. The outcome predicted from a given set of dependent or independent variables is binary, and therefore logistic regression was chosen. This technique concerns the probability that the dependent variable indicates the occurrence or non-occurrence of an event, in this case, a record of a particular diagnosis. Thus, it is a classification algorithm. Here, assuming that the probability of the event occurring is "p" (p ∈ [0,1]), the probability of the event not occurring is (1-p).

[0027] The logistic regression equation is given as follows:

[0028]

number

[0029] It is worth noting that (p / 1-p) is the odds ratio (OR) for the event to occur. An OR value greater than 1 means the probability of the event occurring is greater than 50%, and therefore more likely than the event not occurring.

[0030] Logistic regression, L1 regularized logistic regression, random forest, and XGBoost models can be used with Python (3.8.5) using, for example, the Scikit-learn (0.23.2) and XGBoost (1.2.0) libraries. In this example, 80% of the data was used for training and 20% was used for testing with 5x cross-validation (Figure 22). Hyperparameters can be optimized using grid search. For L1 regularized logistic regression, the regularization strength C can be adjusted. For the random forest algorithm, the space of the number of trees and the maximum depth of each tree combination can be searched. Hyperparameter tuning for XGBoost can include the learning rate and maximum depth of each tree. A machine learning modeling pipeline can be established, and information on missing values ​​fitted and learned from the training data can be applied to the test dataset to avoid information loss. Java-based implementations of J48 decision tree modeling and C4 trees can be run on the WEKA ML Workbench (University of Waikato, Hamilton, New Zealand). 10x cross-validation can be employed with an initial leaf size of 2% of the dataset.

[0031] Exemplary results The population for the case-control study varied depending on the disease stage, ranging from 395,140 cases of cataracts that significantly affect vision to 7,440 cases of OSD treated with Refitegrast or cyclosporine (Table 2). The outcomes of different ML strategies also varied (Figure 1, Table 3 overview, Table 4 details), but XGBoost demonstrated the best results in all cases, with predictive accuracy and AUC of 77.4% and 0.858 for exudative AMD, 79.2% and 0.879 for non-exudative AMD, 78.6% and 0.878 for cataracts affecting visual acuity, 72.2% and 0.803 for OSD requiring medication, 70.8% and 0.785 for glaucoma, 82.2% and 0.911 for type 1 PDR, 85.0% and 0.924 for type 1 NPDR, 82.1% and 0.900 for type 2 PDR, and 81.3% and 0.891 for type 2 NPDR (Table 4). XGBoost identified several clinical attributes important for diagnostic prediction (Figure 3).

[0032] [Table 2]

[0033] [Table 3]

[0034] The top-performing models in this embodiment, which primarily contributed to the prediction of each disease state, are mentioned here and quantified in the box plots in Figures 2-10. The prediction of exudative AMD was associated, in order of importance, with average household income, college education rate, geographical region (Mid-Atlantic, Northeast Central, Southeast Central, New England, Southern Atlantic / South-Central West, Mountainous, Northwest Central, Pacific, Unknown / Other), body fat index (BMI), and Elixhauser score (comorbidity index). (Figure 2) Non-exudative AMD demonstrated similar associations. In order of importance, these were average household income, college education rate, region (Northeast, Midwest, South, West, Other / Unknown), smoking, and Elixhauser score. (Figure 3) The clinical associations for cataracts, in order of importance, included average household income, college education rate, region, BMI, and smoking. (Figure 4) The associations with OSD, in order of importance, included average household income, geographical location, rheumatoid arthritis, connective tissue disease, and region. (Figure 5) The clinical associations for glaucoma, in order of importance, included average household income, college education rate, adrenal or androgen use, BMI, and race. (Figure 6) The associations with diabetic retinopathy varied across different sub-conditions (Type 1 PDR, Type 1 NPDR, Type 2 PDR, Type 2 NPDR), but generally included Elixhauser score, hyperglycemia, BMI, hypertension, chronic lung disease, depression, cardiac arrhythmias, and obesity. (Figures 7-10)

[0035] The overall results of each XGBoost model in this embodiment, including performance and correlation, are shown in Table 4 below.

[0036] [Table 4-1]

[0037] [Table 4-2]

[0038] The details of the AUC in this embodiment are shown for each lesion, one by one, in the relevant ROC curves shown in Figures 11-19.

[0039] Consideration Performance of the embodiment model Starting with EHR data from over 80 million patients, the final study population totaled 1,486,078 patients, 50% of whom were controls. In addition to this enormous patient population, this embodiment demonstrated 90 different AI models for five primary conditions and nine sub-conditions to arrive at the best predictive model for each condition.

[0040] The goal of this project is to identify patients at high risk for the presence of eye disease and to create a digital health tool that does so based solely on the types of non-ophthalmic data that PCPs will have access to. This digital health tool does not propose to make definitive ophthalmic diagnoses or predict the onset of future lesions. Rather, this digital health tool seeks to identify patients whose clinical and demographic context is associated with the presence of AMD, cataracts, clinically significant diabetic retinopathy, glaucoma, or OSD disease to a degree requiring drug therapy.

[0041] The predictive accuracy for the presence of lesions in this embodiment ranged from 71% in glaucoma cases to 87% in type 1 proliferative diabetic retinopathy cases, with an average accuracy of 80% across all groups. Since the objective is to identify patients at risk, these results can be used to determine the disease odds ratio, for example, by the method described by Hogue, Gaylor, and Schulz, as explained in Altman, Douglas G., Practical Statistics for Medical Research, Chapman & Hall (1991). Because the case-control study populations for each condition were uniformly divided into affected and controlled groups, random selection of patients results in a 50% probability of being affected. If the model runs with 80% accuracy, it essentially identifies a population with an 80% risk of developing the disease. The calculation of the odds ratio (θ) is as follows:

[0042]

number

[0043]

number

[0044] Applying this to each model provides a clinically useful measure. The model in this embodiment identifies patients with high prevalence odds ratios ranging from 2.44 for glaucoma cases to 6.58 for type 1 proliferative diabetic retinopathy cases, as shown in Table 3, with a mean odds ratio of 4. The application of such a model in a clinical setting allows PCPs to identify patients who are nearly four times more likely to have an eye disease. Such a tool would provide substantial benefits in triage for referral of at-risk patients to ophthalmic treatment specialists.

[0045] Exemplary Data and Result Processing The data used to generate and test these models in this embodiment were obtained from the Optum Pan-Thrapeutic EHR database (Optum PanTher EHR), but other databases can also be used. This data consists of diagnostic and treatment codes, biometric data such as BMI and vital signs, demographic information including socioeconomic and geographical information, experimental results, and prescribed medications. This information does not include physician notes that may provide the basis for recorded diagnoses. In practice, only a very limited number of diagnoses can be listed in the claims, so some existing diagnoses may not be recorded. On the other hand, since the ICD-10 classification does not distinguish between clinically significant cataracts and OSDs where these lesions are subclinical, diagnoses such as cataracts and OSDs may be overrepresented if these lesions were clinical. In fact, there is little clinical utility in building AI models to detect subclinical cataracts.

[0046] This embodiment demonstrates the challenge of identifying clinically relevant diagnoses from large datasets. A 2018 study in JAMAO pharmology examined the accuracy of ICD-10 codes for uveitis patients and found that 13 of 27 uveitis cases were inaccurately defined and that multiple codes were used to describe the same lesion. A 2020 study of ophthalmic diseases in stroke patients pointed out that there were fewer glaucoma patients than expected, which was attributed to a lack of ophthalmic clinical data. Patients may be taking glaucoma medications without a co-documented ICD code for glaucoma, suggesting that a diagnosis of glaucoma may have been documented in the patient's medical records before incorporation into the dataset. Therefore, the definition of the glaucoma cohort in this embodiment was expanded to include patients meeting one or more of three criteria: H40.1% ICD-10 code (open-angle glaucoma), prescription of glaucoma medication, or presence of a CPT code indicating glaucoma surgery (see Table 1). This definition was developed with the aim of both detecting glaucoma patients without an ICD-10 code and excluding patients who were inappropriately classified as having glaucoma by ICD-10. As a result of this definition, the substantial screening of the glaucoma cohort patients was reduced from 1,368,700 (50% controls) to 385,514. Similar data preprocessing may be required in other databases to include all patients who may be at risk.

[0047] A similar approach can be applied to the populations of cataract and OSD studies. Cataracts and OSDs are the most frequently recorded diagnoses in claims. Cataracts are nearly ubiquitous, especially in elderly patients, and were the most common ophthalmic ICD-10 diagnosis examined in this embodiment. Since only a fraction of these require cataract surgery, detecting cataracts alone is not clinically useful. The ICD-10 code does not distinguish between cataracts requiring surgery and those that do not. However, the CPT code does make this distinction. Therefore, as criteria for clinically significant cataracts, we selected 66984 CPT (cataract extraction with intraocular lens) and 66982 (comorbid cataract extraction). In this embodiment, narrowing the inclusion criteria in this way reduced the patient population for the cataract study from 2,087,836 (50% controls) to 395,140. The OSD code presents further challenges. Numerous ICD-10 codes are available, making it difficult to establish clinical significance. The initial cohort of OSD patients and controls in this database totals 1,182,912 patients. To model the clinical context associated with OSD, a limited criterion of topical cyclosporine or rifegrast prescription was selected. This significantly reduces the OSD population to 7,440 patients, but these patients represent those with clinically significant disease. Outcome engineering techniques were not applied to the AMD group or diabetic retinopathy group as defined by the corresponding ICD-10 code.

[0048] Examples of clinical attribute and feature engineering The initial dataset in this embodiment contained numerous attributes, or (in ML terminology) "features," totaling 380 individual parameters. To generate models that are not burdensome for clinicians to adopt, the number of attributes required by each model was reduced. This reduction and modification of model parameters is referred to as "feature engineering." Several criteria must be met for features to be included in the final model. Features must play a significant role in the model's results. It is self-evident that features that do not substantially contribute to the model can be discarded with little impact on model performance. In the case of the XGBoost model, parameter optimization was performed using a grid search algorithm. The second feature inclusion criterion was that there was no correlation with other features. In some cases, correlations are clear, such as between weight and BMI. However, correlations between other clinical features simply become apparent during analysis. The issue of feature correlation highlights the difference between AI and traditional risk analysis research. When studied individually, certain attributes, such as obesity or socioeconomic status, may be identified as disease risk factors. However, when viewed collectively, if two attributes are highly correlated, the importance of one of them may be reduced. The third feature inclusion criterion was high frequency in the dataset. Some experimental values, particularly serum fibrinogen, are very sparse in this particular dataset, so excluding this feature was preferred as an alternative to sample reduction or interpolation. In this embodiment, two thresholds were used for feature sparsity. The model was built on a dataset from which features with more than 20% missing values ​​were excluded. Feature engineering benefited substantially from guidance from experts in the clinical field, and our feature and outcome engineering was based on clinical information, particularly within the scope of the diagnostic criteria described above.

[0049] Usage data and generalization The data in the above examples does not possess the richness of complete medical records. Therefore, it is impossible to establish criteria for more stringent standards, such as criteria for clinicians to record diagnoses, and consequently, outcome engineering tools, to identify clinically significant cataract patients using CPT codes for cataract surgery. At the same time, models built on these types of data are more specific and perhaps more general and readily available than models built on more specific data sources. These models are precisely the type of data available to PCPs, making them easier to implement than models built on specific medical record systems. Indeed, the availability of this data is illustrated by the above examples, which include over 80 million patients from heterogeneous healthcare systems.

[0050] Hierarchical relationships Furthermore, clinical features identified as associated with each disease state model should be considered correlated, not necessarily causal. It is more appropriate to consider the set of clinical values ​​as part of the patient's clinical environment rather than as a collection of individual risk factors. While it is difficult to imagine that university education itself is a risk factor for disease, its correlation and importance to a given model should not be ignored, as it contributes to predicting the presence of the disease state in the above examples.

[0051] It goes without saying that there may be no causal relationship between some of these features and the modeled disease state. Advanced multidimensional clinical AI research, such as the example above, can identify previously unrecognized factors that directly influence the disease state. However, causal relationships cannot be established by this type of research and will require a more traditional experimental approach. The J-48 decision tree model did not work as well as the GLM or XGBoost strategies in the example case, but it is useful in explaining hierarchical relationships between clinical features. As an example, the J-48 model for glaucoma identifies race, systemic steroid use, and antidiabetic drug use as important clinical features. However, the model dictates the order in which these factors should be considered, evaluating race only after it was determined whether the patient was taking antidiabetic drugs, and evaluating systemic steroid use only after these first two attributes were determined. Establishing such hierarchical relationships between clinical features can be extremely difficult in conventional dimensionality-reducing scientific queries. This Gestalt approach to multidimensional clinical contexts is one of the strengths of this method.

[0052] prediction The purpose of these models is prediction. However, a clear understanding of “prediction” must first be established in order to properly apply the work. These models predict the presence of existing disease conditions. They are important in identifying populations in which these conditions are substantially more prevalent than in the general population. The models should not be used to diagnose individual patients, but rather to identify patients at risk of having undetected AMD, cataracts, diabetic retinopathy, glaucoma, or OSD. Furthermore, these models are built on clinical data in which eye disease is present or absent. That is, these models are not configured to predict the future onset of disease conditions. As defined by the multidimensional features incorporated into the models, a particular clinical context may or may not be able to predict the future onset of disease. It would be inappropriate to use these models for pure diagnostic purposes. These models predict the presence of eye disease based on non-ophthalmic data and may be best suited for triage and referrals from non-ophthalmologists to ophthalmic treatment specialists. Other uses are also intended.

[0053] This disclosure includes at least the following aspects:

[0054] Embodiment 1. A computer execution method for identifying one or more at-risk patients having an undetected ophthalmic disease, comprising: receiving non-ophthalmic data; preprocessing the non-ophthalmic data to generate a selected dataset including a subset of the non-ophthalmic data; training a model based on at least a first portion of the selected dataset; testing the model based on at least a second portion of a computed dataset different from the first portion; receiving non-ophthalmic patient data; using the model to analyze the non-ophthalmic patient data to determine the presence or absence of one or more ophthalmic diseases; and outputting an index of the presence or absence of one or more ophthalmic diseases based on the analysis of the non-ophthalmic patient data.

[0055] Aspect 2. The method of Aspect 1, wherein the non-ophthalmic patient data is based on the target patient, and the non-ophthalmic data is based on one or more subjects different from the target patient. The non-ophthalmic data may be based on one or more subjects excluding the target patient.

[0056] Aspect 3. The method of Aspect 1, wherein the one or more ophthalmic diseases comprise age-related macular degeneration (AMD), cataract, diabetic retinopathy, glaucoma, or ocular surface disease (OSD).

[0057] Aspect 4. The method of Aspect 1, wherein the pre-processing comprises feature engineering.

[0058] Aspect 5. The method of Aspect 4, wherein the feature engineering comprises eliminating or combining highly correlated features.

[0059] Aspect 6. The method of Aspect 1, wherein the pre-processing comprises eliminating one or more attributes having more than 20% of missing values.

[0060] Aspect 7. The method of Aspect 1, wherein the pre-processing comprises replacing values less than the 0.1th percentile value with the 0.1th percentile value, and replacing values greater than the 99.9th percentile value with the 99.9th percentile value.

[0061] Aspect 8. The method of Aspect 1, wherein the model is based on at least a logistic regression model.

[0062] Aspect 9. The model is based on at least the following logistic regression equation:

[0063] [Numerical formula] wherein, Y is a dependent variable, X i is an independent variable, β0 is a Y-intercept of a population, βi is a regression coefficient between the dependent variable and the corresponding independent variable (X iThe method according to embodiment 1, which is the gradient value of the line drawn between ) and ).

[0064] Embodiment 10. A digital health tool for identifying high-risk patients for the presence of eye disease, comprising: a user interface configured to receive patient data including non-ophthalmic data; and one or more processors configured to select a model, use the model to analyze non-ophthalmic patient data to determine the presence or absence of one or more ophthalmic diseases, and to output an indicator of the presence or absence of one or more ophthalmic diseases.

[0065] Embodiment 11. The digital health tool according to Embodiment 10, wherein one or more ophthalmic diseases include age-related macular degeneration (AMD), cataracts, diabetic retinopathy, glaucoma, or ocular surface diseases (OSD).

[0066] Embodiment 12. The digital health tool according to Embodiment 10, wherein the model is based on at least a logistic regression model.

[0067] Embodiment 13. The model is based on at least the following logistic regression equation:

[0068]

number

[0069] Embodiment 14. A computer execution method for identifying one or more at-risk patients for the presence of an eye disease, comprising: selecting a model; using the model to analyze non-ophthalmic patient data to determine the presence or absence of an eye disease; and outputting an indicator of the presence or absence of an eye disease.

[0070] Embodiment 15. The method according to claim 14, wherein the non-ophthalmic patient data is based on the target patient, and the model is based on non-ophthalmic data associated with one or more subjects different from the target patient. The non-ophthalmic data may be based on one or more subjects other than the target patient.

[0071] Embodiment 16. The method according to Embodiment 14, wherein the eye disease includes age-related macular degeneration (AMD), cataracts, diabetic retinopathy, glaucoma, or ocular surface disease (OSD).

[0072] Embodiment 17. The method according to Embodiment 14, wherein the eye disease includes one or more variables of non-ophthalmic data that correlate with the risk of age-related macular degeneration (AMD), cataracts, diabetic retinopathy, glaucoma, or ocular surface disease (OSD).

[0073] Embodiment 18. The method according to Embodiment 14, further comprising preprocessing non-ophthalmic patient data.

[0074] Embodiment 19. The method according to Embodiment 18, wherein preprocessing includes feature engineering.

[0075] Embodiment 20. The method according to Embodiment 19, wherein the feature engineering includes eliminating or combining highly correlated features.

[0076] Embodiment 21. The method according to Embodiment 18, wherein preprocessing includes removing one or more attributes having more than 20% missing values.

[0077] Embodiment 22. The method according to Embodiment 18, wherein the preprocessing includes replacing values ​​below the 0.1 percentile with the 0.1 percentile and replacing values ​​above the 99.9 percentile with the 99.9 percentile.

[0078] Embodiment 23. The method according to Embodiment 14, wherein the model is based on at least a logistic regression model.

[0079] Apparatus 24. The model is based on at least the following logistic regression equation:

[0080]

number

[0081] The embodiments illustrated and described herein are considered to be the most practical and preferred embodiments; however, it will be clear to those skilled in the art that modifications from the specific designs and methods illustrated and disclosed herein are themselves obvious to them and can be used without departing from the spirit and scope of the invention. For example, the systems, devices, and methods for predicting ophthalmic diagnoses described herein are based on non-ophthalmic data. It will be understood by those skilled in the art that the devices and methods described herein are not limited to this area and can be used in other diagnostic areas. The invention is not limited to the specific configurations described and illustrated, but should be configured to be consistent with all modifications that may be included in the appended claims.

[0082] [Implementation Method] (1) A computer execution method for identifying one or more at-risk patients with undetected ophthalmic diseases, Receiving non-ophthalmic data, The non-ophthalmic data is preprocessed to generate a selected dataset that includes a subset of the non-ophthalmic data. The model is trained based on at least the first portion of the selected dataset, Testing the model based at least on a second portion of the selected dataset, which is different from the first portion, Receiving non-ophthalmic patient data, Using the aforementioned model, the non-ophthalmic patient data is analyzed to determine the presence or absence of one or more ophthalmic diseases. A method comprising outputting an indicator of the presence or absence of one or more ophthalmic diseases based on the analysis of the aforementioned non-ophthalmic patient data. (2) The method according to Embodiment 1, wherein the non-ophthalmic patient data is based on the target patient, and the non-ophthalmic data is based on one or more subjects different from the target patient. (3) The method according to Embodiment 1, wherein the one or more ophthalmic diseases include age-related macular degeneration (AMD), cataracts, diabetic retinopathy, glaucoma, or ocular surface diseases (OSD). (4) The method according to Embodiment 1, wherein the preprocessing includes feature engineering. (5) The method according to Embodiment 4, wherein the feature engineering includes eliminating or combining highly correlated features.

[0083] (6) The method according to Embodiment 1, wherein the preprocessing includes eliminating one or more attributes having more than 20% missing values. (7) The method according to Embodiment 1, wherein the pretreatment includes replacing values ​​less than the 0.1 percentile value with the 0.1 percentile value and replacing values ​​greater than the 99.9 percentile value with the 99.9 percentile value. (8) The method according to Embodiment 1, wherein the model is based on at least a logistic regression model. (9) The model is based on at least the following logistic regression equation:

[0084]

number

[0085] (11) The digital health tool according to Embodiment 10, wherein the one or more ophthalmic diseases include age-related macular degeneration (AMD), cataracts, diabetic retinopathy, glaucoma, or ocular surface diseases (OSDs). (12) The digital health tool according to Embodiment 10, wherein the model is based on at least a logistic regression model. (13) The model is based on at least the following logistic regression equation:

[0086]

number

[0087] (16) The method according to Embodiment 14, wherein the eye disease includes age-related macular degeneration (AMD), cataracts, diabetic retinopathy, glaucoma, or ocular surface disease (OSD). (17) The method of Embodiment 14, wherein the eye disease includes one or more variables of the non-ophthalmic data that correlate with the risk of age-related macular degeneration (AMD), cataracts, diabetic retinopathy, glaucoma, or ocular surface disease (OSD). (18) The method according to Embodiment 14, further comprising preprocessing the non-ophthalmic patient data. (19) The method according to embodiment 18, wherein the preprocessing includes feature engineering. (20) The method according to embodiment 19, wherein the feature engineering includes eliminating or combining highly correlated features.

[0088] (21) The method according to Embodiment 18, wherein the preprocessing includes eliminating one or more attributes having more than 20% missing values. (22) The method according to Embodiment 18, wherein the preprocessing includes replacing values ​​less than the 0.1 percentile value with the 0.1 percentile value and replacing values ​​greater than the 99.9 percentile value with the 99.9 percentile value. (23) The method according to Embodiment 14, wherein the model is based on at least a logistic regression model. (24) The above model is based on at least the following logistic regression equation:

[0089]

number

Claims

1. A computer execution method for determining whether a patient is at high risk of developing an undetected ophthalmic disease, Receiving non-ophthalmic data related to one or more subjects other than the target patient, wherein the non-ophthalmic data includes demographic data, diagnostic data, and treatment data. The non-ophthalmic data is preprocessed to generate a selected dataset that includes a subset of the non-ophthalmic data. The model is trained based on at least the first portion of the selected dataset, Testing the model based on at least a second portion of the selected dataset, which is different from the first portion, Receiving non-ophthalmic patient data based on the aforementioned target patient, wherein the non-ophthalmic patient data includes demographic data, diagnostic data, and treatment data. Using the aforementioned model, the non-ophthalmic patient data is analyzed to determine the presence or absence of one or more ophthalmic diseases. A method comprising outputting an indicator of the presence or absence of one or more ophthalmic diseases based on the analysis of the aforementioned non-ophthalmic patient data.

2. The method according to claim 1, wherein the non-ophthalmic data and the non-ophthalmic patient data further include average household income and the percentage of university education in the residential area.

3. The method according to claim 1, wherein the one or more ophthalmic diseases include age-related macular degeneration (AMD), cataracts, diabetic retinopathy, glaucoma, or ocular surface diseases (OSD).

4. The method according to claim 1, wherein the preprocessing includes feature engineering.

5. The method according to claim 4, wherein the feature engineering includes eliminating or combining highly correlated features.

6. The method according to claim 1, wherein the preprocessing includes excluding one or more attributes having more than 20% missing values.

7. The method according to claim 1, wherein the preprocessing includes replacing values ​​less than the 0.1 percentile value with the 0.1 percentile value and replacing values ​​greater than the 99.9 percentile value with the 99.9 percentile value.

8. The method according to claim 1, wherein the model is based on at least a logistic regression model.

9. The aforementioned model is based on at least the following logistic regression equation: [Math 1] During the ceremony, Y is the dependent variable, X i is an independent variable, β 0 This is the Y-intercept of the population, βi is the dependent variable and the corresponding independent variable (X i The method according to claim 1, wherein the gradient value of the line drawn between )

10. A digital health tool for identifying whether a patient is at high risk of developing an undetected ophthalmic disease, Patient data including non-ophthalmic data related to one or more subjects different from the target patient, including demographic data, diagnostic data, and treatment data, and a user interface configured to receive non-ophthalmic data, One or more processors, Select a model trained on at least the aforementioned non-ophthalmic data, Using the aforementioned model, non-ophthalmic patient data based on the target patient, including demographic data, diagnostic data, and treatment data, is analyzed to determine the presence or absence of one or more ophthalmic diseases, and Outputs an indicator of the presence or absence of one or more ophthalmic diseases. A processor configured as follows, A digital health tool equipped with these features.

11. The digital health tool according to claim 10, wherein the one or more ophthalmic diseases include age-related macular degeneration (AMD), cataracts, diabetic retinopathy, glaucoma, or ocular surface diseases (OSD).

12. The digital health tool according to claim 10, wherein the model is based on at least a logistic regression model.

13. The aforementioned model is based on at least the following logistic regression equation: [Math 2] During the ceremony, Y is the dependent variable, X i is an independent variable, β 0 This is the Y-intercept of the population, βi is the dependent variable and the corresponding independent variable (X i The digital health tool according to claim 10, which is the gradient value of the line drawn between ) and ).

14. A method for determining whether a patient is at high risk of developing an undetected ophthalmic disease, Select a model trained on non-ophthalmic data, which includes demographic data, diagnostic data, and treatment data, relating to one or more subjects different from the target patient. Using the aforementioned model, non-ophthalmic patient data based on the target patient, including demographic data, diagnostic data, and treatment data, is analyzed to determine the presence or absence of eye disease. A method comprising outputting an indicator of the presence or absence of the aforementioned eye disease.

15. The method according to claim 14, wherein the non-ophthalmic data and the non-ophthalmic patient data further include average household income and the percentage of university education in the residential area.

16. The method according to claim 14, wherein the eye disease includes age-related macular degeneration (AMD), cataracts, diabetic retinopathy, glaucoma, or ocular surface disease (OSD).

17. The method according to claim 14, wherein the eye disease includes one or more variables of the non-ophthalmic data that correlate with the risk of age-related macular degeneration (AMD), cataracts, diabetic retinopathy, glaucoma, or ocular surface disease (OSD).

18. The method according to claim 14, further comprising preprocessing the non-ophthalmic patient data.

19. The method according to claim 18, wherein the preprocessing includes feature engineering.

20. The method according to claim 19, wherein the feature engineering includes eliminating or combining highly correlated features.

21. The method according to claim 18, wherein the preprocessing includes excluding one or more attributes having more than 20% missing values.

22. The method according to claim 18, wherein the preprocessing includes replacing values ​​less than the 0.1 percentile value with the 0.1 percentile value and replacing values ​​greater than the 99.9 percentile value with the 99.9 percentile value.

23. The method according to claim 14, wherein the model is based on at least a logistic regression model.

24. The aforementioned model is based on at least the following logistic regression equation: [Math 3] During the ceremony, Y is the dependent variable, X i is an independent variable, β 0 This is the Y-intercept of the population, βi is the gradient value of a line drawn between said dependent variable and said corresponding independent variable (X i ), the method according to claim 14.

Citation Information

Patent Citations

  • Biomarkers and predictive methods

    JP2017533428A

  • Method for determining the risk of glaucoma

    JP2020178555A

  • Medical support system and ophthalmologic examination device

    JP2020201531A

  • Medical information processing system and medical information processing method

    JP2021051776A

  • Using deep learning to process images of the eye to predict visual acuity

    WO2021026039A1