A method for predicting, screening, an electronic device and a program product for a risk of fabry disease
By constructing a Fabry disease risk prediction model based on expert knowledge and combining multi-source medical data and natural language processing technology, the problems of data standardization and model generalization in Fabry disease diagnosis are solved, realizing efficient and accurate screening and early warning of Fabry disease, and suitable for standardized promotion in different medical environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHONGSHAN HOSPITAL FUDAN UNIV
- Filing Date
- 2026-05-14
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies for the diagnosis of Fabry disease suffer from difficulties in data standardization, insufficient model generalization ability, poor interpretability, and low cost-effectiveness, making them difficult to apply widely in the real world.
A Fabry disease risk prediction model based on expert knowledge is adopted, combined with multi-source heterogeneous medical data and natural language processing technology. Entity recognition and relation extraction are performed through convolutional neural networks and TextCNN to construct a Fabry disease risk prediction system, which is then integrated into a clinical decision support system to achieve real-time risk assessment and early warning.
It improves the efficiency and accuracy of early screening for Fabry disease, helps doctors provide timely diagnostic advice, reduces the rate of missed diagnoses and saves labor costs, and is suitable for standardized promotion in different medical environments.
Smart Images

Figure CN122494249A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of rare disease screening technology, and in particular to a Fabry disease risk prediction, screening method, electronic device and program product. Background Technology
[0002] Fabry disease is a rare inherited lysosomal storage disorder, listed in the rare disease catalog. It is caused by a deficiency of α-galactosidase A (α-Gal A), leading to the abnormal accumulation of glycosphingolipids in various tissues and organs throughout the body, causing multi-system damage, severely impacting patients' quality of life and shortening their lifespan by 15-20 years. Because the disease can affect multiple organs, including the nervous system, kidneys, heart, skin, gastrointestinal tract, and eyes, its clinical manifestations exhibit significant heterogeneity. Furthermore, due to its low prevalence, doctors have relatively insufficient awareness of the disease, leading to frequent misdiagnosis, missed diagnosis, and delayed diagnosis. Data shows that the time from symptom onset to diagnosis in adult Fabry disease patients can be as long as 10.5 years, and even as long as 14.8 years. Previous efforts to improve diagnostic efficiency through methods such as strengthening physician education and promoting newborn screening have yielded limited results. In recent years, artificial intelligence (AI) methods have provided new options for the diagnosis of rare diseases.
[0003] Existing research has established a large-sample Fabry disease cohort (4978 patients) and matched data from 5000 healthy individuals. Machine learning was used to extract disease phenotypic features to build a diagnostic model. External testing showed that the model achieved an area under the receiver operating characteristic (AUC) of 0.82, demonstrating excellent recognition ability. However, the biggest limitation of this method is the requirement for high-quality, large-sample raw data, making it difficult to generalize in the field of rare diseases.
[0004] In subsequent research, to address the challenge of training models with insufficient rare disease cases, a novel image enhancement method was developed. This method artificially expands limited image data to improve the generalization ability of the algorithm model. This model can identify undiagnosed Fabry disease patients from images of urine sediment, achieving a sensitivity of 0.90 and an AUC of 0.968. However, the application of these machine modeling methods in the real world is still not feasible and remains limited.
[0005] In real-world clinical settings, the complexity and diversity of data can pose challenges to the accuracy and reliability of models. For example, different medical institutions may have different recording methods, making data standardization difficult; patient medical records may be incomplete or biased, affecting the quality of model input; furthermore, the incidence and phenotypic manifestations of Fabry disease may vary in different regions, which could affect the model's generalization ability.
[0006] To evaluate the effectiveness of these models in the real world, large-scale clinical validation studies are needed. This includes testing the models in different regions, populations, and healthcare settings to ensure they maintain high accuracy and stability under various conditions. Simultaneously, the interpretability of the model must be considered, meaning that doctors and patients can understand the model's predictions, which is crucial for the model's acceptance and practical application.
[0007] In summary, while machine learning models have shown great potential in identifying rare diseases such as Fabry disease, their widespread real-world application still faces challenges in areas such as data standardization, model generalization ability, interpretability, real-time updates, and cost-effectiveness. Meanwhile, the rapid development of artificial intelligence (AI) has led to the emergence of relatively mature technologies such as natural language processing, knowledge graphs, and machine learning, which are driving the continuous development of various new Clinical Decision Support Systems (CDSS). CDSS is gradually being applied in fields such as healthcare management and clinical medicine, providing intelligent services such as treatment recommendations, disease warnings, medical order monitoring, and medical record quality control during the medical process. This helps to effectively reduce medical errors and costs, and improve doctors' work efficiency. Among these, CDSS research in disease screening and diagnosis has received considerable attention and has been successfully applied in common disease areas such as venous thromboembolism (VTE), sepsis, and cardiovascular diseases, providing new methods for solving the challenges of disease identification among doctors of different departments and experience levels.
[0008] As an important means to improve patient safety and quality of care, many countries regard CDSS as a key module for the effective use of electronic health record systems and have comprehensively deployed basic medical research, regulatory standards, knowledge base and database construction to improve CDSS research and application. However, due to its relatively short exploration time in disease prediction, especially in the field of rare diseases, its application is lacking. Summary of the Invention
[0009] One aspect of this disclosure is a method for predicting the risk of Fabry disease, which uses a digital screening software system to predict the risk of Fabry disease by employing a Fabry disease risk prediction model constructed based on expert knowledge.
[0010] The risk prediction model includes multiple high-risk signs and their corresponding weights. These high-risk signs include left ventricular hypertrophy, symmetrical hypertrophy, papillary muscle hypertrophy, bilateral sign, reduced longitudinal strain, shortened PR interval, delayed gadolinium enhancement of the inferior lateral wall of the left ventricle, decreased native T1 value, history of dialysis, renal insufficiency, corneal whorled opacity, neuropathic burning pain, neuropathic hearing loss, sweating disorders, keratoma, cerebrovascular complications, and family history.
[0011] One aspect of this disclosure is a Fabry disease risk prediction system, comprising:
[0012] The data acquisition module is used to acquire multi-source heterogeneous medical data of patients across different visits and departments from the hospital information system, including structured data and unstructured text data;
[0013] The natural language processing module has a built-in entity recognition and relation extraction model based on convolutional neural networks and TextCNN, which is used to extract medical entities and standardize the mapping of the unstructured text.
[0014] The risk prediction module has a built-in Fabry disease risk prediction model based on expert knowledge, which includes multiple high-risk signs and their weight configurations. It is used to calculate risk scores and perform risk stratification based on the extracted medical entities.
[0015] The real-time alert module is used to monitor changes in patient data, trigger incremental calculations of risk scores, and push alert information to the doctor's workstation in real time.
[0016] A clinical decision support interface is used to integrate with the hospital's electronic medical record system, displaying risk assessment results in real time while doctors are writing medical records.
[0017] One aspect of this disclosure is a Fabry disease screening method based on a clinical decision support system, comprising:
[0018] When doctors write and save admission records or initial medical records, they collect patients' current medical data and historical medical data in real time.
[0019] Automated risk assessment is performed using the Fabry disease risk prediction method as described in claim 1;
[0020] When the risk score reaches the first preset threshold, the Fabry high-risk screening assessment form and recommendations for echocardiography and cardiac magnetic resonance imaging are displayed in real time on the doctor's workstation interface.
[0021] When the risk score reaches the second preset threshold, the doctor will be simultaneously reminded to perform Fabry enzyme and genetic testing.
[0022] It provides a function to trace the basis of the assessment, displaying the original medical records on which the risk score is based, and supports doctors to manually correct the assessment results.
[0023] This disclosure presents a system and method for early warning of Fabry disease risk. Through expert-based knowledge graph modeling and in-hospital CDSS foundation, a Fabry disease risk prediction system was developed and validated. It is embedded in clinical workflow to provide real-time diagnosis and treatment reminders, improve the efficiency and accuracy of early screening for Fabry disease, and assist doctors in providing timely diagnostic advice to patients. Attached Figure Description
[0024] The above and other objects, features, and advantages of this disclosure will become readily apparent from the following detailed description of exemplary embodiments, taken in conjunction with the accompanying drawings. Several embodiments of this disclosure are illustrated in the drawings by way of example and not limitation, in which:
[0025] Figure 1 A schematic diagram of a system architecture for early warning of Fabry disease risk based on clinical decision support, according to one embodiment of the present disclosure. Detailed Implementation
[0026] The purpose of this disclosure is to develop a model for predicting the risk of Fabry disease, addressing the aforementioned clinical needs and technical challenges, and to combine it with a clinical decision support system that provides efficient screening, real-time alerts, and is easy to standardize and promote. By integrating clinical data and advanced algorithm models, this model aims to achieve efficient and accurate screening and timely early warning for Fabry disease.
[0027] This disclosure is based on the creation of a model using a medical knowledge base. Through comprehensive literature retrieval and screening, the variables and their weights in the model are determined. Data simulation is conducted using confirmed cases and the model is continuously optimized to ultimately form a risk stratification with variables, scores, and quantification.
[0028] This disclosure is based on a medical knowledge base and a predictive model built using data from previous Fabry disease cases. The following relevant parameters were selected, and reference sources are listed below:
[0029] 01. Left ventricular hypertrophy (excluding hypertensive heart disease);
[0030] 02, symmetrical hypertrophy or centripetal hypertrophy;
[0031] 03, Bicompartmental hypertrophy;
[0032] Source: Italian screening study Citro R, et al. Front Cardiovasc Med. 2022 Apr25;9:838200. (Unexplained LVH with at least one red flag sign: neuropathic pain, hypohidrosis / anhidrosis, angiokeratoma, gastrointestinal problems (nausea, vomiting, abdominal pain, diarrhea or constipation), chronic kidney disease or cerebrovascular complications (transient ischemic attack or stroke)), positive rate 10%;
[0033] Esposito R, et al. J. Clin. Med., 2021, 10(9): 1994. Cardiac involvement in Fabry disease can also manifest as right ventricular hypertrophy (RVH), and it is estimated that RVH will occur in 40%-70% of patients.
[0034] 04. Papillary muscle hypertrophy;
[0035] Source: Niemann M, et al. Ultrasound Med Biol. 2011 Jan;37(1):37-43. The sensitivity of papillary muscle area >3.6cm2 and the ratio of papillary muscle area to left ventricular cross-sectional area >0.18 was 75% and the specificity was 86%.
[0036] 05, bilateral features / heterogeneous echoes / ground-glass opacities, etc.;
[0037] Source: Pieroni M, et al. J Am Coll Cardiol. 2006 Apr 18;47(8):1663-71. The sensitivity was 94% and the specificity was 100%.
[0038] 06, longitudinal strain in the outer / rear region is reduced;
[0039] 07, longitudinal strain is reduced;
[0040] Source: Militaru S, et al. Echocardiography. 2019 Nov;36(11):2041-2049. The sensitivity of subbasal lateral LS ≤ -13% was 93%, and the specificity was 50%.
[0041] 08, shortened PR interval / ventricular pre-excitation / pacing rhythm;
[0042] Source: Vitale G, Ditaranto R, Graziani F, et al. Heart. 2022; 108(1):54-60. Shortened PR interval can independently predict the diagnosis of Fabry disease, with a sensitivity of 69% and a specificity of 84% when combined with other parameters (prolonged QRS duration, right bundle branch block (RBBB), left upper limb compression lead (aVL) ≥1.1 mV and inferior ST segment depression).
[0043] 09, Delayed enhancement of gadolinium in the inferior lateral wall of the left ventricle;
[0044] 10. Gadolinium delayed enhancement;
[0045] Source: Perry R, et al. JACC Cardiovasc Imaging . 2019 Jul;12(7 Pt 1):1230-1242. Esposito R, et al. J Clin Med. 2021 May 6;10(9):1994. LGE in Fabry disease is commonly found in the basal segment of the inferior lateral wall of the left ventricle and may be considered a red flag sign for the diagnosis of Fabry disease.
[0046] 11. The native T1 value decreased;
[0047] Source: Niemann M, et al. Ultrasound Med Biol. 2011 Jan;37(1):37-43. The sensitivity of papillary muscle area >3.6cm2 and the ratio of papillary muscle area to left ventricular cross-sectional area >0.18 was 75% and the specificity was 86%.
[0048] 12. History of dialysis;
[0049] Source: Russian screening study Moiseev, S.; et al. Nephron 2019, 141, 249–255. (The positive rate was 0.36% in hemodialysis patients; the study also analyzed the age of patients at the time of dialysis initiation and found that 60% of patients started dialysis before the age of 40).
[0050] 13. Renal insufficiency;
[0051] Source: Lin CJ, et al. Kidney Blood Press Res. 2018;43(5):1636-1645. (The positive rate was 0.59% in screening male patients with CKD of unknown cause).
[0052] 14. Corneal vortex opacity;
[0053] 15. Corneal opacity;
[0054] Source: Japanese screening study Yoshida S, et al. Orphanet J Rare Dis . 2020 Aug 26;15(1):220. (The positive rate was 4.37% in limb pain, limb paresthesia, angiokeratoma, corneal whorled opacity, and hypohidrosis.)
[0055] 16. Clear, burning, neuropathic pain;
[0056] 17. Suspected neuropathic pain: pain in the extremities / limbs;
[0057] Source: Japanese screening study Yoshida S, et al. Orphanet J Rare Dis . 2020 Aug 26;15(1):220. (The positive rate was 4.37% in limb pain, limb paresthesia, angiokeratoma, corneal whorled opacity, and hypohidrosis.)
[0058] 18. A confirmed diagnosis of sensorineural hearing loss;
[0059] 19. Suspected sensorineural hearing loss: tinnitus, hearing loss, etc.;
[0060] Source: Turkish screening study Köroğlu EY, et al. Turk Arch Otorhinolaryngol. 2023 Jun;61(2):52-57. (The positive rate was 1.2% in patients with sensorineural hearing loss).
[0061] 20. Sweating disorders;
[0062] 21. Keratinoma / angiokeratoma;
[0063] 22. Suspected skin rash;
[0064] Source: Japanese screening study Yoshida S, et al. Orphanet J Rare Dis . 2020 Aug 26;15(1):220. (The positive rate was 4.37% in limb pain, limb paresthesia, angiokeratoma, corneal whorled opacity, and hypohidrosis.)
[0065] 23. Cerebral infarction / stroke, etc.;
[0066] Source: Italian screening study Citro R, et al. Front Cardiovasc Med. 2022 Apr25;9:838200. (Unexplained LVH with at least one red flag sign: neuropathic pain, hypohidrosis / anhidrosis, angiokeratoma, gastrointestinal problems (nausea, vomiting, abdominal pain, diarrhea or constipation), chronic kidney disease or cerebrovascular complications (transient ischemic attack or stroke)), positive rate 10%.
[0067] 24. Family history of heart, brain, or kidney disease and premature death;
[0068] Source: Japanese screening study Yoshida S, et al. Orphanet J Rare Dis . 2020 Aug 26;15(1):220. (The positive rate was 23.4% in people with a family history of Fabry disease).
[0069] 25. Myeloid bodies / zebra bodies are visible under electron microscopy;
[0070] Source: Chinese Expert Consensus on the Diagnosis and Treatment of Fabry Disease (2021 Edition) (Histopathological biopsy has auxiliary diagnostic significance, and "myeloid bodies" under electron microscopy are a typical pathological feature).
[0071] This embodiment employs a clinical decision support system. Based on IDCNN (Dilated Convolutional Network) technology, named entity recognition of patient medical data is performed using CNN (Convolutional Neural Network), and TextCNN is used to extract entity relationships from documents, forming a standardized Fabry screening dataset for the hospital. After configuring the above parameters in the clinical decision support system, a model analysis was performed using logistic regression based on historical patient data, yielding the following prediction formula. * :
[0072]
[0073] f(x) is the patient's final risk score, i is the parameter number (0-24), and ax is the parameter weight.
[0074] According to one or more embodiments, this disclosure proposes a Fabry disease risk prediction system based on artificial intelligence and big data analysis. Through machine learning algorithms and integrated clinical information modeling, it achieves intelligent assessment of the disease risk in a target population. The method adopts a modular development approach in its architecture design, integrating multi-source heterogeneous data processing technology, natural language processing technology using deep learning models, and personalized risk prediction result generation, possessing high accuracy, practicality, and scalability.
[0075] Firstly, regarding disease risk prediction, this disclosure introduces a hybrid model combining convolutional neural networks (CNN) and random forests for natural language processing. This model, trained on a large amount of multi-dimensional data including patient phenotypic features and family history, can automatically identify characteristics closely related to Fabry disease, such as myocardial hypertrophy and electrocardiogram findings of Wolff-Parkinson-White syndrome, for comprehensive risk prediction. Furthermore, with the refinement of guidelines and expert opinions, the extracted features and model parameters are adjusted to form the final method for predicting Fabry disease risk.
[0076] This disclosure combines the characteristics of rare disease incidence with clinical workload requirements to complete model structure optimization, training, and validation. During model structure optimization, considering the completeness of confirmed patient data and their relatively high scores, the weighting of core clinical features is optimized to reduce interference from non-specific indicators. A dynamic correction mechanism for repeatedly hospitalized confirmed cases is introduced to adapt to the score stability requirements of long-term follow-up populations. Based on population incidence and clinical screening efficiency, the score range is calibrated to achieve clinical fit. Model training and validation employ a combination of retrospective testing and prospective validation: a retrospective test set is used with previously diagnosed Fabry disease patients, whose model scores are all in the 10-30 range, completing model weight fitting and scoring range calibration for confirmed cases; a prospective test set is run on all hospitalized patients for one consecutive month, including confirmed patients, whose model scores are in the 16-32 range. Based on the above training and validation results, and taking into account both the incidence rate in the population and the clinical workload, the clinical risk stratification thresholds were determined: 8-9 points indicate intermediate-risk patients, and a review is recommended; ≥10 points indicate high-risk patients, and the Fabry disease diagnosis process is initiated. Finally, the model structure was optimized and its performance was solidified, achieving Fabry disease screening that balances screening sensitivity and clinical workload.
[0077] Secondly, this disclosure combines this Fabry disease risk prediction method with a clinical decision support system. The backend of the clinical decision support system is built on a microservice architecture, possessing excellent concurrency processing capabilities and system scalability, adaptable to the needs of large-scale population screening. This allows this disclosure to perform real-time dynamic data monitoring and risk alerts during doctors' daily clinical practice. Simultaneously, all user data is encrypted and access controlled to ensure privacy and security.
[0078] The variables in the model built by Fabry disease experts are further logically judged and accurately collected by data processing experts, and extracted from the patient's electronic medical record information using AI technology. AI automatically collects, calculates, and predicts the patient's risk of developing Fabry disease, and provides real-time alerts at the doctor's workstation, improving the efficiency of Fabry disease screening and increasing the diagnosis rate.
[0079] The Fabry disease risk prediction system disclosed in this embodiment specifically includes the following functional modules and technical parameters:
[0080] The module for collecting and processing data related to high-risk signs of Fabry disease collects and processes information related to high-risk signs of Fabry disease in patients, including:
[0081] (1) Patient's basic information;
[0082] (2) Present medical history: electrocardiogram, echocardiogram, cardiac magnetic resonance imaging, and speckle tracking imaging results;
[0083] (3) Past medical history: manifestations of kidney involvement, ocular involvement, peripheral nerve involvement, skin involvement, and brain involvement;
[0084] (4) Family history;
[0085] (5) Biopsy results.
[0086] The technical parameters of this module include:
[0087] (1) Implement data governance services required for auxiliary diagnosis within the hospital's closed data environment, including data exploration, data collection, data cleaning, dictionary standardization, primary data structuring, terminology unification, secondary data structuring, and data quality control.
[0088] (2) For real-time data processing during diagnosis and treatment, the number of data processed shall not be less than 500 / s and the response speed shall not be higher than 3s.
[0089] Fabry disease risk prediction model:
[0090] (1) Fabry disease risk prediction model based on previous guidelines: about 25-30 high-risk signs (including cardiac lesion features, kidney involvement features, eye involvement features, peripheral nerve involvement features, skin involvement features, brain involvement features, family history and biopsy results).
[0091] (2) The basic model for predicting the risk of Fabry disease is deployed and connected to the hospital's clinical auxiliary decision-making system, and the model is generalized and the rules are optimized based on local data;
[0092] (3) Based on governance data, calculate the risk of Fabry disease and trigger logic processing, and remind doctors of the risk of patients in real time at the doctor's workstation to assist doctors in Fabry disease screening.
[0093] The technical parameters of this module include,
[0094] (1) Ensure data stability and reliability, with real-time data processing count not less than 500 / s and response speed not exceeding 3s.
[0095] Fabry Risk Screening and Statistics Platform:
[0096] (1) Visual display of screening data, supporting the viewing of specific process data, such as: total number of people screened, total number of people at medium risk, total number of people at high risk, number of Fabry disease patients, and the results of specific screening patient assessment forms.
[0097] The technical parameters of this module include:
[0098] (1) Provides visual display and supports continuous updating of statistical data.
[0099] The clinical validation process disclosed herein included establishing a Fabry disease screening cohort study based on CDSS, obtaining scientific review and ethical approval from a hospital, and screening a total of 584,081 inpatients in the cardiology department. The system identified 1,287 patients as medium- to high-risk, with 98 classified as high-risk. Of these, 44 were sent for Fabry enzyme / substrate / genetic testing, ultimately confirming 9 cases. Simultaneously, family screening of the confirmed patients further confirmed 6 cases. A total of 582,794 patients were assessed as low-risk and not identified by the system. Of these, 101 patients were empirically tested by clinicians for Fabry enzyme / substrate / genetic testing, ultimately resulting in 0 confirmed cases.
[0100] Comparison of Fabry disease screening and diagnosis data in 2024
[0101]
[0102] In 2025, model optimization and system upgrades were carried out, resulting in further improvements in the submission rate and positive rate compared to the same period in 2024.
[0103] Comparison of 2025 and the same period in 2024
[0104] Verification of the embodiments disclosed herein demonstrates that the batch screening of high-risk Fabry disease patients using the model disclosed herein is far more efficient than manual screening, significantly saving labor costs, improving diagnostic efficiency, and reducing missed diagnoses. Furthermore, during patient treatment, any changes in information trigger real-time calculation of the latest risk score and simultaneous risk alerts, helping doctors adjust treatment strategies immediately. Because it is based on big data and combined with an expert-designed Fabry disease prediction model, it breaks down the barriers preventing specialists from accessing cross-departmental medical records, achieving more accurate identification of Fabry disease risk.
[0105] According to one or more embodiments, such as Figure 1 As shown, a system architecture and process for early warning of Fabry disease risk based on clinical decision support is presented. It comprises offline and online components. The offline component is primarily responsible for constructing a Fabry disease profile and establishing a Fabry disease prediction model, while the online component is mainly responsible for real-time Fabry disease risk warnings. The boundary between the two components is defined by the patient's overall profile, forming a collaborative mechanism of offline pre-computation and online incremental updates. Dashed lines in the diagram represent cross-domain data or model transfer, while solid lines represent processing flows between modules.
[0106] The offline component handles historical data mining and expert model building, processing the complete historical data of confirmed Fabry disease positive and negative patients. Data annotation is divided into Fabry disease positive and Fabry disease negative patients. Each patient's data is presented along a timeline based on their medical visits, including multiple visit records from visit 1 to visit n. Each visit's data covers six dimensions: diagnosis, examination, laboratory tests, surgery, medication, and medical records. The raw, multi-source, heterogeneous data first undergoes NLP model processing, employing a CNN combined with a TextCNN architecture: first, a convolutional neural network performs medical named entity recognition (e.g., extracting "left ventricular posterior wall" / location, "thickness" / attribute, "14" / value, and "mm" / unit from "left ventricular posterior wall thickness: 14mm"); then, TextCNN extracts entity relationships, combining discrete entities into structured medical events to transform unstructured text into a computable data structure. The NLP-processed data is then fed into the patient profile module, forming a standardized data structure containing comprehensive information including vital signs, symptoms, disease, examinations, laboratory tests, medical orders, and medical records. Fabry disease-related features are further extracted from the patient's comprehensive profile to construct a Fabry disease-specific profile (dataset). This dataset corresponds to a configuration set of 25-30 high-risk vital signs parameters. The Fabry disease-specific factor (rule engine) module performs feature engineering processing on the disease-specific dataset, including time-series comparisons across medical visit data and array operations, outputting structured predictive factors. These predictive factors undergo an expert factor scoring process, where clinical experts assign weight coefficients based on a medical knowledge base, ultimately forming an expert model. This model is expressed using a logistic regression formula, outputting a Fabry disease risk score and defining stratification thresholds for low, intermediate, and high risk.
[0107] The online component provides real-time risk warnings and clinical decision support. Its core design lies in incremental calculation mechanisms, avoiding repetitive full processing of historical data. The online process is triggered by a change in the latest patient information. This event can manifest as any data change operation, such as a doctor saving admission records, adding test reports, or updating diagnostic information. Upon triggering, the system performs incremental calculations only on the latest patient profile, using the same NLP processing flow as the offline component to extract medical entities from the new data. Subsequently, this incremental profile is merged with historical patient profiles (pre-stored data such as Patient 1 profile, Patient 2 profile, etc.) to form a complete patient profile. The merged profile then undergoes Fabry disease profile (dataset) filtering and Fabry disease factor (rule engine) feature extraction before being input into an expert model for risk score calculation, outputting a Fabry risk score and corresponding hazard stratification.
[0108] Specifically, the offline architecture includes:
[0109] (1) Data source: Rare diseases are often characterized by recurrent attacks, multiple visits, and visits to multiple departments, so it is necessary to integrate information from multiple visits and multiple departments. At the same time, since it is necessary to differentiate from other diseases, it is necessary to consider information from all aspects, not only the formatted diagnosis, medication, and surgical data, but also textual information such as documents and pathology reports. Therefore, in terms of data source, this disclosure has the characteristics of cross-visit, cross-department, and rich data dimensions.
[0110] (2) In the process of constructing a panoramic patient profile, a large amount of document data needs to be processed. In this process, the NLP entity recognition algorithm is first used to extract medical entities, such as “left ventricular posterior wall thickness: 14mm”. The “left ventricular posterior wall” (location), “thickness” (attribute), “14” (value), and “mm” (unit) are extracted first. Then, the entities are combined using the NLP relation discovery algorithm. Here, the combination of “left ventricular posterior wall”, “thickness”, “14”, and “mm” is obtained. Finally, the medical knowledge is configured and mapped to the medical concept of “left ventricular wall thickness” with the attribute of 14mm.
[0111] (3) By combining the patient profiles from multiple visits, a comprehensive patient profile can be obtained. The patient profile mainly includes the following modules: vital signs, symptoms, disease, examinations, tests, detailed examination items, detailed test items, surgery or procedure, allergies, assessment form, detailed assessment items, medication orders, examination orders, test orders, surgical orders, nursing orders, and documents.
[0112] (4) Based on the overall patient profile, a Fabry disease dataset is further constructed, and the content related to Fabry in the profile is screened and reorganized.
[0113] (5) Based on the Fabry disease dataset and the global patient profile, further extract Fabry-related disease factors. As shown below, in the processing of the longitudinal strain reduction disease factor, cross-visit and array comparison operations will be performed to produce predictive factors related to Fabry disease.
[0114] (6) After obtaining the predictive factors, the risk factors are evaluated by the expert team, and corresponding risk scores are assigned to the risk factors. The thresholds for different risk stratifications are determined according to the scores, and the risk levels are divided into three levels: low risk, medium risk, and high risk.
[0115] The online architecture includes:
[0116] (1) The online process is largely the same as the offline process, with the main modifications made to meet the need for real-time early warning across medical visits. First, the patient's historical medical profile is pre-calculated and stored on the hard drive, and will not be triggered by changes in the latest medical information.
[0117] (2) When the latest medical information changes, such as when a new document is added in the latest hospitalization, the document will be calculated separately and merged into the patient profile of the latest visit. Then the patient profile of the latest visit will be merged with the patient profile of the historical visits to obtain the same panoramic patient profile as in the offline process.
[0118] (3) After obtaining the patient's panoramic profile, the Fabry disease dataset is calculated in the subsequent process. The Fabry disease predictor factors are consistent with those in the offline process.
[0119] (4) After obtaining the predictive factors, the risk score of the patient's Fabry disease is obtained according to the expert model, and the patient is automatically identified as high-risk, intermediate-risk, or low-risk. The results are then sent to the doctor's client in real time.
[0120] In predicting the risk of Fabry disease in patients, this disclosure considers not only the patient's current medical records but also the patient's past medical records (including information from other departments); and uses the patient's full in-hospital data, including medical documents, pathology reports, examination / laboratory reports, etc.
[0121] Therefore, the method disclosed herein has the following advantages compared with other Fabry machine learning algorithms:
[0122] (1) It is more conducive to the promotion of standardization.
[0123] This disclosure, through professional medical knowledge and combined with existing Fabry disease cohort research, uses commonly used clinical indicators to form a standardized formula for predicting the risk of Fabry disease in patients. It helps to solve the problem of inconsistent understanding of rare diseases among clinicians, and is especially suitable for large-scale promotion in primary hospitals. It helps primary hospitals to make more effective use of limited resources, improve the overall quality of medical care, and contribute to the rational allocation of medical resources.
[0124] (2) It is cost-effective to integrate with clinical diagnosis and treatment auxiliary decision-making systems.
[0125] The integration and deployment of the model take into account cost-effectiveness and ease of operation. In resource-constrained healthcare institutions, the deployment and maintenance costs of the model can be a significant factor. Therefore, this disclosure is easy to integrate, cost-effective, and readily applicable in a wide range of healthcare settings.
[0126] The section involving clinical procedures includes the following content.
[0127] The Fabry High-Risk Screening Assessment Form is used for automatic assessment and alerts.
[0128] (1) Scenario: When a patient suspected of Fabry disease is admitted to the hospital, the doctor fills in the [Admission Record] and [Initial Medical Record]. After clicking to save the document, the system will make a judgment based on the doctor's recorded diagnosis, the patient's medical record, and the in-hospital test results, and automatically calculate the Fabry disease high-risk screening score. If the score is ≥8 points, the doctor will be reminded to view, modify and confirm.
[0129] (2) The operation process includes:
[0130] Click on "Treatment" in the medical record system to view reminders and triage results;
[0131] In the large window, you can view the assessment results automatically calculated by the system's AI based on medical records, diagnoses, test results, and other data; click [Confirm] to accept the system's assessment results;
[0132] Click on the "Assessment Form Name" [Fabry High-Risk Screening Assessment] or [System Assessment Results] to enter the "Assessment Details and Basis Page";
[0133] Display the results of AI-generated calculations, as well as the scores and basis for the matched variables;
[0134] Clicking on "Assessment Basis" (e.g., for left ventricular hypertrophy, see "Ward Rounds - Ward Rounds Records") will take you to the "Assessment Basis Source Page";
[0135] The original text can be viewed on the "Assessment Basis Traceability Page," where the basis is highlighted in red.
[0136] In the "Assessment Form Editing Page", doctors can modify the test values and options to obtain the final assessment results;
[0137] (3) Treatment recommendations include:
[0138] Scenario 1: When a patient suspected of having Fabry disease is admitted to the hospital, after the doctor fills out the [Admission Record] and [Initial Medical Record] and clicks to save the document, the system will make a judgment based on the doctor's recorded diagnosis, the patient's medical records, and in-hospital laboratory test results, and automatically calculate the Fabry disease high-risk screening score. When the score is ≥8 points but <10 points, in addition to the above assessment form, the system will simultaneously remind the doctor to perform echocardiography and cardiac MRI examinations. If the patient has already undergone echocardiography and cardiac MRI examinations, no further reminders will be given.
[0139] Scenario 2: When a patient suspected of having Fabry disease is admitted to the hospital, the doctor fills out the "Admission Record" and "Initial Medical Record". After clicking "Save Document", the system will make a judgment based on the doctor's recorded diagnosis, the patient's medical records, and the hospital's laboratory test results, and automatically calculate the Fabry disease high-risk screening score. When the score is ≥10, in addition to the above assessment form, the doctor will be reminded to conduct Fabry disease enzyme and gene tests.
[0140] Operation process: Click on the [Diagnosis and Treatment] section of the medical record system to view the specific recommended examinations; click to complete and accept the system's recommended results.
[0141] However, in practice, this disclosure has found that due to the extreme scarcity of Fabry disease positive samples (<0.001%), and the existence of ambiguities in medical numerical expressions in medical record texts (such as "left ventricular posterior wall thickness 14mm" versus "14 millimeters") and interference from negative expressions (such as "no corneal opacity" being incorrectly augmented to "with corneal opacity"), conventional NLP data augmentation can generate medically unreasonable spurious samples, leading to overfitting of the predictive screening model. To address this issue, this disclosure proposes solutions for model convergence and clinical usability of screening models in small sample scenarios in clinical settings.
[0142] According to one or more embodiments, a method for predicting the risk of Fabry disease includes the following steps:
[0143] S101, based on historical confirmed patient data, uses a Bayesian network to calculate the posterior probability of various medical sign rule combinations and constructs a dynamic weighted rule engine;
[0144] S102, rule-based screening is performed on unlabeled medical records in the past. When the posterior probability of the rule combination exceeds the first threshold, a suspected positive pseudo-label is generated. When there is a rule conflict, the conflict is resolved based on the shortest path algorithm of the medical knowledge graph. Feature combinations that are semantically closer to the pathological mechanism of Fabry disease are retained first to construct the initial training set.
[0145] S103, perform medical-specific text data augmentation on a very small number of real positive samples. The augmentation includes: perturbing numerical vital signs within the error range allowed by medical guidelines and simultaneously adjusting related derived descriptions; locking negative expressions and their modified objects through a negative word detector and prohibiting semantic transformation of the locked segments; and reorganizing sentences while keeping medical entities and negative relationships unchanged to generate medically reasonable augmented samples.
[0146] S104, a language model pre-trained based on a general medical corpus, is fine-tuned hierarchically using the enhanced samples and the initial training set, freezing the general parameters at the bottom layer, fine-tuning only the recognition weights of specific entities in the top layer for Fabry disease, and introducing a class reweighting loss function to solve class imbalance.
[0147] S105 deploys the finely tuned model into the clinical decision support system to perform risk prediction on real-time collected patient medical data.
[0148] S106, Based on the uncertainty sampling strategy, cases with predicted probabilities in the fuzzy range are pushed to doctors for review;
[0149] S107 feeds back the doctor's review results to the training set, triggering an incremental learning mechanism that fine-tunes only the parameters of the top-level classifier to complete the model's iterative optimization.
[0150] Furthermore, the calculation of the posterior probability of each combination of medical sign rules based on a Bayesian network includes:
[0151] A conditional probability graph containing 25-30 high-risk vital signs nodes is constructed, and the conditional probability is calculated through parameter learning based on historical confirmed patient data. ,in For different vital signs nodes. This refers to the random event that a patient is diagnosed with Fabry disease;
[0152] The first threshold for posterior probability is set to 0.7-0.8. When the posterior probability of the rule combination exceeds this threshold, a suspected positive false label is generated. Cases below this threshold but above the second threshold (0.3-0.4) are marked as pending observation, and those below the second threshold are marked as negative.
[0153] The shortest path algorithm based on medical knowledge graphs resolves conflicts, including:
[0154] When a combination of rules pointing to Fabry disease conflicts with rules pointing to other diseases, the semantic distance from the conflicting node to the core pathological node is calculated in the medical knowledge graph.
[0155] Feature combinations with shorter semantic distances (<3 hops) are prioritized for retention, while conflicting features with longer semantic distances (≥3 hops) are discarded to generate the final pseudo-labels.
[0156] The interval perturbation of logarithmic vital signs within the error range allowed by medical guidelines includes:
[0157] Identify numerical medical entities and their units in medical record texts, and generate reasonable numerical variations within a set range based on the measurement error range determined by medical guidelines.
[0158] Based on the medical classification range in which the numerical variants are located, the corresponding qualitative description text is adjusted synchronously to ensure consistency between the numerical values and the descriptions.
[0159] The method of locking negative expressions using a negative word detector includes:
[0160] Construct a negative term library (including "none", "not seen", "deny", "not heard of", etc.) and identify negative expressions in medical record texts through regular expression matching;
[0161] Medical entities modified by negative words are extracted as locked fragments. During the data augmentation process, synonym replacement, sentence restructuring, or numerical perturbation of the locked fragments are prohibited to prevent negative medical records from being mistakenly augmented into positive samples.
[0162] The uncertainty-based sampling strategy includes:
[0163] Set a fuzzy range as the predicted probability, and prioritize pushing high-uncertainty cases within this range for review, rather than only pushing high-scoring cases;
[0164] The system monitors the number of medical records that doctors are currently processing in real time. When a doctor's current writing workload exceeds the preset limit or is in the surgical period, the system will temporarily postpone the process or assign it to a doctor with a lower workload.
[0165] The triggering mechanism for incremental learning includes:
[0166] Once 100 doctor review results have been accumulated, incremental training of the model will be triggered.
[0167] Freeze the parameters of the feature extraction layer, fine-tune only the weights of the top classifier, and keep the time for a single incremental training session within 5 minutes to avoid affecting the real-time service of the system.
[0168] The following example is given to further illustrate the technical solution of the embodiments of this disclosure.
[0169] A method for predicting the risk of Fabry disease includes the following steps:
[0170] S201, Dynamic allocation of rule weights based on Bayesian posterior probability. Assuming only 9 initial real positive samples, the screening model cannot be directly trained. Therefore, a conditional probability graph for Fabry disease is constructed here, where nodes include 25 signs such as "left ventricular hypertrophy (LVH)," "corneal whorled opacity," "acromial pain," and "renal insufficiency," and edges represent the strength of pathological associations. Based on data from 9 historical confirmed patients, the posterior probability of each rule combination is calculated through Bayesian network parameter learning:
[0171] ,in Different physical signs;
[0172] A dynamic threshold is set. When the posterior probability of the rule combination is >0.75, a pseudo-label of "suspected positive" is generated. When there is a rule conflict (such as "LVH + corneal opacity" pointing to Fabry disease, but "history of hypertensive heart disease" pointing to other diseases), the shortest path algorithm based on knowledge graph is used to resolve the conflict: the semantic distance of the conflicting nodes in the medical knowledge graph is calculated, and feature combinations that are closer to the pathological mechanism of Fabry disease (α-galactosidase deficiency leading to glycosphingolipid accumulation) are retained first.
[0173] S202, augmentation of text data specific to the medical field. For 9 positive medical records, targeted augmentation under medical rationality constraints was implemented, specifically including:
[0174] To constrain the range perturbations of numerical signs, for example, for key values such as "left ventricular posterior wall thickness: 14mm", instead of arbitrary synonym replacement, reasonable numerical variations between "12mm" and "16mm" are generated based on the measurement error range allowed by medical guidelines (±2mm), and related derived descriptions are adjusted simultaneously (e.g., "mild thickening" becomes "moderate thickening" as the value is adjusted) to avoid generating medically impossible pseudo-samples; a locking mechanism for negative expressions is set up, using a negative word detector (based on regular expressions + a negative word library, such as "none", "not seen", "deny") to identify negative expressions such as "no corneal vortex opacity" and "deny of limb pain", freezing negative words and their modified objects ("corneal vortex opacity", "limb pain") during the enhancement process, prohibiting synonym replacement or sentence restructuring of these segments, and preventing the erroneous enhancement of negative medical records into positive samples; sentence restructuring is performed through sentence templates, adjusting the sentence structure (active to passive, sentence merging) while keeping the medical entities and negative relationships unchanged, to generate diverse but medically semantically consistent text. Through the aforementioned constraint enhancements, the 9 original positive cases were expanded into 180 medically reasonable enhanced samples. Furthermore, after sampling verification by medical experts, the medical accuracy of the enhanced samples reached 98.2%.
[0175] S203, based on BioBERT, performs domain-specific fine-tuning (freeze-thaw strategy). Through a hierarchical fine-tuning strategy, the underlying general medical semantic parameters of BioBERT (layers 1-10) are frozen, while only the recognition weights of Fabry disease-specific entities (such as "bilateral sign" and "myeloid bodies") at the top level (layers 11-12) are fine-tuned to prevent catastrophic forgetting in small sample sizes; class reweighting is performed, introducing class weights into the loss function. (The square root of the ratio of negative samples to positive samples) solves the gradient vanishing problem caused by extreme class imbalance (1:50000). N neg N is the total number of negative samples (i.e., samples that do not have Fabry disease) in the training dataset. pos This represents the total number of positive samples (i.e., samples diagnosed with Fabry disease) in the training dataset.
[0176] S204 employs an uncertainty sampling active learning approach based on balancing the workload of physician reviewers. It prioritizes fuzzy cases (the least uncertain samples in the model) with predicted probabilities between 0.4 and 0.6, rather than only pushing high-scoring cases, maximizing the information gain from manual annotation. It monitors the number of medical records pending processing by physicians in real time, and postpones pushing new review tasks when a physician's current writing workload exceeds 20 records per hour or is in surgery. It uses a load balancing algorithm to distribute review tasks to physician terminals with lower daily workloads. After accumulating review results for hundreds of physicians, it only fine-tunes the parameters of the top-level classifier and freezes the feature extraction layer to avoid the loss of learned features due to full retraining. A single incremental training session takes less than 5 minutes, ensuring that it does not affect the system's real-time service.
[0177] Therefore, in this embodiment of the disclosure, a targeted small-sample solution is proposed to address the dilemma of extremely scarce Fabry disease positive samples (less than 0.001%) and the ease with which conventional data can generate medically unreasonable pseudo-samples. In summary,
[0178] First, a dynamic weighted rule engine is built based on Bayesian network. The posterior probability of each high-risk sign rule combination is calculated using historical confirmed patient data. Suspected positive pseudo-labels are generated by screening a large number of unlabeled medical records. When rule conflicts occur, the shortest path algorithm of medical knowledge graph is used to resolve the conflicts. Feature combinations that are semantically closer to the core pathological mechanism of Fabry disease are retained first, thus constructing an initial training set without a large amount of labeled data.
[0179] Secondly, text data augmentation in the medical field is implemented. Numerical vital signs are perturbed only within the measurement error range allowed by medical guidelines, and related qualitative descriptions are adjusted simultaneously to ensure medical consistency. Negative word detectors are used to lock negative expressions such as "none," "not seen," and "deny" and their modified objects, and any semantic transformation of the locked segments is prohibited to prevent negative medical records from being mistakenly augmented into positive samples. Sentence reorganization is carried out while keeping medical entities and negative relationships unchanged, thereby expanding a medically reasonable augmented sample set based on a very small number of real positive samples.
[0180] Furthermore, a hierarchical fine-tuning strategy based on a pre-trained language model using a general medical corpus is adopted. The underlying general medical semantic parameters are frozen, and only the recognition weights of specific entities for Fabry disease at the top level are fine-tuned. A class reweighting loss function is introduced to solve the problem of extreme class imbalance and avoid catastrophic forgetting and gradient vanishing in small sample scenarios.
[0181] Finally, an uncertainty sampling strategy is introduced during the model deployment phase. Cases with predicted probabilities in the fuzzy range are prioritized for doctor review. An incremental learning mechanism is triggered based on the doctor's feedback, and the model iteration is completed by only fine-tuning the parameters of the top classifier. Each training session is controlled within a few minutes, continuously improving model performance without affecting the real-time service of the system.
[0182] It should be understood that in the embodiments of this disclosure, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.
[0183] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0184] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this disclosure, and these modifications or substitutions should all be covered within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.
Claims
1. A method for predicting the risk of Fabry disease, characterized in that, This method predicts the risk of developing Fabry disease using a Fabry disease risk prediction model built based on expert knowledge. The risk prediction model includes multiple high-risk signs and their corresponding weights. These high-risk signs include left ventricular hypertrophy, symmetrical hypertrophy, papillary muscle hypertrophy, bilateral sign, reduced longitudinal strain, shortened PR interval, delayed gadolinium enhancement of the inferior lateral wall of the left ventricle, decreased native T1 value, history of dialysis, renal insufficiency, corneal whorled opacity, neuropathic burning pain, neuropathic hearing loss, sweating disorders, keratoma, cerebrovascular complications, and family history.
2. The method according to claim 1, characterized in that, Including steps, A1. Obtain patient medical data, including structured data and unstructured text data including medical records and pathology reports; A2, the unstructured text is subjected to entity recognition and relation extraction to extract medical entities and their attribute values related to Fabry disease, and mapped to standardized medical concepts based on a medical knowledge base; A3, using the Fabry disease risk prediction model, calculates a risk score based on the standardized medical concepts stated above.
3. The method according to claim 2, characterized in that, Training the Fabry disease risk prediction model includes the following steps: B1, based on historical confirmed patient medical data, uses a Bayesian network to calculate the posterior probability of various medical sign rule combinations and constructs a dynamic weighted rule engine; Unlabeled medical records in the past are filtered by rules. When the posterior probability of a combination of rules exceeds the first threshold, a suspected positive pseudo-label is generated. When there is a rule conflict, the conflict is resolved by the shortest path algorithm based on the medical knowledge graph. Feature combinations that are semantically closer to the pathological mechanism of Fabry disease are retained first to construct the initial training set. B2, performing medical-specific text data augmentation on a very small number of real positive samples, including: perturbing numerical signs within the error range allowed by medical guidelines and simultaneously adjusting related derived descriptions; By using a negation word detector to identify negative expressions and their modified objects, semantic transformation of the identified segments is prohibited. While maintaining the medical entities and negation relationships, sentence structure is reorganized to generate medically reasonable augmented samples; B3 is a language model pre-trained on a general medical corpus. It uses augmented samples and the initial training set for hierarchical fine-tuning, freezes the general parameters at the bottom layer, fine-tunes only the recognition weights of Fabry disease-specific entities at the top layer, and introduces a class reweighting loss function to solve class imbalance.
4. The method according to claim 3, characterized in that, The training of the Fabry disease risk prediction model further includes the following steps: B4. Deploy the finely tuned model clinically and perform risk prediction on real-time collected patient medical data. Based on the uncertainty sampling strategy, cases with predicted probabilities in the fuzzy range are pushed to doctors for review; The doctor's review results are fed back to the training set, triggering an incremental learning mechanism to fine-tune the parameters of the top-level classifier and complete the model iterative optimization.
5. The method according to claim 2, characterized in that, By using natural language processing and employing convolutional neural networks combined with the TextCNN model, we can complete named entity recognition and entity relationship extraction of patient medical data to form a Fabry disease screening dataset.
6. The method according to claim 2, characterized in that, Before calculating the risk score, a comprehensive patient profile is constructed. This profile integrates data from all of the patient's previous medical visits, including vital signs, symptoms, disease diagnosis, examination and test results, surgical procedure records, allergy history, assessment forms, medical orders, and medical records.
7. The method according to claim 2, characterized in that, The risk score calculation formula for the risk prediction model is as follows: Where f(x) is the patient's final risk score, i is the parameter number (0-24), and a i x i It is the parameter weight, x i The value of the i-th high-risk vital sign parameter, a i The weighting coefficient corresponding to the i-th high-risk vital sign parameter.
8. A Fabry disease screening method based on a clinical decision support system, characterized in that, include: When doctors write and save admission records or initial medical records, they collect patients' current medical data and historical medical data in real time. Automated risk assessment is performed using the Fabry disease risk prediction method as described in any one of claims 1 to 7; When the risk score reaches the first preset threshold, the Fabry high-risk screening assessment form and recommendations for echocardiography and cardiac magnetic resonance imaging are displayed in real time on the doctor's workstation interface. When the risk score reaches the second preset threshold, the doctor will be simultaneously reminded to perform Fabry enzyme and genetic testing. It provides a function to trace the basis of the assessment, displaying the original medical records on which the risk score is based, and supports doctors to manually correct the assessment results.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor runs the computer program to implement the Fabry disease risk prediction method as described in any one of claims 1 to 7, or the Fabry disease screening method as described in claim 8.
10. A computer program product, comprising a computer program, characterized in that, The computer program is executed by a processor to implement the Fabry disease risk prediction method as described in any one of claims 1 to 7, or the Fabry disease screening method as described in claim 8.