Machine learning model training method and device for predicting individual abnormal state
By using a machine learning model that integrates multi-source data to screen causal lifestyle features and combining multi-gene risk scores and historical disease characteristics, the problem of insufficient accuracy in predicting complex diseases in existing technologies has been solved, enabling more accurate prediction of individual abnormal states and personalized intervention.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2026-03-27
AI Technical Summary
Existing complex disease risk prediction models mainly rely on a single source of information and fail to fully consider the interaction between abnormal states and individual lifestyle habits, resulting in insufficient prediction accuracy.
A machine learning model integrating multi-source data was used to screen lifestyle characteristics through Mendelian randomization analysis and combined with polygenic risk scores and historical disease characteristics. The model was trained using a Transformer-Hawkes model and a stepwise logistic regression model to predict the probability of an individual's abnormal state.
It improves the accuracy and comprehensiveness of risk prediction for complex diseases, and can provide personalized lifestyle improvement suggestions to reduce individual disease risk.
Smart Images

Figure CN121745336A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of intelligent medical treatment, in particular to a machine learning model training method and device for predicting an abnormal state of an individual, and more particularly to a machine learning model training method and device, a device for predicting a probability of an abnormal state of an individual, a server, and a computer readable storage medium. BACKGROUND
[0002] Risk prediction of complex diseases is crucial for today's clinical medical diagnosis, which is a kind of assistant diagnosis tool for primary care physicians, so as to more accurately identify high-risk patients and provide them with targeted dietary, lifestyle intervention opinions and drug treatment in time, and also to avoid over-treatment of low-risk population by doctors and cause additional waste of medical resources.
[0003] In the past three decades, the number of disease risk prediction models has increased dramatically. One type of model is based on a plurality of specific effective risk factor variables to predict diseases, for example, the commonly used features for predicting cardiovascular diseases include age, gender, total cholesterol, low-density lipoprotein, high-density lipoprotein, systolic blood pressure, diastolic blood pressure, smoking, and diabetes. Common models include logistic regression, random forest, support vector machine, and proportional hazards regression, etc. Another type of model uses genomic variations (single nucleotide polymorphism sites) to predict the risk of individuals suffering from diseases (polygenic risk score). The risk score calculation method of this type of model is usually based on risk variants found in genome-wide association study data, and uses weighted summation to calculate the risk of a specific disease of an individual. Common models include P+T, PRSice2, and LDPred2, etc.
[0004] However, studies have shown that most complex diseases are not only affected by a specific factor in one aspect, and different diseases usually interact with each other, that is, the patient's past medical history will also have a certain impact on the production of future diseases. Therefore, using single information to predict complex diseases has limitations and is not accurate enough.
[0005] Therefore, there is an urgent need in the art to develop a complex disease risk prediction model combining multi-source data. SUMMARY
[0006] The present application is proposed by the inventors based on the following problems and facts:
[0007] Currently, there are some biases in the judgment of the abnormal state of an individual, mainly due to the influence of multiple factors on complex abnormal states, including the interaction between abnormal states, the patient's past medical history, and the patient's unique living habits, etc.
[0008] In order to accurately predict the occurrence and development of various abnormal states, and then provide targeted intervention suggestions and drug treatment methods. The application proposes a risk prediction model and a machine learning explanation model for complex diseases based on multi-source data integration.
[0009] The present application aims to at least partially solve one of the above technical problems.
[0010] To this end, in a first aspect of the present application, a machine learning model training method is proposed, which is used to predict the occurrence probability of an individual abnormal state. According to an embodiment of the present application, the method comprises: obtaining feature information of a training sample, the feature information comprising at least a multi-gene risk score feature, a lifestyle habit feature and a historical disease feature; and inputting the feature information into a machine learning model, using the known abnormal state of the training sample as a label, so as to supervise the training of the machine learning model, wherein the lifestyle habit is determined to have a causal relationship with the abnormal state through Mendelian randomization analysis.
[0011] According to an embodiment of the present application, the method can more accurately and efficiently predict the individual abnormal state by integrating different data sources and using lifestyle habit features having a causal relationship with the abnormal state for modeling. The machine learning model can be trained by the above machine learning model training method, which can accurately and comprehensively predict the occurrence probability of the individual abnormal state.
[0012] It should be noted that the abnormal state includes but is not limited to a disease.
[0013] It should be noted that after the Mendelian randomization analysis, a plurality of Mendelian randomization results are obtained, and through meta-analysis (statistical analysis method) of the Mendelian randomization results, the lifestyle habit feature meeting all of the following conditions (a-e) is identified as a lifestyle habit having a causal relationship with the abnormal state.
[0014] a. The significance P value of any lifestyle habit feature Mendelian randomization result is less than 0.05;
[0015] b. The fixed effect 95% confidence interval of the meta-analysis result does not contain 0;
[0016] c. The fixed effect significance P value of the meta-analysis result is less than 0.05;
[0017] d. The I square of the meta-analysis result is less than 0.05;
[0018] e. The equivalence P value of the meta-analysis result is greater than 0.05.
[0019] According to an embodiment of the present application, the machine learning model training method further comprises at least one of the following additional technical features:
[0020] According to an embodiment of the present application, the historical disease feature is obtained by processing disease information using a preprocessing model, the disease information comprising type information and time information; the preprocessing model is trained by the following steps: obtaining the disease information of the training sample; inputting the disease information into an encoding layer to digitally process the disease information; inputting the result of the digital processing into the preprocessing model to obtain converted disease features; inputting the disease features into a machine learning model, and training the preprocessing model and the machine learning model using known disease information of the training sample.
[0021] According to an embodiment of the present application, the inventors use a Transformer-Hawkes model to fit historical disease information. Based on the event prediction layer in the model, the historical disease is re-predicted, and based on the time prediction layer in the model, the diagnosis time of the historical disease is re-predicted. Finally, the overall model is trained by using a gradient descent training method to minimize the Hawkes loss function, thereby updating the model parameters. The output of the preprocessing model Transformer model is the historical disease feature.
[0022] According to an embodiment of the present application, the preprocessing model is a Transformer model, and the Transformer model comprises a masking multi-head attention layer, a first normalization layer, a position-wise feed-forward network, and a second normalization layer. According to an embodiment of the present application, using a Transformer model can speed up the calculation process.
[0023] According to an embodiment of the present application, the multi-gene risk score feature is obtained by processing single nucleotide polymorphism site scores using an LDPred2 model.
[0024] According to an embodiment of the present application, the abnormal state comprises coronary heart disease, and the lifestyle features comprise at least one of the following:
[0025]
[0026]
[0027] According to an embodiment of the present application, the lifestyle features related to coronary heart disease are features determined to have a causal relationship with the coronary heart disease through Mendelian randomization analysis. The features are obtained by the inventors through a causal inference model screening. Too many lifestyle features not only increase the computational load of the machine learning model, but also produce more noise, affecting the effect of model training. Fewer lifestyle features will result in low accuracy of coronary heart disease analysis results, leading to poor representation.
[0028] According to an embodiment of the present application, the abnormal state is type 2 diabetes, and the lifestyle features include at least one of the following:
[0029]
[0030]
[0031] According to an embodiment of the present application, the lifestyle features related to type 2 diabetes are features determined to have a causal relationship with the type 2 diabetes through Mendelian randomization analysis. The features are obtained by the inventors through a causal inference model screening. Too many lifestyle features will increase the computational load of the machine learning model, and fewer lifestyle features will result in low accuracy of coronary heart disease analysis results, leading to poor representation.
[0032] According to an embodiment of the present application, the machine learning model is a stepwise logistic regression model. According to an embodiment of the present application, the stepwise logistic regression model automatically selects features with the largest contribution from the data preprocessed by the LDPred2 model, the causal inference model, and the Transformer-Hawkes model, thereby reducing the dimension and complexity of the features. This can improve the efficiency and generalization ability of the model and avoid overfitting.
[0033] In a second aspect of the present application, a device for predicting the probability of occurrence of an abnormal state of an individual is provided. According to an embodiment of the present application, the device comprises: a feature information acquisition unit configured to acquire feature information of the individual, the feature information comprising a polygenic risk score feature, a lifestyle feature, and a historical disease feature; and a prediction unit connected to the feature information acquisition unit and configured to predict the probability of occurrence of an abnormal state of the individual based on the feature information using a machine learning model trained using the method of the first aspect of the present application.
[0034] According to an embodiment of the present application, the device for predicting the probability of occurrence of an abnormal state of an individual can process data from multiple data sources, thereby accurately predicting the probability of occurrence of an abnormal state of an individual.
[0035] It should be noted that, from a structural point of view, Figure 1 The feature information acquisition unit S100 of the device is connected to the prediction unit S200.
[0036] According to the embodiments of the present application, the device for predicting the probability of the individual abnormal state can further include at least one of the following additional technical features:
[0037] According to the embodiments of the present application, the device further includes an evaluation unit for evaluating the individual living habits and / or generating individual living habit improvement information. The evaluation device can provide personalized improvement suggestions according to the individual abnormal state characteristics and evaluation results. These suggestions include measures such as changing the living habits, regular checkups, and the like, thereby helping the individual reduce the risk of disease.
[0038] It should be noted that, from a structural point of view, Figure 2 The prediction unit S200 of the device is connected to the evaluation unit S300.
[0039] In a third aspect of the present application, a device for training a machine learning model is provided. According to the embodiments of the present application, the device includes a feature information acquisition unit for acquiring feature information of a training sample, the feature information including at least a multi-gene risk score feature, a living habit feature, and a historical disease feature; and a training unit connected to the feature information acquisition unit, for inputting the feature information into a machine learning model, using the known abnormal state of the training sample as a label, so as to supervise the training of the machine learning model; wherein the living habit is determined to have a causal relationship with the abnormal state through Mendelian randomization analysis.
[0040] According to the embodiments of the present application, the device can train a machine learning model for predicting the probability of an individual abnormal state.
[0041] It should be noted that, from a structural point of view, Figure 3 The feature information acquisition unit S400 of the device is connected to the training unit S500.
[0042] In a fourth aspect of the present application, a server is provided. According to the embodiments of the present application, the server includes a processor and a memory, and the memory stores a computer program, when the computer program is executed by the processor, the machine learning model training method of the first aspect is implemented.
[0043] It should be noted that the machine learning model training method of the first aspect can also be in the form of a program package, which is convenient for transmission and carrying.
[0044] In a fifth aspect, the present application provides a computer readable storage medium containing a computer program, characterized in that when the computer program is executed by one or more processors, the machine learning model training method of the first aspect is implemented. The computer program contained in the computer readable storage medium can more accurately contain complex information in the multi-gene risk score feature, the life habit feature and the historical disease feature, and based on a large amount of training in advance, a more accurate result can be given. Moreover, the computer readable storage medium can store a large amount of data information, avoid human error, and ensure the accuracy and reliability of the data.
[0045] It should be noted that in the present application, the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequence list of executable instructions for implementing logical functions, which can be specifically implemented in any computer readable medium for use by or in conjunction with an instruction execution system, device or apparatus, such as a computer-based system, a system including a processor or other system that can fetch and execute instructions from the instruction execution system, device or apparatus. It should be understood that parts of the present application can be realized by hardware, software, firmware or their combination. In the above embodiments, a plurality of steps or methods can be realized by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if realized by hardware, it can be realized by any one or their combination of the following technologies known in the art: discrete logic circuit with logic gate circuit for implementing logical functions on data signals, application specific integrated circuit with suitable combination of logic gate circuit, programmable gate array (PGA), field programmable gate array (FPGA) and the like.
[0046] Additional aspects and advantages of the present application will be in part apparent and in part pointed out hereinafter in the description of embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0047] The above and / or additional aspects and advantages of the present application will become apparent and be readily appreciated from the following description of embodiments, taken in conjunction with the accompanying drawings.
[0048] Figure 1 is a schematic diagram of a device for predicting the probability of occurrence of an abnormal state of an individual according to an embodiment of the present application;
[0049] Figure 2 is a schematic diagram of a device for predicting the probability of occurrence of an abnormal state of an individual according to an embodiment of the present application (including an evaluation unit);
[0050] Figure 3 is a schematic diagram of a device for training a machine learning model according to an embodiment of the present application;
[0051] Figure 4 is a whole structure diagram of a complex disease risk prediction model of multiple source data combined with a multi-gene risk score, life habit characteristics and historical disease information according to an embodiment of the present application;
[0052] Figure 5 A flowchart of causal inference of life habits and diseases;
[0053] Figure 6 A structure diagram of a Transformer-Hawkes model;
[0054] Figure 7 A Manhattan plot of association study quantitative data of a single nucleotide polymorphism site of a European with coronary heart disease downloaded from “GWAS Catalog”—“https: / / www.ebi.ac.uk / gwas / home”;
[0055] Figure 8 A Manhattan plot of association study quantitative data of a single nucleotide polymorphism site of a European with type 2 diabetes downloaded from “GWAS Catalog”—“https: / / www.ebi.ac.uk / gwas / home”;
[0056] Figure 9 An AUROC evaluation diagram of a model predicting coronary heart disease;
[0057] Figure 10 An AUROC evaluation diagram of a model predicting type 2 diabetes;
[0058] Figure 11 Mean of absolute SHAP values of the top 20 global life habits of coronary heart disease;
[0059] Figure 12 SHAP values of the top 20 global life habits of coronary heart disease;
[0060] Figure 13 Mean of absolute SHAP values of the top 20 global life habits of type 2 diabetes;
[0061] Figure 14 SHAP values of the top 20 global life habits of type 2 diabetes;
[0062] Figure 15 SHAP values of the top 10 personalized life habits of a sample (2559491) with coronary heart disease;
[0063] Figure 16Shapley values of the top 10 personalized lifestyle habits for samples without coronary heart disease (2445247);
[0064] Figure 17 Shapley values of the top 10 personalized lifestyle habits for samples with type 2 diabetes (3069707);
[0065] Figure 18 Shapley values of the top 10 personalized lifestyle habits for samples without type 2 diabetes (3816331). DETAILED DESCRIPTION
[0066] Embodiments of the present application are described in detail below with reference to the attached drawing figures, wherein like reference numerals identify like elements or components throughout the figures. The embodiments described below are merely examples and do not limit the application, which can be practiced with various modifications and alterations.
[0067] Definitions and Descriptions
[0068] In this application, the "polygenic risk score" is a measure of an individual's risk of developing a complex disease. By using genomic variation (single nucleotide polymorphism sites, SNPs) data, the risk of developing a specific disease in an individual is calculated in a weighted sum manner. The calculation method of the polygenic risk score is usually based on whole genome association study data to find risk variants (such as SNPs) associated with a specific disease. By combining the results of these risk variants and weighting the contribution of each variant to the disease, a comprehensive polygenic risk score is obtained. This score can help predict the probability of an individual developing the disease and assess the individual's relative risk level in the population. Common polygenic risk score models include P+T, PRSice2, and LDPred2, etc.
[0069] In this application, the term "computer readable medium" can be any means that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer readable medium include the following: an electrical connection having one or more wires (electrical devices), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber device, and a portable compact disc read-only memory (CD-ROM). In addition, the computer readable medium can even be paper or other suitable medium on which the program is printed, such as by optical scanning, then by editing, interpreting, or otherwise processing the program to be electronically obtained, and then stored in the computer memory. The various computer readable storage media described in this application can represent one or more devices and / or other machine-readable storage media for storing information.
[0070] In this application, the term "machine readable storage medium" can include but not limited to wireless channel and various other media capable of storing, containing and / or carrying instructions and / or data.
[0071] It should be noted that the toolkit and model used in the technical solution of the present application are realized based on R and Python language. The machine learning model proposed in the present application has strong scalability, and based on the Python package, an interface for adding additional data information is provided, that is, the model can also be trained by adding other modal data or data sources (such as image information, text information and audio information, etc.), thereby producing more comprehensive and accurate analysis results.
[0072] Those skilled in the art of the present technology can understand that all or part of the steps carried out by the above-mentioned embodiment method can be completed by a program instructing the relevant hardware, and the program can be stored in a computer readable storage medium. When the program is executed, it includes one of the steps of the method embodiment or a combination thereof.
[0073] In addition, each functional unit in each embodiment of the present application can be integrated in one processing module, or each unit can exist physically independently, or two or more units can be integrated in one module. The above integrated module can be realized in the form of hardware or in the form of software functional module. When the integrated module is realized in the form of software functional module and sold or used as an independent product, it can also be stored in a computer readable storage medium.
[0074] Machine learning model training method
[0075] In an aspect of the present application, a machine learning model training method is provided, which is used to predict the probability of occurrence of an individual abnormal state. According to an embodiment of the present application, the method comprises: obtaining feature information of a training sample, the feature information comprising at least a multi-gene risk score feature, a lifestyle habit feature and a historical disease feature; and inputting the feature information into a machine learning model, using a known abnormal state of the training sample as a label, so as to perform supervised training on the machine learning model, wherein the lifestyle habit is determined to have a causal relationship with the abnormal state through Mendelian randomization analysis.
[0076] Specifically, in order to facilitate understanding, the training method is described in detail through the following specific description.
[0077] According to an embodiment of the present application, as shown in Figure 4 , the construction of the model for predicting the probability of occurrence of an individual abnormal state mainly includes five modules: a lifestyle habit feature screening module having a causal relationship with a specified abnormal state, a multi-gene risk score calculation module, a historical disease feature learning module, a probability of occurrence of an individual abnormal state calculation module and a user personalized lifestyle habit evaluation module (optionally added). The above modules are described in detail as follows:
[0078] 1. Lifestyle habit feature screening module having a causal relationship with a specified abnormal state Figure 5
[0079] 1.1 Data download
[0080] According to an embodiment of the present application, association measurement data of a single nucleotide polymorphism site (SNP) and an exposure variable (lifestyle habit), and association measurement data of multiple single nucleotide polymorphism sites and an outcome variable (specified abnormal state) are downloaded from “GWAS Catalog”—“https: / / www.ebi.ac.uk / gwas / home” or “Open GWAS”—“https: / / gwas.mrcieu.ac.uk / ”.
[0081] The above two kinds of data are processed into the format required by R language “TwoSampleMR” (for reference, https: / / mrcieu.github.io / TwoSampleMR / articles / exposure.html and https: / / mrcieu.github.io / TwoSampleMR / articles / outcome.html).
[0082] 1.2 Data preprocessing
[0083] According to the embodiments of the present application, the single nucleotide polymorphism sites (SNPs) strongly correlated with exposure variables (lifestyle habits) are extracted, and a threshold of significance P value is set, and the single nucleotide polymorphism sites higher than the threshold will be filtered out;
[0084] According to the embodiments of the present application, the method of “clump_data” in the R language “TwoSampleMR” package is used to ensure that the remaining single nucleotide polymorphism sites are independent of each other and to screen out single nucleotide polymorphism sites with a frequency of minor alleles greater than 0.01.
[0085] According to the embodiments of the present application, the R language “PhenoScanner” package is used to set a threshold of significance P value to detect whether the single nucleotide polymorphism sites are strongly correlated with multiple exposure variables, and if so, the data is filtered out.
[0086] It should be noted that filtering out data strongly correlated with exposure variables is to prevent collinearity between exposure variables. Generally speaking, collinearity may cause redundancy between explanatory variables in the model, making the stability and reliability of the model decline. In addition, the estimation of the coefficient may become unstable and difficult to interpret; and collinearity may cause redundancy between variables in the model, making statistical inference of the model difficult.
[0087] 1.3 Causal relationship analysis
[0088] According to the embodiments of the present application, the association of one single nucleotide polymorphism site in step 1.2 with lifestyle habits and the association of multiple single nucleotide polymorphism sites with specified abnormal states are matched one-to-many to form multiple pairs of data;
[0089] According to the embodiments of the present application, a Mendelian randomization model (commonly used models include “MR Egger” and “IVW”) is selected and the “mr” method in the R language “TwoSampleMR” package is used for Mendelian randomization analysis. Since there are multiple pairs of data, multiple Mendelian randomization results are obtained, and a meta-analysis is performed on the multiple Mendelian randomization results. The final result is considered to have a causal relationship between the lifestyle habit and the specified abnormal state if it satisfies all the following conditions:
[0090] a. The significance P value of one of the Mendelian randomization results is less than 0.05;
[0091] b. The fixed effect 95% confidence interval of the meta-analysis result does not contain 0;
[0092] c. The fixed effect significance P value of the meta-analysis result is less than 0.05;
[0093] d. The I square of the meta-analysis result is less than 0.05;
[0094] e.The P-value for the equivalence of the results of the meta-analysis is greater than 0.05.
[0095] It should be noted that the "meta-analysis" means a statistical analysis method for synthesizing and summarizing the results of multiple independent studies to obtain more comprehensive, accurate and reliable conclusions. By calculating the weighted average of different study effect sizes, meta-analysis can provide more stable and accurate effect estimates and a confidence interval for assessing the confidence level of the overall effect.
[0096] 2. Multi-gene risk score calculation module
[0097] 2.1 Data download
[0098] According to the embodiments of the present application, the association quantitative data of single nucleotide polymorphism sites and specified abnormal states of a large sample are downloaded from "GWAS Catalog"— "https: / / www.ebi.ac.uk / gwas / home" or "OpenGWAS"— "https: / / gwas.mrcieu.ac.uk / "; the genotype data of the population for which the multi-gene risk score is to be calculated is prepared in the file format of bed, bim, and fam (for reference "https: / / www.cog-genomics.org / plink / ").
[0099] 2.2 Data quality control
[0100] According to the embodiments of the present application, the association quantitative data of single nucleotide polymorphism sites and specified abnormal states obtained in step 2.1 are matched with the single nucleotide polymorphism sites in the International Human Genome Haplotype Map Project.
[0101] It should be noted that the International Human Genome Haplotype Map Project is a multi-country cooperative project aimed at developing a haplotype map of the human genome to describe the common patterns of human genetic variation and explore the differences in individual responses to human health, disease, drugs and environmental factors for different races.
[0102] 2.3 Multi-gene risk score calculation
[0103] According to the embodiments of the present application, the score of the single nucleotide polymorphism site is calculated by using the R language "bigsnpr" package (for specific steps, please refer to "https: / / privefl.github.io / bigsnpr / articles / LDpred2.html");
[0104] The score of the single nucleotide polymorphism site generated by the "bigsnpr" package described above and the "--score" command of the "PLINK" software (for details, please refer to "https: / / www.cog-genomics.org / plink / 1.9 / score") is used to calculate the polygenic risk score of the population.
[0105] 3. History disease feature learning module
[0106] 3.1 Training data acquisition
[0107] According to an embodiment of the present application, the training data contains two categories, the first category is the ICD-10 code (the first three digits) of the historical disease of the training population, and the second category is the time when the training population suffers from the historical disease. The time of the first historical disease diagnosis is set to 0, and the diagnosis time of the subsequent disease is the interval days from the first disease diagnosis time. The historical disease data is sorted according to the diagnosis time from small to large.
[0108] It should be noted that the ICD code refers to the International Classification of Diseases (ICD) coding system, which is an international standard for classifying and coding diseases, causes of diseases, causes, behavioral disorders and other health problems.
[0109] The ICD coding system uses a series of letters and numbers to represent different disease classifications and subdivisions. The ICD-10 coding system is the most commonly used version, in which the code is divided into multiple levels, from Chapter to Category, Subcategory, Subdivision, etc. Each code usually consists of several letters and numbers, and each code represents a specific disease or health problem.
[0110] 3.2 Structure of Transformer-Hawkes model and its training
[0111] According to an embodiment of the present application, the Transformer model and the Hawkes model are used to fit the historical disease information. For convenience of description, the two models are represented by the Transformer-Hawkes model below. The Transformer-Hawkes model specifically includes three main structures Figure 6):a.Transformer's encoding layer: a novel position encoding containing time information is used to improve the self-attention layer to learn the relationship between events occurring at different times, and the self-attention layer uses a masked multi-head attention model, which is more in line with the reality that events occurring at a certain time point will only be affected by previous events and not by subsequent events; b. Event prediction layer: used to predict the probability of an event occurring at a certain time point; c. Time prediction layer: used to predict the time of the next occurrence of an event;
[0112] According to an embodiment of the present application, the Transformer-Hawkes model is trained using the gradient descent method, the objective function is as follows, and is minimized,
[0113] where λ represents the intensity function, t represents time, and H represents events.
[0114] After the model is trained, the patient's historical disease data is input into the model, and the output of the encoding layer of the Transformer is the historical disease characteristics required by the subsequent disease risk prediction model.
[0115] 4、Individual abnormal state occurrence probability calculation module
[0116] 4.1 Training data integration
[0117] According to an embodiment of the present application, the data obtained in steps 1-3 is combined, and the age feature is added, which is the independent variable required for training the disease risk prediction model. Whether the test sample is an abnormal state is the dependent variable of the machine learning model.
[0118] 4.2 Individual abnormal state occurrence probability prediction model training
[0119] According to an embodiment of the present application, a stepwise logistic regression model is used to fit the data, and the output of the model is the occurrence probability of the individual abnormal state.
[0120] It should be noted that the stepwise logistic regression model is a statistical modeling method based on logistic regression, used to predict probability or risk scores in binary or multi-class classification problems. In this application, a stepwise logistic regression model is used to fit the data. By progressively selecting variables (polygenic risk scores, lifestyle habits, and historical diseases), the most relevant variables are chosen from the initial variable set and incorporated into the model. Its advantage lies in automatically selecting the most relevant variables, reducing redundant and irrelevant variables, and improving the model's generalization and interpretability. By progressively selecting variables, it helps identify the variables with the greatest predictive power for the target variable, thereby constructing a more accurate and concise predictive model.
[0121] 5. User Personalized Lifestyle Habit Assessment Module (Optional)
[0122] 5.1 Data and Model Preparation
[0123] According to embodiments of this application, the model uses data (training set and test set) consistent with the stepwise logistic regression model data;
[0124] 5.2 Machine Learning Explanation Model Training
[0125] According to embodiments of this application, the parameters of the stepwise logistic regression model are evaluated using the "LinearExplainer" function from the "SHAP" package in Python (see "https: / / shaplrjball.readthedocs.io / en / latest / generated / shap.LinearExplainer.html") in conjunction with the training and test sets, thereby obtaining a global lifestyle influence score. Figure 11 , Figure 12 , Figure 13 and Figure 14 Inputting individual data yields a personalized score reflecting the impact of lifestyle habits. Figure 15 , Figure 16 , Figure 17 and Figure 18 );
[0126] 5.3 Suggestions for Improving Lifestyle Habits
[0127] Using the "SHAP" interpreter obtained in the previous step, individual data can be input to generate personalized lifestyle habit impact scores. Individuals are then advised to improve lifestyle habits with high impact scores to reduce their risk of developing abnormal conditions.
[0128] It should be noted that the above data are all from open-source databases, and the corresponding abnormal state site information data can be selected according to training needs.
[0129] Advantages
[0130] According to the embodiments of the present application, the inventors found that the data source is single in the existing disease risk prediction model, and the relevant data of features such as dietary habits and exercise are lacking in the training, which greatly affects the accuracy and comprehensiveness of the prediction of complex diseases. In the present application, a causal inference model (Mendelian randomization) is used to screen life habit features that have a causal relationship with the specified abnormal state for training, so as to enhance the accuracy of model analysis and prediction. By integrating different data sources (genetic variation data, life habits and historical diseases) to build a model, the accuracy of disease prediction is improved. Moreover, the prediction model constructed in the present application can effectively calculate the contribution of each part of the information to the prediction of the individual abnormal state, and generate targeted solutions to improve life habits and other factors to prevent the occurrence or development of individual abnormal states.
[0131] It should be noted that the features and technical effects described in different aspects in this paper can be mutually referenced, and will not be repeated here.
[0132] The present application is illustrated by way of examples below, but this should not be understood as limiting the scope of the subject matter of the present application to the examples below. Any technology implemented based on the above description of the present application falls within the scope of the present application. The compounds or reagents used in the following examples can be obtained commercially or prepared by conventional methods known to those skilled in the art; the experimental instruments used can be obtained commercially.
[0133] Example 1: Machine learning model training and evaluation
[0134] In this embodiment, the machine learning model of the present application is trained using the UK Biobank European White Data (multigenetic risk score, life habit feature and historical disease feature related data), and there is no kinship among the European whites. The training and evaluation of the machine learning model are mainly based on the following five main stages: obtaining training data, training data preprocessing, disease risk prediction model training, model prediction ability evaluation and machine learning explanation model evaluation. The specific steps are as follows:
[0135] 1. Obtaining training data
[0136] 1.1 Analyzing and obtaining life habit data having a causal relationship with coronary heart disease and type 2 diabetes
[0137] 1.1.1 Download the association statistics between single nucleotide polymorphisms (SNPs) and lifestyle habits of Europeans (a total of 268 different lifestyle data points) from "openGWAS"—"https: / / gwas.mrcieu.ac.uk / " and "GWASCatalog"—"https: / / www.ebi.ac.uk / gwas / home" (268 different lifestyle data points were downloaded in total) and the association statistics between two SNPs and coronary heart disease and type 2 diabetes. Figure 7 and Figure 8 );
[0138] 1.1.2 Based on the data preprocessing and data cleaning steps described in the specific embodiments of this application, the above data is processed, and the lifestyle data that has a causal relationship with coronary heart disease and type 2 diabetes is obtained through the causal relationship analysis steps described in the specific embodiments of this application;
[0139] 1.2 Preprocessing and Extraction of Lifestyle Habits Data
[0140] 1.2.1 Screen out some useless lifestyle characteristics (e.g., whether you drank alcohol yesterday, whether you ate cheese yesterday, etc.);
[0141] 1.2.2 Remove lifestyle habits with a missing number greater than 44259 (since the number of European whites involved is 442591);
[0142] 1.2.3 Handle missing values: for numerical data, use the median value to fill in the missing values; for categorical data, use the most frequent value to fill in the missing values.
[0143] 1.2.4 Hot coding was performed on unordered categorical data with more than two categories to obtain 133 lifestyle features;
[0144] 1.2.5 The data on lifestyle habits retained in step 1.2.4 were intersected with the data on lifestyle habits causally related to the diseases obtained in step 1 (based on the same name), and then data extraction was performed. Ultimately, 65 and 45 lifestyle habit features causally related to coronary heart disease and type 2 diabetes were obtained, respectively (Tables 1 and 2).
[0145] Table 1: Lifestyle Habits with a Causal Relationship to Coronary Artery Disease
[0146]
[0147]
[0148] Table 2: Lifestyle characteristics with causal relationship to type 2 diabetes
[0149]
[0150]
[0151] 1.3 Extracting disease patients and dividing training set population and validation set population
[0152] 1.3.1 Set ICD-10 codes specifying abnormal states (coronary heart disease: I21, I22, I23, I241, I252; type 2 diabetes: E11) and match in historical disease diagnosis data of participants in UK Biobank;
[0153] 1.3.2 Set disease codes specifying abnormal states in self-reported non-cancer disease data in UK Biobank (coronary heart disease: 1075; type 2 diabetes: 1223) and match in self-reported non-cancer disease data in UK Biobank;
[0154] 1.3.3 Filter out disease patients without genotype data and historical diagnosis information (calculate polygenic risk scores using genotype data and historical disease information data and train disease risk prediction models), finally extract 9727 European white people with coronary heart disease and 2643 European white people with type 2 diabetes;
[0155] 1.3.4 Divide the self-reported disease patients and the patients whose disease diagnosis time is before the division point into the training set population, and the patients whose disease diagnosis time is after the division point into the validation set population, with the time of participating in the UK Biobank study as the division point. Finally, randomly select the same number of people as the number of patients with diseases from the European white people without diseases into the training set and the validation set;
[0156] 1.4 Calculate the "LDPred2" polygenic risk scores of coronary heart disease and type 2 diabetes respectively using the genotype data of European white people in UK Biobank (see the specific embodiment part for details);
[0157] 1.5 Train the Transformer-Hawkes model and use the output of the encoding layer of the Transformer as the historical disease feature;
[0158] 1.5.1 Model training (see the specific embodiment part for details);
[0159] 1.5.2 Historical disease feature output of the participant:
[0160] Process the historical diagnosis data of the participant. If it is a disease patient, extract the historical disease diagnosis information before the diagnosis of the specified abnormal state; if it is a non-disease patient, no additional processing is required;
[0161] The extracted diagnostic data is input into the model, and the encoding layer output of the Transformer at the last time node in the diagnostic data is output as the historical disease characteristics of the tracker.
[0162] 2. Preprocessing of training data
[0163] 2.1 Standardize the multi-gene risk score and the lifestyle habit data of continuous type, i.e. make the mean of the data 0 and the variance 1;
[0164] 2.2 Calculate the correlation coefficient between lifestyle habits (eliminate the collinearity problem of lifestyle habit data), find that arm, leg and whole body fat data have strong correlation with body weight, therefore, only the data of body weight characteristics can be retained. Then splice the multi-gene risk score, age and lifestyle habit data;
[0165] 2.3 Splice the data of step 2.2 with the historical disease information data to obtain the target training data.
[0166] 3. Training of disease risk prediction model and evaluation of prediction ability
[0167] 3.1 Use the stepwise logistic regression model to fit the target training data of step 2.3;
[0168] 3.2 Use the "roc_curve" and "auc" methods in the "Sklearn" package to evaluate the prediction results of the model on the validation set, and compare the AUROC scores of the model prediction ability when different types of data are input ( Figure 9 and Figure 10 ).
[0169] The results show that the AUROC scores of the trained machine learning model in predicting coronary heart disease and type 2 diabetes can reach 0.81 and 0.86 respectively.
[0170] 4. Evaluation of machine learning explanation model
[0171] 4.1 Combine the training data of step 2.3, the validation data of step 1.3.4 and the disease risk regression model, and use the "LinearExplainer" in the "SHAP" package to obtain a global model parameter interpreter;
[0172] The results show that large waist circumference (body weight), old age, frequent smoking, drowsiness and lack of exercise increase the risk of coronary heart disease ( Figure 11 and Figure 12 ); large body weight, old age, long time watching TV and lack of exercise increase the risk of type 2 diabetes ( Figure 13 and Figure 14 ).
[0173] 4.2 Using the global model parameter interpreter, by inputting individual data, obtaining personalized lifestyle impact scores;
[0174] The results show that the lifestyle features with high impact scores for samples with coronary heart disease are large waist circumference, frequent drinking and older age, while the impact scores of samples without coronary heart disease are relatively low on these features Figure 15 and Figure 16 Therefore, it is recommended that high-age users do moderate exercise to reduce fat and reduce the frequency of drinking; the lifestyle features with high impact scores for samples with type 2 diabetes include overweight and long-time TV watching, while the impact scores of samples without type 2 diabetes are relatively low on these lifestyle habits Figure 17 and Figure 18 Therefore, it is recommended that users reduce TV usage time, avoid sitting for a long time and do more moderate exercise.
[0175] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example" or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0176] Although the embodiments of the present application have been shown and described, those skilled in the art can understand that various changes, modifications, replacements and variations can be made to the embodiments without departing from the principles and purposes of the present application, and the scope of the present application is defined by the claims and their equivalents.
Claims
1. A machine learning model training method, the machine learning model being used to predict a probability of occurrence of an individual abnormal state, characterized in that, The method comprises: obtaining feature information of a training sample, the feature information comprising at least a polygenic risk score feature, a lifestyle habit feature, and a historical disease feature; and inputting the feature information into a machine learning model, using known abnormal states of the training sample as labels to supervise training of the machine learning model, wherein the lifestyle habit is determined to have a causal relationship with the abnormal state through Mendelian randomization analysis.
2. The method of claim 1, wherein, The historical disease feature is obtained by processing disease information using a preprocessing model, the disease information comprising type information and time information. The preprocessing model is trained through the following steps: obtaining the disease information of the training sample; inputting the disease information into an encoding layer to digitally process the disease information; inputting the result of the digital processing into the preprocessing model to obtain converted disease features; inputting the disease features into a machine learning model and training the preprocessing model and the machine learning model using known disease information of the training sample.
3. The method of claim 2, wherein, The preprocessing model is a Transformer model, which comprises a masked multi-head attention layer, a first normalization layer, a position-wise feed-forward network, and a second normalization layer.
4. The method of claim 1, wherein, The polygenic risk score feature is obtained by processing single nucleotide polymorphism site scores using an LDPred2 model.
5. The method of claim 1, wherein, The abnormal state comprises coronary heart disease, and the lifestyle habit feature comprises at least one of the following:
6. The method of claim 1, wherein, The abnormal state is type 2 diabetes, and the lifestyle habit feature comprises at least one of the following:
7. The method of claim 1, wherein, The machine learning model is a stepwise logistic regression model.
8. An apparatus for predicting a probability of occurrence of an abnormal state of an individual, characterized by comprising: Comprise: a feature information obtaining unit configured to obtain feature information of an individual, the feature information comprising a polygenic risk score feature, a lifestyle habit feature, and a historical disease feature; a prediction unit connected to the feature information obtaining unit and configured to predict a probability of the individual developing an abnormal state based on the feature information using a machine learning model trained using the method of any one of claims 1-7.
9. The apparatus of claim 8, wherein, Further comprise: an evaluation unit configured to evaluate the individual's lifestyle habits and / or generate information for improving the individual's lifestyle habits.
10. An apparatus for training a machine learning model, the apparatus comprising: Comprise: a feature information obtaining unit configured to obtain feature information of a training sample, the feature information comprising at least a polygenic risk score feature, a lifestyle habit feature, and a historical disease feature; and a training unit connected to the feature information obtaining unit and configured to input the feature information into a machine learning model, using known abnormal states of the training sample as labels to supervise training of the machine learning model; wherein the lifestyle habit is determined to have a causal relationship with the abnormal state through Mendelian randomization analysis.
11. A server, characterized by The server comprises a processor and a memory, and the memory has stored thereon a computer program which, when executed by the processor, implements the machine learning model training method of any one of claims 1-7.
12. A computer readable storage medium embodying a computer program, wherein, The computer program, when executed by one or more processors, implements the machine learning model training method of any one of claims 1-7.