A multi-state model prediction system based on environment-gene interaction
Through a multi-state model prediction system based on environmental-gene interaction, the problem of being unable to dynamically predict the probability of disease metastasis and explain the differences in comorbidities in the prior art, and the accurate assessment of multiple diseases and screening of susceptible populations is achieved.
Patent Information
- Application Number
- CN202211336214.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-28
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-10-28
AI Technical Summary
The prior art cannot dynamically predict the probability of metastasis of different diseases and cannot explain the individual differences in the development trajectory of comorbidities.
It provides a multi-state model prediction system based on environmental-gene interaction, including a database module, data preprocessing module, model analysis module, disease assessment and prediction module and information feedback module. It analyzes the probability and health effects of disease state transfer through multi-state models, and combines environmental exposure assessment data, gene susceptibility data and medical diagnosis and treatment records for evaluation and prediction.
It can more accurately assess the incidence and outcome of multiple diseases, comprehensively explain individual differences in the development trajectory of comorbidities, screen susceptible populations, and provide scientific basis for the assessment of disease risk.
Smart Images

Figure CN115497591B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of disease state prediction, and more particularly, to a multi-state model prediction system based on environment-gene interaction. Background Art
[0002] Multiple diseases or comorbidities refer to an individual suffering from two or more chronic diseases simultaneously. Comorbidities can reduce the quality of life of patients, increase the difficulty of disease treatment and management, as well as medical expenses, increase the risks of disability and death, and impose a huge burden and challenge on patients and the medical and health service system. Currently, comorbidities have gradually become a major threat to residents' health and an important public health issue. Understanding the occurrence and development process of comorbidities and identifying modifiable risk factors are of great significance for formulating reasonable comorbidity management strategies and measures.
[0003] Previous studies have shown that environmental factors such as air pollution play an important role in the progression of chronic diseases, and the impact of environment-gene interaction on disease occurrence has also received extensive attention. By studying the role of exposure factors at different stages of disease progression, it is possible to further clarify their impact at different stages of the occurrence and development of comorbidities, which helps to systematically understand the association between comorbidities and their potential influencing factors. Considering that the occurrence and development of comorbidities is a long-term, multi-stage, and reversible process, using longitudinal follow-up data to establish a multi-state model to study the natural history of the occurrence and development of comorbidities has the following advantages: (1) it can describe in detail the process of disease state changes over time; (2) it can dynamically predict the transition probabilities between different disease states; (3) it can evaluate the health effects of influencing factors at different stages; (4) it can estimate the residence time of an individual in a specific disease state.
[0004] The prior art discloses a disease prediction method and system based on a prediction model. This method constructs a disease prediction model by selecting a disease-specific database based on the disease prediction purpose; trains the disease prediction model; validates the trained disease prediction model; and uses the validated disease prediction model for disease prediction. The defect of this solution is that it cannot dynamically predict the transition probabilities of different diseases and cannot explain the individual differences in the occurrence and development trajectories of comorbidities.
[0005] Therefore, in combination with the above requirements and the defects in the prior art, the present application proposes a multi-state model prediction system based on environment-gene interaction. Summary of the Invention
[0006] The present invention provides a multi-state model prediction system based on environment-gene interaction, which can make full use of data from fields such as clinical, experimental, environmental, and financial, more accurately evaluate the incidence and prognosis of multiple diseases, be able to comprehensively explain the individual differences in the occurrence and development trajectories of comorbidities, screen susceptible populations, and provide a scientific basis for evaluating the risk of disease occurrence.
[0007] The primary objective of the present invention is to solve the above-mentioned technical problems, and the technical solution of the present invention is as follows:
[0008] The present invention provides, in a first aspect, a multi-state model prediction system based on environment-gene interaction. This system includes: a database module, a data preprocessing module, a model analysis module, a disease assessment and prediction module, and an information feedback module; the database module is used to store, read and write, and manage the personal basic information and health-related data of the research object; the data preprocessing module is used to obtain the research object data in the database module and perform preprocessing to obtain a feature data set; the model analysis module uses a multi-state model to analyze the feature data set, obtains the transition probabilities between different disease states, and calculates the health effect estimates of the influencing factors at each transition stage; the disease assessment and prediction module uses the transition probabilities and health effect estimates to evaluate the incidence risks and dynamic progression trajectories of multiple diseases in the feature data set, integrates the evaluation data with the feature data set into a sample training set, uses the sample training set to train the multi-state model, and uses the trained multi-state model to evaluate and predict the object to be predicted; the information feedback module is used to store the disease assessment and prediction data in the database module and feedback the data to the research object after visualizing the data.
[0009] Further, the personal basic information of the research object includes: age, gender, family income, and marital status, and the health-related data includes: environmental exposure assessment data, gene susceptibility data, and medical treatment records.
[0010] Among them, the environmental exposure assessment data includes: the average concentrations of particulate matter and gaseous pollutants to which the research object is exposed daily in the air; the gene susceptibility data includes: single nucleotide polymorphisms, polygenic risk scores for specific diseases.
[0011] Among them, the medical treatment records include: the incidence or history of multiple diseases, prognosis, medication history, surgical history, blood and urine biochemical indexes, and organ function indexes.
[0012] Further, the specific process of the preprocessing is: obtaining the personal basic information and health-related data of the research object, cleaning the research object data, filling in the missing values of the data by means of linear interpolation and mean filling, converting the measurement units and data formats of the research object data, and merging and integrating the information of different sub-databases.
[0013] Further, the transition probability is calculated by the following formula:
[0014] P hj(s, t) = Prob(X(t) = j | X(s) = h)
[0015] where P hj (s, t) is the transition probability, which is used to indicate the possibility that an individual in state h at time s is in state j at a future time t.
[0016] Furthermore, the specific value of the health effect estimate is: calculating the hazard ratio HR of each single outcome event and its 95% confidence interval through the Cox proportional hazards model; the specific Cox proportional hazards model and hazard ratio HR are as follows:
[0017] h(t) = h0(t) × exp(β l X1 + β2X2 + … + β3X3)
[0018]
[0019] j = 1, 2, …, n
[0020] where t is the time node, X p is the independent variable included in the model, β p is the partial regression coefficient of the variable, h0(t) is the hazard function at time t when all independent variables are 0, and h(t) is the hazard function at time t.
[0021] Furthermore, the specific method for evaluating the risk of multiple diseases and the dynamic progression trajectory in the feature dataset is: marking different outcome events or states, where the outcome events or states include: healthy, diseased, dead, and others, which are respectively marked as 1, 2, 3, and more numbers; defining state progression variables according to the progression direction between different states and reconstructing the dataset, that is, adding three variable information of the starting state, the terminal state, and the state progression variable to the original dataset respectively to construct a multi-state model and calculating the hazard ratio at different transition stages.
[0022] Furthermore, the specific process of using the trained multi-state model to evaluate and predict the object to be predicted is: training the multi-state model based on the sample training set to obtain stable multi-state model parameters, inputting the current personal basic information and health-related data of the unknown research object as the initial value into the trained multi-state model, outputting the predicted value of the disease onset risk of the research object at a future time point, as well as the transition probability between different states, and performing summary analysis based on the continuous dynamic transition probability to obtain the disease dynamic progression trajectory.
[0023] Further, the environment-gene interaction includes additive interaction and multiplicative interaction between environmental exposure information and genetic susceptibility. The additive interaction is used to explore whether the combined effect of environmental exposure information and genetic susceptibility is equal to the sum of the independent effects of the two factors, and the multiplicative interaction is used to explore whether the combined effect of environmental exposure information and genetic susceptibility is equal to the product of the independent effects of the two factors.
[0024] Among them, if the additive interaction does not exist, when two factors act jointly on a certain disease, their effect is equal to the sum of the independent effects of these factors, that is, R 11 -R 00 =(R 10 -R 00 )+(R 01 -R 00 ); among them, R 00 represents the incidence rate and other frequency indicators not exposed to the two factors, R 01 represents the incidence rate and other frequency indicators only exposed to factor 1, R 10 represents the incidence rate and other frequency indicators only exposed to factor 2, R 11 represents the incidence rate and other frequency indicators exposed to the two factors simultaneously.
[0025] Among them, if the multiplicative interaction does not exist, when two factors act jointly on a certain disease, their effect is equal to the product of the independent effects of these factors, that is, R 11 / R 00 =(R 10 / R 00 )×(R 01 / R 00 ).
[0026] Further, the evaluation indexes of the environment-gene interaction include: the effect estimate value and 95% confidence interval of the interaction term in the multi-state model, the excess relative risk of interaction RERI, the attributable proportion of interaction AP, and the synergy index SI.
[0027] Among them, when using the multiplicative model to test the interaction, a product term is included in the regression model to test the significance of the product term coefficient; if the significance P of the product term coefficient estimate value > 0.05, it is prompted in the information feedback module that the product interaction term has no statistical significance; when using the additive model to test the interaction, several indexes and their 95% confidence interval CI are used, and the indexes include: the excess relative risk of interaction RERI, the attributable percentage of interaction AP, and the synergy index SI.
[0028] Further, the mathematical expression forms of the indexes are as follows:
[0029] RERI = R11 -R 01 -R 10 +1
[0030] AP = RERI / R 11
[0031] SI = (R 11 -1) / [(R 01 -1)+(R 10 -1)]
[0032] Among them, R 00 , R 01 , R 10 , R 11 respectively represent the effect estimates of non-exposure, exposure only to factor 1, exposure only to factor 2, and simultaneous exposure to both factors; when the 95% confidence intervals CI of the excess relative risk of interaction RERI and the attributable percentage of interaction AP do not contain 0, and the 95% confidence interval CI of the synergy index SI does not contain 1, it indicates that there is an additive interaction effect between the two factors.
[0033] Furthermore, the personal basic information is obtained through a personal information questionnaire; the environmental exposure assessment data is obtained through public data sources or non-public data sources. The public data sources include: air pollution concentration, temperature, and humidity data published in real time by environmental monitoring stations under the China National Environmental Monitoring Centre and the National Meteorological Science Data Centre; the non-public data sources include data obtained from environmental quality monitoring stations set up by the scientific research project team; the gene susceptibility data is obtained by obtaining basic information through laboratory gene sequencing technology and then performing genotyping on genes and genetic markers in combination with genome-wide association studies GWAS; the medical diagnosis and treatment records are obtained through a medical history questionnaire and the hospital health registration system.
[0034] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0035] The present invention provides a multi-state model prediction system based on environment-gene interaction. By preprocessing the data of the research object through a data preprocessing module and further analyzing the data of the research object in multiple aspects by using a model analysis module and a disease assessment and prediction module, it can more accurately evaluate the incidence and prognosis of multiple diseases, comprehensively explain the individual differences in the occurrence and development trajectories of comorbidities, screen susceptible populations, and provide a scientific basis for evaluating the risk of disease occurrence. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a schematic diagram of a multi-state model prediction system based on environment-gene interaction of the present invention.
[0037] Figure 2Schematic diagram of preprocessing the data of the research object by the present invention to obtain a feature dataset.
[0038] Figure 3 Schematic diagram of the multi-state model framework for disease state transition in an embodiment of the present invention.
[0039] Figure 4 Example diagram of the number and probability of different disease state transitions in an embodiment of the present invention.
[0040] Figure 5 Example diagram of predicting the disease onset risk and dynamic progression trajectory of unknown research objects in an embodiment of the present invention. Detailed implementation manners
[0041] In order to more clearly understand the above objects, features and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other.
[0042] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0043] Embodiment 1
[0044] As Figure 1 shown, the present invention provides a multi-state model prediction system based on environment-gene interaction. This system includes: a database module, a data preprocessing module, a model analysis module, a disease evaluation and prediction module, and an information feedback module; the database module is used to store, read and write, and manage the personal basic information and health-related data of the research object; the data preprocessing module is used to obtain the research object data in the database module and perform preprocessing to obtain a feature dataset; the model analysis module uses a multi-state model to analyze the feature dataset to obtain the transition probabilities between different disease states, and calculates the health effect estimates of the influencing factors in each transition stage; the disease evaluation and prediction module uses the transition probabilities and health effect estimates to evaluate the multiple disease onset risks and dynamic progression trajectories in the feature dataset, and integrates the evaluation data with the feature dataset into a sample training set, uses the sample training set to train the multi-state model, and uses the trained multi-state model to evaluate and predict the object to be predicted; the information feedback module is used to store the disease evaluation and prediction data in the database module, and feedback the data to the research object after visualizing the data.
[0045] Further, the personal basic information of the research object includes: age, gender, family income, and marital status, and the health-related data includes: environmental exposure assessment data, gene susceptibility data, and medical diagnosis and treatment records.
[0046] Among them, the environmental exposure assessment data includes: the average concentration of particulate matter and gaseous pollutants to which the research object is exposed daily in the air; the gene susceptibility data includes: single nucleotide polymorphisms, polygenic risk scores for specific diseases.
[0047] Among them, the medical diagnosis and treatment records include: the onset or disease history, prognosis, medication history, surgical history, hematuria biochemical indicators, and organ function indicators of multiple diseases.
[0048] Further, the specific process of the preprocessing is: obtaining the personal basic information and health-related data of the research object, cleaning the data of the research object, filling in the missing values of the data by means of linear interpolation and mean filling, converting the measurement units and data formats of the data, and merging and integrating the information of different sub-databases.
[0049] In a specific embodiment, statistical methods such as linear interpolation and mean filling are used for the missing data values to improve the integrity of the collected information; the data conversion includes the conversion of measurement units and data formats such as long data or wide data, making the data formats standardized and unified, with good repeatability; the data integration includes the merging of information from different sub-databases, and the multiple data sources increase the richness of the information, which can improve the accuracy and precision of the disease prediction model.
[0050] Further, the transition probability is calculated by the following formula:
[0051] P hj (s, t) = Prob(X(t) = j|X(s) = h)
[0052] Among them, P hj (s,t) is the transition probability, which is used to indicate the possibility of an individual being in state h at time s and in state j at a future time t.
[0053] Further, the specific health effect estimate is: calculating the hazard ratio HR of each single outcome event and its 95% confidence interval through the Cox proportional hazards model; where the Cox proportional hazards model and the hazard ratio HR are specifically:
[0054] h(t) = h0(t) × exp(β1X1 + β2x2 + … + β3X3)
[0055]
[0056] j = 1, 2, …, n
[0057] where t is the time node, X p is the independent variable incorporated into the model, β p is the partial regression coefficient of the variable, h0(t) is the risk function at time t when all independent variables are 0, and h(t) is the risk function at time t.
[0058] Furthermore, the method for evaluating the incidence risk and dynamic progression trajectory of multiple diseases in the feature dataset is specifically as follows: different outcome events or states are marked, where the outcomes or states include: healthy, diseased, dead, and others, which are marked as 1, 2, 3, and more numbers respectively; state progression variables are defined according to the progression direction between different states and a dataset is reconstructed, and the reconstructed dataset is to add three variable information of the starting state, the terminal state, and the state progression variable respectively on the basis of the original dataset to construct a multi-state model and calculate the risk ratio of different transition stages.
[0059] Furthermore, the specific process of using the trained multi-state model to evaluate and predict the object to be predicted is as follows: the multi-state model is trained based on the sample training set to obtain stable multi-state model parameters, and the current personal basic information and health-related data of the unknown research object are input into the trained multi-state model as initial values, and the predicted value of the disease incidence risk of the research object at a future time point and the transition probability between different states are output, and the disease dynamic progression trajectory is obtained through summary analysis according to the continuous dynamic transition probability.
[0060] Furthermore, the environment-gene interaction includes the additive interaction and multiplicative interaction between environmental exposure information and genetic susceptibility. The additive interaction is used to explore whether the combined effect of environmental exposure information and genetic susceptibility is equal to the sum of the independent effects of the two factors, and the multiplicative interaction is used to explore whether the combined effect of environmental exposure information and genetic susceptibility is equal to the product of the independent effects of the two factors.
[0061] where, if the additive interaction does not exist, when the two factors act on a certain disease together, their effect is equal to the sum of the independent actions of these factors, that is, R 11 -R 00 =(R 10 -R 00 )+(R 01 -R 00 ); where, R 00 represents the incidence rate and other frequency indicators not exposed to the two factors, R 01 represents the incidence rate and other frequency indicators only exposed to factor 1, R 10 represents the incidence rate and other frequency indicators only exposed to factor 2,11 Indicates the incidence rate and other frequency indicators of simultaneous exposure to two factors.
[0062] Among them, if the multiplicative interaction does not exist, when two factors act jointly on a certain disease, its effect is equal to the product of the independent effects of these factors, that is, R 11 / R 00 =(R 10 / R 00 )×(R 01 / R 00 ).
[0063] Furthermore, the evaluation indexes of the environment-gene interaction include: the effect estimate value and 95% confidence interval of the interaction term in the multi-state model, the excess relative risk of interaction RERI, the attributable proportion of interaction AP, and the synergy index SI.
[0064] Among them, when using the multiplicative model to test the interaction, a product term is incorporated into the regression model to test the significance of the product term coefficient; if the significance P of the product term coefficient estimate value > 0.05, it is prompted in the information feedback module that the product interaction term has no statistical significance; when using the additive model to test the interaction, several indexes and their 95% confidence interval CI are used, and the indexes include: the excess relative risk of interaction RERI, the attributable percentage of interaction AP, and the synergy index SI.
[0065] Furthermore, the mathematical expression forms of the indexes are as follows:
[0066] RERI = R 11 -R 01 -R 10 +1
[0067] AP = RERI / R 11
[0068] SI = (R 11 -1) / [(R 01 -1)+(R 10 -1)]
[0069] Among them, R 00 、R 01 、R 10 、R 11 respectively represent the effect estimate values of non-exposure, exposure to factor 1 only, exposure to factor 2 only, and simultaneous exposure to two factors; when the 95% confidence interval CI of the excess relative risk of interaction RERI and the attributable percentage of interaction AP does not contain 0, and the 95% confidence interval CI of the synergy index SI does not contain 1, it indicates that there is an additive interaction effect between the two factors.
[0070] Furthermore, the personal basic information is obtained through a personal information questionnaire; the environmental exposure assessment data is obtained from public data sources or non-public data sources. The public data sources include: real-time air pollution concentration, temperature, and humidity data published by environmental monitoring stations under the China National Environmental Monitoring Center and the National Meteorological Science Data Center. The non-public data sources include data obtained from environmental quality monitoring stations independently set up by the research project team; the gene susceptibility data is obtained by obtaining basic information through laboratory gene sequencing technology and then performing genotyping on genes and genetic markers in combination with genome-wide association studies (GWAS); the medical diagnosis and treatment records are obtained through a medical history questionnaire and the hospital health registration system.
[0071] Example 2
[0072] Based on the above Example 1, combined with Figures 2 to 4 , this example details the process of using this system to conduct a multiple disease state transition study on the research object.
[0073] In a specific embodiment, the original data of the research object is extracted from the database module. After the data preprocessing module performs preprocessing processes such as cleaning, transformation, and integration on the research object data, an example diagram of the feature dataset obtained is as Figure 2 shown. This dataset contains the health follow-up data of 500 research objects. The health outcomes of the follow-up include diabetes, cardiovascular disease, and death; the personal basic information includes age and gender; the health-related data includes the average concentration of nitrogen dioxide (NO2) to which the research object is exposed to in the air every day, and the gene susceptibility score, that is, the polygenic risk score (PRS) for a specific disease. Among them, the average concentration of nitrogen dioxide (NO2) has been converted to the logarithm of the actual concentration.
[0074] Construct a custom multi-state model framework for disease state transition, mainly including specified disease transition states such as healthy, diseased, and dead; transition paths such as Path A from healthy to diseased, and Path B from diseased to dead. In a specific embodiment, as Figure 3 shown, the constructed multi-state model framework includes 4 specified disease transition states and 5 transition paths. Among them, the 4 specified disease transition states include: healthy, diabetes, cardiovascular disease, and dead; the 5 transition paths include: ① from healthy to diabetes ② from diabetes to cardiovascular disease ③ from diabetes to dead ④ from cardiovascular disease to dead ⑤ from healthy to dead.
[0075] In a specific embodiment, based on the feature dataset obtained by the data preprocessing module, multi-state model analysis is performed using the mstate and survival packages of R 4.1.0 software and the multistate package in Stata / SE 16.0 software. The results are as Figure 4As shown, an example graph of the number and probability of different disease state transitions is obtained.
[0076] Example 3
[0077] Based on the above Example 1 and Example 2, combined with Figure 5 , this example elaborates in detail the process of predicting the disease onset risk and dynamic progression trajectory of an unknown research object.
[0078] In a specific example, the personal health information of this unknown research object is obtained through a questionnaire survey. She is an adult female over 40 years old, and this research object is exposed to high levels of nitrogen dioxide (NO2) daily, with a high gene susceptibility score (indicating a relatively high disease onset risk).
[0079] As Figure 5 shown, the prediction results of this system show that by the 10th year of follow-up, this research object has approximately a 1% probability of developing cardiovascular disease, a 19% probability of developing diabetes, and a 75% probability of death.
[0080] The icons describing the structural position relationships in the drawings are only for illustrative purposes and should not be construed as limitations on this patent.
[0081] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention and are not limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the claims of the present invention.
Claims
1. A multi-state model prediction system based on environment-gene interaction, characterized in that, It includes: a database module, a data preprocessing module, a model analysis module, a disease assessment and prediction module, and an information feedback module; The database module is used to store, read, write, and manage the personal basic information and health-related data of the research object; The data preprocessing module is used to obtain the research object data in the database module and perform preprocessing to obtain a feature data set; The model analysis module uses a multi-state model to analyze the feature data set, obtains the transition probabilities between different disease states, and calculates the health effect estimates of the influencing factors at each transition stage; The specific health effect estimate is: calculate the hazard ratio HR and its 95% confidence interval of each single outcome event through the Cox proportional hazards model; The specific Cox proportional hazards model and hazard ratio HR are: h(t) = h0(t) × exp(β1X1 + β2X2 + … + β p X p ) j = 1, 2, …, n Among them, X1, X2, ……, X p , X 11 , X 12 , ……, X 1p and X j1 , X j2 , ……, X jp are all independent variables incorporated into the model, β 1, β2, ……, β p are the partial regression coefficients of the variables; h0(t) is the risk function at time t when all independent variables are 0, h(t) is the risk function at time t, h1(t) is the risk function at time t when all independent variables are 1, and h j (t) is the risk function at time t when all independent variables are j; The disease assessment and prediction module uses the transition probability and health effect estimate to evaluate the multiple disease incidence risks and dynamic progression trajectories in the feature data set, integrates the evaluation data with the feature data set into a sample training set, uses the sample training set to train the multi-state model, and uses the trained multi-state model to evaluate and predict the object to be predicted; The specific method for evaluating the multiple disease incidence risks and dynamic progression trajectories in the feature data set is: mark different outcome events or states, define state progression variables according to the progression directions between different states and reconstruct the data set, that is, add three variable information of the starting state, the terminal state, and the state progression variable to the original data set respectively, construct a multi-state model, and calculate the hazard ratio of different transition stages; The specific process of using the trained multi-state model to evaluate and predict the object to be predicted is: train the multi-state model based on the sample training set to obtain stable multi-state model parameters, input the current personal basic information and health-related data of the unknown research object as the initial value into the trained multi-state model, output the predicted value of the disease incidence risk of the research object at a future time point, and the transition probabilities between different states, and perform summary analysis based on the continuous dynamic transition probabilities to obtain the disease dynamic progression trajectory; The information feedback module is used to store the evaluation and prediction data in the database module and visualize the data.
2. The multi-state model prediction system based on environment-gene interaction according to claim 1, wherein The personal basic information of the research object includes: age, gender, family income, and marital status, and the health-related data includes: environmental exposure assessment data, gene susceptibility data, and medical diagnosis and treatment records; Among them, the medical diagnosis and treatment records include: the incidence or history of multiple diseases, prognosis, medication history, surgical history, hematuria biochemical indicators, and organ function indicators.
3. The multi-state model prediction system based on environment-gene interaction according to claim 1, wherein The specific process of the preprocessing is: obtain the personal basic information and health-related data of the research object, clean the research object data, use linear interpolation and mean filling methods to fill in the missing values of the data, perform conversion on the measurement units and data formats of the research object data, and merge and integrate the information of different sub-databases.
4. The multi-state model prediction system based on environment-gene interaction according to claim 1, wherein The transition probability is calculated by the following formula: P hj (s, t) = Prob(X(t) = j | X(s) = h) where, P hj (s, t) is the transition probability, which is used to indicate the possibility that an individual in state h at time s will be in state j at a future time t.
5. The multi-state model prediction system based on environment-gene interaction according to claim 2, wherein The environmental-gene interaction includes additive interaction and multiplicative interaction between environmental exposure information and genetic susceptibility; The additive interaction is used to explore whether the combined effect of environmental exposure information and genetic susceptibility is equal to the sum of the independent effects of the two factors. If there is no interaction, when the two factors act jointly on a certain disease, their effects are equal to the sum of the independent effects of these factors, that is, R 11 -R 00 =(R 10 -R 00 )+(R 01 -R 00 ); where R 00 represents the incidence rate and other frequency indicators of not being exposed to the two factors, R 01 represents the incidence rate and other frequency indicators of being exposed only to factor 1, R 10 represents the incidence rate and other frequency indicators of being exposed only to factor 2, R 11 represents the incidence rate and other frequency indicators of being exposed to the two factors simultaneously; The multiplicative interaction is used to explore whether the combined effect of environmental exposure information and genetic susceptibility is equal to the product of the independent effects of the two factors; if there is no interaction, when two factors act jointly on a certain disease, their effect is equal to the product of the independent effects of these factors, that is, R 11 / R 00 =(R 10 / R 00 )×(R 01 / R 00 ).
6. The multi-state model prediction system based on environment-gene interaction according to claim 5, wherein The evaluation indicators of the environmental-gene interaction include: the effect estimate of the interaction term in the multi-state model, the excess relative risk of interaction RERI, the attributable proportion of interaction AP, and the synergy index SI; Among them, when using the multiplicative model to test the interaction, a product term is included in the regression model to test the significance of the coefficient of the product term; if the significance P of the coefficient estimate of the product term > 0.05, it is prompted in the information feedback module that the product interaction term is not statistically significant; when using the additive model to test the interaction, several indicators and their 95% confidence intervals CI are used, and the indicators include: the excess relative risk of interaction RERI, the attributable percentage of interaction AP, and the synergy index SI; the mathematical expression forms of the indicators are as follows: RERI = R 11 -R 01 -R 10 +1 AP = RERI / R 11 SI = (R 11 - 1) / [(R 01 - 1)+(R 10 - 1)] Among them, R 00 , R 01 , R 10 , R 11 respectively represent the effect estimates of non-exposure, exposure only to factor 1, exposure only to factor 2, and simultaneous exposure to both factors; when the 95% confidence intervals CI of the excess relative risk of interaction RERI and the attributable percentage of interaction AP do not contain 0, and the 95% confidence interval CI of the synergy index SI does not contain 1, it indicates that there is an additive interaction effect between the two factors.
7. The multi-state model prediction system based on environment-gene interaction according to claim 2, characterized in that, The personal basic information is obtained through a personal information questionnaire; the environmental exposure assessment data is obtained through public data sources or non-public data sources. The public data sources include: air pollution concentration, temperature, and humidity data published in real time by environmental monitoring stations under the China National Environmental Monitoring Centre and the National Meteorological Science Data Centre; the non-public data sources include data obtained from environmental quality monitoring stations set up by the scientific research project team; the genetic susceptibility data is obtained by obtaining basic information through laboratory gene sequencing technology and then performing genotyping on genes and genetic markers in combination with genome-wide association studies GWAS; the medical diagnosis and treatment records are obtained through a medical history questionnaire and the hospital health registration system.
Citation Information
Patent Citations
Dynamic multi-level system modeling and state prediction method based on hybrid cognition
CN110489898A
Analysis and verification of models derived from clinical studies data extracted from a database
US20210183523A1